Skip to content

Survey scripting best practices in market research

CAVEMAN MODE ACTIVE

Fixed problems 1, 2, 3, 4, 5, 6, 7, 8. Cut unsourced completion/phone claims (not just numbers), cut Forsta vendor sentence, fixed Pew split-sample description, de-duplicated the PMC 89/52 finding across bullet and body, split the run-on, rewrote the opener to lead with the action decision, added the health-economics limitation into the figure caption itself. Everything else untouched.

A research director scoping a study has one real decision to make before scripting starts: clean up the wording of a single survey question, or build a randomized experiment that varies attributes and can support a causal claim. Survey scripting best practices (shorter instruments, mobile-first layout, randomized question and answer order) are worth doing, but they only improve the first option. They do nothing for the second: a single stated-preference question, however well scripted, was never built to isolate why people choose or which action would change their behavior. That takes a randomized, attribute-varying experiment, not a better-worded survey item.

- Scripting fixes the mechanics of a question: length, mobile-first layout, wording. It does not fix what the question measures.
- Randomizing question and answer order cancels primacy and recency effects, which is why Pew builds it into its own methodology ([Pew, Writing Survey Questions](https://www.pewresearch.org/writing-survey-questions/)).
- None of this fixes measurement validity: a systematic review and meta-analysis of discrete choice experiments in health found only 52% specificity against real-world choices, meaning even well-designed hypothetical questions misclassify roughly half of the people who would not actually choose what they said they would ([PMC, 2024](https://pmc.ncbi.nlm.nih.gov/articles/PMC11714376/)).
- The fix for a causal question is a randomized experiment analyzed with discrete choice models, not tighter wording.

## What good scripting actually fixes

Survey scripting guidance in 2026 runs on three levers, and all three are legitimate. Brevity is the standard recommendation: keep the instrument short. Mobile-first layout is the default case for scripting, not an edge case, since grid design, response scales, and load time all have to work on a phone screen. And order randomization, of both questions and answer categories, exists specifically because Pew's own methodology work shows order changes response distributions on its own ([Pew, Writing Survey Questions](https://www.pewresearch.org/writing-survey-questions/)).

The newer addition is AI embedded directly in the scripting layer to flag straightlining, speeding, and fraud as the respondent answers, rather than catching it in post-field cleanup. That's a real improvement on panel contamination. It's still a data-quality fix, not a measurement-validity fix.

All of it is aimed at the same target: a cleaner, faster, less biased version of the same instrument. None of it touches what that instrument is measuring.

## Does cleaner wording change what a respondent actually believes?

No. It changes what they report, which is a different thing. Pew ran the cleanest possible test of this: a split-sample experiment, same topic, one word different, with each respondent randomly assigned to only one wording. Asking whether "jobs" were plentiful got 60% agreement; asking about "good jobs" got 48%, a 12-point swing from a single qualifier ([Pew, 2019](https://www.pewresearch.org/short-reads/2019/01/29/good-jobs-vs-jobs-survey-experiments-can-measure-the-effects-of-question-wording-and-more/)). A well-scripted, unbiased, randomized-order version of either question would still produce a number that moves with phrasing. Scripting hygiene reduces noise in how a question is asked. It does not touch whether the answer reflects a stable underlying preference, because the question is still measuring framing, not behavior.

## Can a well-scripted question predict what people will actually do?

Not reliably. The best available evidence on this comes from health economics, where discrete choice experiments get checked against real uptake more often than in most other fields. The same systematic review and meta-analysis found 89% sensitivity alongside that 52% specificity ([PMC, 2024](https://pmc.ncbi.nlm.nih.gov/articles/PMC11714376/)): the experiments are good at catching people who really would choose the option, and close to a coin flip at correctly identifying the people who would not. No scripting change fixes that gap, because it isn't a wording problem.

![Bar chart comparing sensitivity at 89 percent and specificity at 52 percent for discrete choice experiments predicting real-world choice, from a systematic review and meta-analysis of health-related studies.](/images/authority/survey-scripting-best-practices-market-research.svg "A systematic review of discrete choice experiments in health found 89% sensitivity and 52% specificity against real-world choice; the finding is scoped to health-related studies and may not generalize to other categories.")

## Why stated preference runs high

This is the say-do gap, and it has a direction: stated preference is inflated relative to revealed behavior, not randomly noisy around it. The hypothetical bias literature documents this consistently in willingness-to-pay research, where stated WTP runs systematically higher than what people actually pay when the choice is real, and reviews the mitigation methods researchers use to try to close the gap ([ScienceDirect, hypothetical bias review](https://www.sciencedirect.com/science/article/abs/pii/S1755534521000555)). A respondent asked to imagine paying for something has no cost, no opportunity cost, and no memory of tradeoffs made under budget pressure. Scripting can pilot-test the wording of that question until it reads perfectly. It cannot give the respondent a reason to answer as if the choice were real.

## The decision this article owns

Scripting best practices and causal measurement are not the same investment, and treating them as one decision is the mistake. Discrete choice models like McFadden's multinomial logit, Mixed Logit, and ICLV are estimators. They tell you how to fit preference weights to choice data. They are not, on their own, causal methods. The causal identification comes from the randomized, attribute-varying design underneath the estimator: vary price, features, or messaging independently across a controlled experiment, and the resulting effect estimate can be read as causal. Run the same estimator on a single static question with no manipulation, and you get a preference weight with no causal claim attached to it.

This also matters for what the model can tell you afterward. A flat multinomial logit carries the independence of irrelevant alternatives (IIA) assumption, which breaks down when new, similar options enter a market. Mixed Logit relaxes that by letting preferences vary across the population; ICLV goes further and models the latent attitude driving the choice directly. Neither fix is about better wording. Both are about whether the underlying experiment was designed to support the question being asked.

![Decision path showing that fact and attitude questions can be answered with a well-scripted survey question, while causal questions about why people choose or which action drives an outcome require a randomized, attribute-varying experiment.](/images/authority/survey-scripting-best-practices-market-research-2.svg "A scripted question can tell you what people report; only a randomized experiment can tell you why they choose or which lever moves the outcome.")

## When is a scripted survey enough, and when do you need a randomized experiment?

A scripted question is enough when the buyer needs a fact, a self-reported attitude, or a tracking number over time, and is not trying to attribute a cause to it. It fails the moment the business question becomes "why did this happen" or "which action should we take," because a single static question has no manipulated variable to attribute the outcome to. That's the gap a randomized experiment closes: hold the population constant, vary the attribute of interest, hold everything else fixed, and the resulting difference in choice is attributable to the thing you varied.

## What validated causal simulation looks like

Subconscious runs randomized, attribute-varying experiments on a simulated population and checks them against real human studies. In that validation, 93 percent replication accuracy measures how often the simulated study reproduces the direction and outcome of the original human study, as reported at [go.subconscious.ai/paper](https://go.subconscious.ai/paper). That number is a validation-set result, not a guarantee for a new market. It carries the same caveat every model-based replication does: some of the published studies used to validate it may exist in a model's training data. That's exactly why the replication protocol treats direction-and-outcome matching, not memorization, as the bar. Independent studies and their replication results are tracked on the [leaderboard](/leaderboard). The methodology behind that scoring, along with the broader case for why estimator choice is not the same as causal design, is covered on the [methods and validation](/blog/methods-and-validation) hub.

Any confidence interval that comes out of a simulated experiment describes the effect within that simulated population under the tested conditions. It does not bound the real market unconditionally, and it should be read that way in any report that cites it.

If you're scripting a tracker right now, the smallest useful move is to separate the two questions on your own survey plan: mark which items are self-report or tracking (keep scripting them well) and which are trying to answer why or which action (flag those as candidates for a randomized experiment instead of a better-worded item). For a working example of how that split plays out on a real study, the [leaderboard](/leaderboard) has published replications to check against. When you want to see a randomized experiment run against your own market, [meet](/meet) with the team.