5 Best Practices in Concept Testing: Launch With Confidence
Concept testing exists to answer one question before money moves: will this specific product decision change what people actually buy, and by how much. The best practice isn't a smarter survey question, a bigger panel, or a faster AI simulation. It's replacing "would you buy this" with a randomized discrete-choice experiment that isolates which concept attribute moves purchase intent, with a confidence interval attached to the answer. A VP of product deciding whether to greenlight a launch needs that causal answer, not a top-box score with a margin of error nobody reports.
- The core practice: run a randomized discrete-choice experiment, not a rating-scale survey, so you can attribute intent shifts to a specific attribute rather than to sampling noise or question order.
- The metric to drop: top-box purchase intent, because roughly half of "definitely buy" respondents and only 10-20% of "probably buy" respondents actually purchase (Analysis Group, Befurt & Silk).
- The metric to add: a confidence interval on the effect of each concept attribute, scoped to the population you simulated or sampled, not the whole market.
- The speed tradeoff: AI synthetic panels solve the 2-3 week fielding cycle but substitute one unvalidated simulation for one uncontrolled survey unless the vendor publishes replication accuracy against real human studies.
- The launch gate: don't ship on a single concept score. Ship on a causal estimate of which attribute change moved intent, tested against a holdout.
Why does stated purchase intent fail as a launch signal?
Stated purchase intent fails because the scale you're scoring collapses a continuum of hesitation into a binary "will buy" signal that doesn't map to behavior. Quirks has flagged top-box scoring as marketing research's "top mistake," arguing that discarding the full scale in favor of top-2-box percentages throws away statistical power and reliability that the underlying continuum actually carries (Quirks). Analysis Group's Befurt and Silk quantified the gap directly: respondents who pick the top rating on a purchase-intent scale have roughly a 50% chance of actually buying, and respondents one notch down have only a 10-20% chance (Analysis Group). A concept that scores well on top-box is not a concept validated to sell. It's a concept that made people comfortable saying yes to a hypothetical.
This matters because the stakes are asymmetric. An often-cited estimate, attributed to Clayton Christensen and reported by Inc., puts new-product failure at roughly 95 percent of the 30,000-plus products launched each year; the underlying methodology isn't published, so treat the figure as a directional stake, not a precise rate (Inc.). If the pre-launch gate is a survey metric with a documented say-do gap, that gate is not doing the job it was built for.
What the traditional five-practice checklist gets right, and where it stops
The standard concept-testing checklist is not wrong, it's incomplete. Define objectives, size the sample at 100-300 respondents per concept, choose monadic versus sequential monadic exposure, test price and packaging alongside the core idea, and iterate early. Each of these is a real improvement to survey design. None of them touches the underlying instrument.
Sequential monadic designs, in particular, introduce order and fatigue effects that no amount of sample size fixes. A respondent's third exposure carries the residue of the first two. A bigger N gives you a tighter confidence interval around a biased number. That's the trap: better survey execution produces more precise estimates of the wrong thing. Stated intent, however cleanly measured, is a correlational proxy. It tells you what people say they'd do, not what changes their behavior when a specific attribute moves.
Survey, synthetic panel, or randomized experiment: which fits your launch decision?
| Approach | What it measures | Known weakness | Best for |
|---|---|---|---|
| Traditional stated-preference survey (monadic/sequential monadic, 100-300 respondents, 2-3 week field) | Self-reported purchase intent on a rating scale | Hypothetical bias: stated intent runs higher than revealed behavior (Wiley Health Economics); top-box conflates polite agreement with real intent (Quirks, Analysis Group) | Teams that need a directional read on message fit with no causal claim attached |
| AI synthetic-respondent panel (e.g., Aaru, Synthetic Users, Outset, Attest) | Simulated concept reactions from an LLM-based population, returned in minutes | Agreement with human benchmarks is claimed in the 80-95% range across vendor and industry reports, without a published protocol or holdout naming the studies; still a correlational read unless the underlying design randomizes attributes | Teams that need fast directional screening before committing to a fielded study |
| Randomized discrete-choice experiment (McFadden discrete choice, Mixed Logit, ICLV) | Which concept attribute causes purchase intent to move, with a confidence interval on the effect | The estimator identifies preference structure; causal identification comes from the randomized manipulation in the design, not from the model itself | Teams making a real launch or pricing decision that needs an attributable, not just correlated, answer |
Are AI synthetic panels a shortcut to validity?
No. They're a shortcut to speed. Platforms like Aaru, Synthetic Users, Outset, and Attest compress the traditional 2-3 week fielding cycle into minutes by simulating respondent reactions instead of collecting them. Published performance against human benchmarks varies by vendor, and few name the protocol behind the number. Agreement with a human panel on a stated-preference question is not a causal estimate. It's a faster version of the same instrument. If the underlying question is still "would you buy this," a synthetic panel returns a faster, cheaper version of a number that already has a known say-do gap. Speed is being solved faster than validity is. Before trusting any panel, ask the vendor for a published replication rate against real human studies and the failure mode when the market being tested wasn't in the training data.
The practice that isolates cause: randomized discrete-choice experiments
This is the one practice the standard checklists skip. Instead of asking respondents to rate a fixed concept, a discrete-choice experiment randomly varies concept attributes (price, feature set, positioning) across choice tasks and forces a trade-off, then estimates which attribute change moved the choice. The randomization is what earns the word "causal." McFadden discrete choice, Mixed Logit, and ICLV are the estimators used to recover preference weights from that data; they are not, by themselves, causal methods. The causal claim comes from the randomized manipulation in the experiment design, not from the choice model.
Mixed Logit also relaxes the independence-of-irrelevant-alternatives assumption that a flat multinomial logit imposes, which matters for any preference-share or substitution question. A standard logit will predict that a new concept variant steals share proportionally from every existing option, which is rarely how real substitution works.
How much replication accuracy is enough to trust a simulated concept test?
Enough is a published number tied to a public protocol, not a vendor's internal claim. Subconscious reports 93 percent replication accuracy, defined as how often a simulated study reproduces the direction and outcome of the original human study, measured against a validation set and published at go.subconscious.ai/paper. That number is a validation-set result, not a guarantee for a market you haven't tested yet, and it carries a real limitation worth stating plainly: some of the human studies used for validation were published before the underlying models were trained, so the replication protocol has to account for the possibility that a model has seen the original result rather than independently reproducing it. The full method comparisons, including how this stacks against other simulation approaches, are on the leaderboard and in the methods and validation hub. Ask any panel vendor claiming "80-95% agreement with human benchmarks" the same question: agreement on what protocol, against which holdout, published where.
Where willingness-to-pay estimates still need a caveat
Any willingness-to-pay number pulled from a concept test, stated-preference survey or discrete-choice experiment alike, runs high relative to what people actually pay unless the design is incentive-aligned, meaning respondents face a real consequence for their choice rather than a hypothetical one. This is documented hypothetical bias, and the direction is consistent: stated WTP overstates real WTP (Wiley Health Economics). A discrete-choice experiment doesn't eliminate this bias by default. It gives you a defensible way to isolate which attribute drives the WTP shift, which is a different and more useful claim than "respondents said they'd pay $12."
A launch-ready concept testing checklist
Before a concept moves from test to launch, a VP of product at the launch gate should be able to answer:
- What attribute was randomized, and what was held constant?
- What is the confidence interval on the attribute's effect, and what population does it cover?
- Was the willingness-to-pay estimate incentive-aligned, or does it carry the standard upward hypothetical bias?
- If a synthetic panel was used, what replication rate is published against real human studies?
- Does the preference-share estimate rely on a flat logit's IIA assumption, or does it use Mixed Logit to relax it?
A concept that clears every item on the standard five-practice list but can't answer these five questions has been surveyed, not tested. See how this plays out in practice in the comparisons hub and in published case studies.
Next step: take the concept variant you're least sure about and design one randomized choice task around its single most important attribute, before you run the next full survey wave. If you want a second read on the design, talk to us.