Synthetic Responses vs. Causal Experiments: Choosing a Market Research Tool
Choose a market research tool by the evidence it produces for the decision at hand. A synthetic response helps a team explore what people might say; a controlled experiment fits better when deciding whether a price, position, or launch message will change an outcome. The procurement question is which result was checked against human behavior, not which report sounds most convincing.
Fluency is not behavioral evidence
A simulated respondent can give a coherent explanation of a preference, but that output reflects a model's prediction, not what people will do when they face a real price, tradeoff, or switching cost.
Research on large language models as virtual survey respondents finds generated answers can resemble aggregate human patterns for some questions and diverge for others (Large Language Models as Virtual Survey Respondents). A review of model limitations documents the gap between fluent output and reliable grounding (LLLMs: A Data-Driven Survey of Evolving Research on Limitations of Large Language Models).
"Our evaluation reveals consistent performance trends across model families, highlights failure modes in structured output generation, and demonstrates how context and prompt design affect simulation fidelity."
Zhao and colleagues, "Large Language Models as Virtual Survey Respondents," arXiv:2509.06337 (source)
That gap matters before a launch: a plausible panel can support the wrong price or message just as confidently as the right one, and the cost appears only after budget or roadmap capacity is committed.
How does each approach's evidence chain compare?
The useful comparison is how each approach connects a proposed action to a measurable outcome.
| Buyer question | Synthetic-response approach | Causal experiment with human-baseline validation |
|---|---|---|
| What is being estimated? | A plausible stated reaction from a simulated respondent | The change in an outcome when one controlled alternative replaces another |
| What should the buyer inspect? | Whether the response was checked against relevant human behavior | How the experiment's result was compared with a matched human study |
| What reduces uncertainty before commitment? | A separate human check designed around the same decision | A buyer-configured real-human test or validation step that preserves the causal question |
| What does a strong result mean? | The response is useful for exploration, within the limits of the model and prompt | The tested alternative moved the defined outcome under the study conditions |
What replication benchmark should you ask for?
Our best configuration reaches 87% of the measured human ceiling on one study: 0.832 rank correlation against the published human result, where two independent samples of real humans reach 0.959. Across all 43 studies that pass design filters the mean is 0.73, drawn from roughly 300 replicated human studies across 9 domains. Naming the failure mode lets a buyer check it before committing budget. It is a validation result, not a guarantee for a new market. See the causal fidelity paper. Buyers can inspect the research behind the benchmark and the individual results on the replication leaderboard.
A number without its limits is marketing. This benchmark supports a specific claim about reproducing studied human outcomes, not a guarantee for every market, segment, or decision, and it remains separate from audience reach: an audience graph is not a recruitable participant panel.
When more certainty is needed, a simulated experiment can move to real-human testing or validation as a distinct, buyer-configured step that keeps the same alternatives and outcome rather than substituting a different study.
How should you compare cost and scope together?
Subconscious does not publish a self-service price list. Compare total cost for the specific decision rather than assuming a public price.
A historical planning example for agency-led research used three to four weeks for the brief, participant recruitment, and final report, treated as historical calendar context, not a current Subconscious estimate or a promise from any vendor.
The misses sit on the same leaderboard as the hits. Not every study includes confidence intervals or segment-level heterogeneity by default. Real-human validation is configured separately, not automatic. A validation result does not become an observed usability session, clinical trial, or automatic proof of market performance.
Bring one decision into the vendor conversation
Use one upcoming pricing, positioning, or launch decision to make this concrete:
- Define the alternatives that could actually ship.
- Name the behavioral outcome that would change the decision.
- Ask whether the method estimates a causal difference or generates a plausible reaction.
- Request the human baseline, replication definition, study count, and domain coverage behind any accuracy claim.
- Confirm whether real-human testing preserves the same alternatives and outcome.
- Ask which uncertainty measures and segment analyses are included for this study.
- Price the complete decision path, including any separate human validation.
The causal workflow shows how alternatives, outcomes, and validation fit together. To evaluate the method against a live decision, book a working session with those specifics ready.