Synthetic Respondents: What They Are and When to Trust One
A synthetic respondent is a language model conditioned on a demographic or behavioral profile, then asked to answer survey and experiment questions the way that profile of person would. The decision a research lead faces is whether a given vendor has proven, against real human behavior, that its output can carry a real decision. Trust a synthetic respondent when its output has been scored in a blind replication of a published randomized experiment, against a measured human-to-human ceiling, with the misses published alongside the wins.
- A synthetic respondent is a model-generated stand-in for a survey or experiment participant, built to answer as a defined population would.
- Accuracy claimed against a vendor's own survey tests imitation. The vendor wrote both the test and the answer key.
- The credible standard is blind replication of a published randomized experiment, scored against a measured human-to-human agreement ceiling of 0.959.
- Subconscious's best configuration reaches 0.832 rank correlation on the Hainmueller immigration conjoint - 87% of the 0.959 ceiling measured on that same study - and separately averages 0.73 across the 43 published studies that passed design filters (causal fidelity paper).
- Independent research is mixed: LLM agents reproduce qualitative economics results but understate variance and shift with prompt wording.
What is a synthetic respondent, and when should you trust one?
A synthetic respondent is a language model prompted with a persona - age, income, prior purchase behavior, stated preferences - and asked to complete a survey or a discrete choice task as that person. Every vendor in this category makes some version of that model. The difference between them is what happens after the model answers.
Trust follows from validation. A synthetic respondent that writes fluent, plausible answers has proven only that the model can write fluent, plausible answers. Trust begins when a vendor replicates a published randomized experiment blind: the study is held out, meaning the model was not scored against it before this run, and the vendor reports the correlation between the simulated result and the real one. A buyer who skips that step has evaluated a demo. It has not been validated as an instrument for a real decision.
How did the market get here?
The category spans several distinct products built for different jobs. Simile and Aaru ground simulated markets in real behavioral and trading data rather than persona role-play alone. Synthetic Users and similar tools generate personas for qualitative product research, mostly aimed at directional usability findings rather than causal estimates. Survey incumbents including Qualtrics and SurveyMonkey have added AI panel features on top of existing survey infrastructure, positioned as a faster version of the same product they already sell.
None of this is dishonest by default. It means the category label "synthetic respondents" covers products built for different jobs, validated to different standards, and a buyer has to ask which standard applies before comparing prices.
What does independent research say about synthetic respondents?
Independent research supports the idea that LLM agents can reproduce real behavioral patterns, and separately warns that the same agents distort the details a serious decision depends on. John Horton's NBER working paper found that LLM agents conditioned on demographic profiles reproduced the qualitative findings of classic behavioral-economics experiments, though he cautions against trusting the exact magnitudes (NBER working paper 31122).
Bisbee, Clinton, Dorff, Kenkel, and Larson, publishing in Political Analysis, found the opposite failure mode in a different task: synthetic feeling-thermometer responses understated the variance present in real human samples, which produced artificially precise standard errors, and the results shifted with prompt wording and model version (Political Analysis, 2024). A model can get the direction of an effect right and still get its precision wrong in a way that misleads a buyer who reads only the point estimate.
A more recent result narrows the gap for a specific task. Maier and colleagues report that semantic similarity rating reached 90% of human test-retest reliability on purchase-intent replication, with response distributions that looked realistic rather than compressed (arXiv, 2025). Read together, these three results show real directional capability alongside real limits in variance and prompt sensitivity. A buyer should ask which of these failure modes applies to the specific decision at hand.
Why a self-reported accuracy number does not prove anything
The conventional pitch in this category is a single accuracy percentage, benchmarked against a survey the vendor ran itself. That number answers a narrow question: can the model imitate the answer the vendor already expected. It does not answer whether the model predicts behavior in a study the vendor did not design and was not scored against ahead of time.
A published human-to-human ceiling exists precisely to keep this claim honest. When two independent samples of real humans answer the same study, they agree with each other at 0.959 rank correlation (causal fidelity paper). No simulation should be reported as beating that number, because it means the simulation is claiming to agree with humans more than humans agree with themselves. Any vendor number that arrives without a stated human ceiling and a stated study set is not comparable to anything.
What does blind replication against a human ceiling actually measure?
Subconscious's own validation follows this discipline. On its best configuration, run against the Hainmueller immigration conjoint, it scores 0.832 rank correlation against the published human result, 87% of the 0.959 human-to-human ceiling measured on that same study. Across the 43 published studies that passed design filters, the mean drops to 0.73 (causal fidelity paper). Those numbers, and every miss behind them, are published on the leaderboard rather than filtered down to a favorable headline.
One limitation applies to any replication protocol in this category, including this one: published studies can sit inside a model's training data, so a strong replication score is not automatically proof the model reasoned rather than recalled. Design filters and holdout selection lower that risk. They do not remove it, and a buyer evaluating any vendor's leaderboard should ask what the filter excludes.
How do you tell a validated simulation from roleplay?
Ask for the protocol before the pitch: which published studies were replicated blind against a stated human baseline, and whether the losses are public. The comparison below maps the honest differences across the category.
| Category | Grounded in | What's published | Best for |
|---|---|---|---|
| Prediction/market simulators (Simile, Aaru) | Real market and trading behavior | Backtests against historical market outcomes | Buyers forecasting market-level outcomes with strong historical analogues |
| Persona-based synthetic panels (Synthetic Users) | Generated personas, no stated human baseline | Qualitative usability findings | Early-stage product teams running directional usability checks |
| Survey incumbents' AI panel add-ons (Qualtrics, SurveyMonkey) | Existing survey infrastructure | Vendor-run comparisons against their own survey | Teams already running traditional surveys who want a lower-cost add-on |
| Causal replication protocols scored against a human baseline | Blind replication of published randomized experiments | Full leaderboard, including misses | Buyers who need a causal estimate defensible to a client or a regulator |
A vendor that cannot answer these questions, or that only offers a comparison against its own survey, belongs in one of the top three rows. See case studies for examples of the replication protocol applied to specific decisions, and the methods and validation hub for the underlying estimators.
Where caveats still apply
Randomized experiments, not the estimator, are what produce a causal claim. McFadden discrete choice, Mixed Logit, and ICLV are estimators used to analyze a randomized experiment's results; identification comes from the manipulation in the design, not from the statistical model layered on top. When a discrete choice task uses a flat logit specification, preference-share and substitution results carry the independence-of-irrelevant-alternatives assumption, which can distort estimated switching between similar options; Mixed Logit relaxes that assumption at the cost of more parameters to estimate.
Willingness-to-pay estimates from any hypothetical choice task run high relative to what people actually pay, unless the design is incentive-aligned - meaning respondents' choices carry real stakes. A confidence interval produced from a simulated experiment covers the estimated effect within the simulated population under study, not the real market unconditionally; treat it as a bound on the model's estimate, not a guarantee about the world.
None of these caveats argue against using synthetic respondents. They argue for asking a vendor to state them. A vendor unwilling to name its own method's limits has skipped the validation protocol.
A concrete next step: pull the study list behind any vendor's accuracy claim, check whether it includes a stated human-to-human ceiling and a denominator, and check the leaderboard for a protocol that publishes both the hits and the misses before you buy against a real decision. If you want to walk through how the replication protocol applies to your specific market, get in touch.