Synthetic Respondents: What They Are and When to Trust One
A synthetic respondent is a model-generated stand-in for a participant in a survey or experiment. The decision for a research lead is whether the model has been validated for the population, task and outcome their team needs. Plausible answers are useful for exploring a question. A decision about real spending needs evidence that the estimates agree with human results on a relevant evaluation, with uncertainty and failures reported.
- Define the decision before comparing methods: describing an attitude and estimating the effect of an intervention require different evidence.
- Ask which population, task, sample and scoring rule a validation result covers.
- Check whether the evaluation was held out and how training-data leakage was assessed.
- Subconscious's public working paper reports mean rank correlation of 0.55 across roughly 300 replicated studies and 0.73 across 43 studies passing design filters. The filtered result describes that cohort, not every new customer decision (causal fidelity paper).
- Validate a proposed study against the decision it must support, then state when human research is still needed.
What is a synthetic respondent, and when should you trust one?
A language model can be conditioned on a demographic or behavioral profile and asked to answer a survey or discrete choice task. Other systems can incorporate observed behavior or additional models. The category label alone does not tell a buyer how the system was built or what it has demonstrated.
Trust depends on the intended use. For early question development, a generated response may expose ambiguities worth investigating. For a price, product or policy decision, ask for a relevant human comparison and the full evaluation protocol. A convincing demonstration is not sufficient evidence that an estimate will transfer to your market.
What does independent research say about synthetic respondents?
The evidence depends on the task. John Horton's NBER working paper investigates language models as simulated economic agents and reports qualitative parallels with behavioral experiments. That result does not establish reliable magnitudes for every new decision (NBER working paper 31122).
Bisbee, Clinton, Dorff, Kenkel and Larson found that synthetic survey responses could understate human variation and change with prompt wording and model version. Those failures can make an estimate appear more precise than its human comparison warrants (Political Analysis, 2024). Read validation by task and outcome rather than treating fluency as evidence of accuracy.
What must an accuracy number tell you?
An accuracy number needs a scoring rule, a study set and a human comparison. Agreement on answer distributions, recovery of treatment effects and prediction of purchase choices measure different things. A vendor-run study can provide useful evidence if its design, comparator and results are available; who ran it does not by itself settle whether it is valid.
A human-to-human agreement estimate also belongs to a particular study, sample size and scoring method. It is not a universal maximum for other populations or tasks. Compare like with like and ask how the result changes under alternative samples and scoring choices. The methods and validation hub explains the design and estimator questions behind those comparisons.
What does Subconscious's public replication evidence show?
The working paper scores Spearman rank correlation on estimated choice parameters. It reports a mean of 0.55 across roughly 300 replicated studies and 0.73 across the 43 studies that passed its design filters. Filtering changes the evaluated cohort; the higher number is not a universal improvement claim. Read the eligibility criteria, the scores and the misses together (causal fidelity paper).
Published studies can appear in a language model's training data. A replication score alone cannot distinguish generalization from recall. A buyer should ask how the protocol investigates contamination and which held-out evaluations support transfer. The leaderboard and methodology case study provide the existing evidence path.
How should a buyer compare validation protocols?
Ask for the evidence behind the intended use rather than assigning a method a single accuracy label.
| Intended use | Evidence to request | Decision it can support |
|---|---|---|
| Explore wording or possible reactions | Examples, known failure modes and human review of the questions | Which questions deserve further research |
| Estimate survey response distributions | Human comparisons for the population, instrument and response distribution | Whether the estimates reproduce the measured responses |
| Estimate the effect of an intervention | Identified experimental design, relevant human effects and uncertainty | Whether an option changes the stated outcome in the tested setting |
| Predict behavior in a new market | Held-out outcome comparisons and evidence of transfer | Whether the result applies to the population and decision being considered |
A product may address more than one row. Ask which uses it has evaluated, what it measured and which conclusions the result supports.
Where caveats still apply
A statistical estimator does not establish a causal effect on its own. Identification depends on the design and assumptions. In a randomized experiment, the assigned intervention supplies the comparison; a simulation still needs evidence that its responses transfer to the human decision being studied.
Discrete choice estimates also depend on the model specification. A flat logit model carries the independence-of-irrelevant-alternatives assumption, which can distort switching between similar options. Mixed Logit relaxes that assumption, with additional parameters to estimate. Hypothetical willingness-to-pay can differ from actual spending when choices have no real stakes.
A confidence interval within a simulated population describes uncertainty under that model and design. It does not include every possible source of error in the real market. Ask how sensitivity to prompts, model versions, population definitions and missing behavior was tested.
For a proposed purchase, bring the decision, options, population and consequences of a wrong answer to a decision review. Request a study design, a relevant validation comparison and a clear account of what would still require human research before committing budget.