What to Demand Before You Trust a Synthetic-Respondent Vendor's Accuracy Claim
A vendor's self-reported accuracy number is a marketing claim until you can reproduce it. The question is not what number the vendor publishes, but what evidence backs it and whether you can replicate it against your own historical data. It has to survive test-retest checks, item-level scrutiny, and an independent replication before a decision leans on it.
Why the aggregate number alone is not proof
A single correlation figure between synthetic and real respondent distributions can hide a wide spread. A platform that reports one clean number might be averaging a long tail of low-agreement items against a smaller set of near-perfect ones, a different pattern than a methodology that holds steady across every item. Treat any headline accuracy figure as a starting point for questions, not the answer.
What to ask a vendor before you believe the number
| Evidence type | What it shows | Red flag if the vendor can't produce it |
|---|---|---|
| Test-retest reliability | Whether the same panel run twice against the same stimulus produces the same result | Only reports one-time accuracy, no repeat-run data |
| Item-level correlation | Whether accuracy holds question by question, not just averaged across a study | Reports only a single aggregate figure |
| Independent replication | Whether the method reproduces a study the vendor didn't design or select | Cites only its own internal validation studies |
| Boundary conditions | Which question types, audiences, and decision contexts the accuracy claim does not cover | Presents the accuracy figure as universal, with no stated limits |
A vendor that reports aggregate accuracy without test-retest reliability data is reporting half the story: reliability is the precondition for any accuracy claim being meaningful.
What the published research actually supports
The academic literature on LLM-conditioned synthetic respondents is worth reading directly, not secondhand from a vendor's own summary of it. Argyle, Busby, Fulton, Gubler, Rytting, and Wingate pulled respondent demographics from the American National Election Studies, built those into backstories for a language model, and found the resulting response distributions correlated strongly with survey distributions across political-attitude batteries, including within demographic subgroups (Argyle et al., 2023, Political Analysis).
A related NBER working paper by Horton tested whether a language model conditioned on agent profiles would reproduce classic behavioral-economics results: ultimatum games, social-preference tasks, willingness-to-pay measures. The synthetic agents reproduced the qualitative findings, with quantitative effect sizes landing within roughly 10 to 20 percent of the published human baselines across most experiments.
A 2024 Political Analysis paper by Bisbee, Clinton, Dorff, Kenkel, and Larson tested whether synthetic respondents could replicate full published survey studies, not just isolated question batteries. Replication held up well on conventional stated-preference questions, and the accuracy gap widened on questions where the real human distribution was itself unusual: heavy-tailed, bimodal, or tied to genuinely novel behavior.
That finding matters more than any headline correlation number: the research consensus is not "synthetic respondents are accurate," it's "synthetic respondents are accurate on conventional stated-preference questions and measurably less accurate outside that zone."
Where the accuracy gap is largest
Across the published studies, the pattern is consistent:
- Novel-behavior prediction. Questions about new product categories or behaviors outside what a language model has meaningful signal about show the widest gaps, in some published comparisons 30 to 50 percent off the human baseline.
- Niche audiences with thin public signal. Accuracy depends on the model having seen enough data about a population, and small, specialized B2B roles or industries are where that signal thins out.
- Regulatory and compliance-substantiation studies. Synthetic data is not a substitute for real-human-respondent data on the record, regardless of how strong the correlation is elsewhere.
- High-stakes, real-consequence decisions. A synthetic respondent answers a hypothetical question; a real respondent facing real consequences is not measuring the same thing, even when the topic is identical.
How to run the validation yourself
A procurement team does not have to take any vendor's word for its accuracy number. The workflow is the same regardless of platform:
- Find a past study in your own archive with a known outcome distribution, ideally built on stated-preference methods, such as a concept test or message test.
- Recreate the demographic, role, and segment specifications that defined the original sample.
- Run the equivalent question battery using the same stimuli and framing as the original study.
- Compare the new distribution to the original, both in aggregate and item by item.
- Decide from your own numbers, not the vendor's marketing page, whether the method is accurate enough for the decision.
This is the only version of "accuracy" that should carry weight in a purchase decision: your data, your replication, your result.
Where a causal-experiment approach fits
Subconscious's position is that validation belongs inside the experiment design, not bolted on afterward as a marketing statistic. A randomized experiment is built against a simulated population with a human baseline and holdout built in, rather than validated once and reused as a blanket claim across every future study. Subconscious can test or validate studies with real human participants, so a team can move from a simulated run to real-human validation on the same causal question, without redesigning the study.
Limitations
Simulation is not a substitute for real-human validation on high-stakes, regulatory, or commitment-context decisions. No vendor's accuracy figure should be accepted as proof for a specific question type or audience without independent replication against your own data. Current replication results are published on the leaderboard.
Next step
Before signing with any synthetic-respondent vendor, ask for test-retest reliability data, item-level correlation, and one study you can independently replicate against your own historical results. If a vendor can't produce all three, start a conversation about which of your decisions are ready for a controlled comparison instead.