Skip to content

What reviewers should ask about synthetic market research data

A research director vetting a synthetic-respondent vendor has one decision to make: sign off on the data or send it back. That decision turns on a single question: can the vendor produce a calibration study, a case where their synthetic estimate was checked against a real human study and came with an error number attached? Ask for that study and that number before anything else on the checklist. Buyers are being asked to sign off on this data with no shared standard for what "validated" means, even as synthetic respondents move from conference-talk novelty to a paid line item at research vendors. If a vendor can't produce a calibration study, the rest of the standard checklist (panel size, LLM version, prompt design) is secondary.

Why representativeness checklists test the wrong thing

The conventional buyer checklist asks about panel provenance, demographic balance, model transparency, and data handling, the questions ESOMAR's buyer-vetting guidance also centers. That's the right list for vetting a traditional panel and the wrong list for vetting a causal claim. A synthetic sample can match census age, income, and geography and still get the underlying relationship backwards. That's not hypothetical: it's what a 2024 study in Political Analysis found when researchers compared ChatGPT-generated survey responses to real ones on matched questions, where the demographics lined up but nearly half the coefficients didn't (Political Analysis). Representativeness is a property of the sample. Causal accuracy is a property of how the model responds to a manipulation, an intervention, a price change, a feature add. Those are different things to test, and a checklist built for the first one won't catch a failure in the second.

Why do synthetic panels get demographics right and the causal effect wrong?

They get demographics right because matching a population's composition is a sampling problem, and they get causal effects wrong because predicting how that population would respond to a specific change is a different, harder problem that most LLM-based panels have never been checked against. Greenbook's replication of the Paxton and Yang benchmark makes the gap concrete: general-purpose models like GPT and Gemini deviated from human survey answers by 0.87 to 0.88 standard deviations, while a model calibrated specifically for this task deviated by 0.07, in this single replication of one benchmark, and the uncalibrated models systematically inflated top-two-box scores and correlations between variables (Greenbook). That inflation matters for a buyer because it doesn't just add noise, it adds a consistent upward bias that would make a mediocre concept look validated.

Bar chart showing general-purpose LLMs deviating from human survey data by 0.87 to 0.88 standard deviations, compared to 0.07 standard deviations for a calibrated model.
A calibrated model deviated from real human survey answers by 0.07 standard deviations; general-purpose LLMs deviated by 0.87 to 0.88, in Greenbook's replication of the Paxton and Yang benchmark.

What should you actually ask a synthetic research vendor?

Ask for a calibration study: a specific instance where the vendor ran the same randomized experiment on a synthetic population and on a real human baseline, and report the error between the two. Subconscious reports a 93 percent replication accuracy, meaning that in its validation set, a simulated study reproduced the direction and outcome of the matched human study 93 percent of the time (go.subconscious.ai/paper). That number describes a validation set, not a guarantee for a market you haven't tested yet, and it's worth pressing any vendor on the same distinction: what set was this measured on, and how often is it re-run. It's also worth asking directly whether any of the published studies used for calibration could have been in the underlying model's training data, since a model that has memorized a study's published result isn't demonstrating replication, it's demonstrating recall. A credible vendor will have a protocol for that risk, not a denial that it exists.

Two ways to vet a synthetic sample

Representativeness checklistCalibration study
What it testsWhether the sample's demographics, panel size, and sourcing resemble the target populationWhether the model's estimate of a causal effect matches a real human baseline on a held-out study
What it can missA sample can match census figures and still produce a coefficient with the wrong signDoesn't catch demographic gaps in the underlying panel used to build the calibration
Example findingGeneral-purpose LLM panels built to resemble a population still deviated from human answers by 0.87 to 0.88 standard deviations in one benchmark replication ([Greenbook](https://www.greenbook.org/insights/data-science/testing-synthetic-data-against-academic-benchmarks-a-replication-study))48% of coefficients from LLM-generated survey data differed significantly from human data, sign flipped in 32% of those cases, in one comparison built on a single model's outputs ([Political Analysis](https://www.cambridge.org/core/journals/political-analysis/article/synthetic-replacements-for-human-survey-data-the-perils-of-large-language-models/B92267DC26195C7F36E63EA04A47D2FE))
Best forScreening a vendor's basic legitimacy and data handlingDeciding whether to trust a specific causal claim before you act on it

Neither check replaces the other. A vendor should pass both, but only the second one tells you whether the number they're handing you is safe to make a decision with.

Where is it safe to trust a synthetic respondent today?

It's safest for early-stage, upstream work, screening concepts, drafting a questionnaire, exploring which messages are worth testing further, and riskiest for confirmatory claims a real budget decision depends on. Sarstedt and coauthors, writing in Psychology & Marketing, make this the center of their guidance: silicon samples are defensible for pretesting and pilot studies, and risky when treated as a substitute for a confirmatory study (Wiley). That maps onto the calibration question directly. A vendor with no calibration study might still be useful for a first pass at message testing. The same vendor without a calibration study has no business producing the number that decides which action a team takes with real money behind it.

What a calibration number needs to specify

A calibration number is only as useful as what it says it covers. A confidence interval from a simulated experiment describes the effect within that simulated population; it doesn't bound the real market unless the vendor can show the simulated population tracks the real one on the dimensions that matter for this decision. If a substitution or preference-share question runs through a standard multinomial logit, it carries the independence of irrelevant alternatives assumption, meaning adding or removing an option shouldn't change the relative odds between the others, an assumption that breaks down often enough in real markets that it's worth asking whether the vendor's estimator relaxes it. Mixed Logit does; a flat logit doesn't. And the estimator itself (McFadden discrete choice, Mixed Logit, ICLV) isn't what makes a claim causal. Identification comes from the randomized manipulation built into the experiment design; the estimator just describes the response once the intervention has run.

Checking the vendor's track record before you sign

Ask to see the vendor's calibration results across more than one study, not just the case study they lead with in a sales deck. Subconscious publishes its accuracy scores across studies on a public leaderboard, which is the kind of artifact worth requesting from any vendor claiming causal validity: a running, checkable record rather than a single anecdote. Compare it against how the vendor talks about their own numbers. A vendor that reports a single flattering case study without a holdout methodology behind it is asking you to trust an anecdote. A vendor that publishes ongoing accuracy against human baselines, and names the limitation on the studies where it fell short, is giving you something you can actually audit.

The concrete next step: before your next synthetic research engagement, ask the vendor for their calibration study, the human baseline it was checked against, the error number, and whether that number covers the class of decision you're about to make. If they can't produce it, treat the output as a pretest, not a confirmatory result. More on how the underlying replication protocol works is on the methods and validation page. If you want to walk through a specific decision with someone who can show the calibration behind it, reach out.