Skip to content

Understanding the margin of error in simulations

A VP of consumer insights deciding whether to trust a synthetic panel's read on a pricing decision needs one number before signing off: the real margin of error on the study about to run, not a resemblance score from someone else's study. True margin of error comes from the response variance inside the randomized experiment you design and field this quarter, and it shrinks predictably as you add sample, landing in a confidence interval the study itself can defend. A high correlation to a prior human survey describes how well a model reproduced old data. It carries no provable bound on the choice set, segment, or question you are about to test.

What does margin of error actually mean in a simulation?

It means the range around an estimated effect, at a stated confidence level, derived from response variance inside the specific randomized experiment you fielded. In a discrete choice study, that variance comes from how respondents (human or simulated) traded off attributes across the choice tasks you built. Feed more respondents into that same design and the interval narrows in a way you can compute and audit. That is what a margin of error is: a property of this study, not a compliment paid to a different one. A correlation coefficient against a past survey does not narrow with more respondents in your current study, because it is not measuring your current study at all.

Why a 0.90 correlation to last year's survey is not a margin of error

The most-cited validation case in the market right now is EY's partnership with Aaru, which recreated EY's 3,600-respondent Global Wealth Research Report and reported a median Spearman rank correlation of 0.90 across 53 questions, with an average RMSE of 7.1 percentage points (EY, "How AI simulation accelerates growth in wealth and asset management"). That number answers a real question: did the model's synthetic population resemble the 2024-vintage human sample closely enough to trust the reproduction. It does not answer the question a buyer actually has going into a new study: what is the error bound on the specific decision I am about to make. Resemblance to a fixed, published dataset does not move when you change the question, the population, or the choice set. It is fixed the moment the benchmark study is fixed.

The thin-segment failure: what resemblance scores hide

Buried inside that same 0.90 median was a -0.38 rank correlation on inheritance planning, a single topic where the simulation and the human sample diverged sharply while the aggregate score stayed reassuring (EY/Aaru). A separate study aggregating 29 real-world design-preference tests across 2,073 human participants found consistent, systematic discrepancies between LLM-simulated and real preferences, including position and order bias; the paper reports these discrepancies directionally rather than as a single effect size, so they are not directly comparable in magnitude to the correlation and divergence figures cited above ("Distorted Perspectives of LLM-Simulated Preferences," arXiv:2605.18311). Neither failure is a sample-size problem. A larger synthetic panel does not fix an order-bias artifact or a segment the model represents poorly; it just reproduces the bias with more decimal places. That distinction, bias versus variance, is the one a bare correlation number cannot make for you.

Individual-level prediction carries higher error than population-level prediction

A cross-domain benchmark comparing synthetic and human survey responses found aggregate-level Jensen-Shannon divergence of 0.011 to 0.046, versus 0.056 to 0.090 for single-answer, individual-level prediction on the same underlying data ("When Can Digital Personas Reliably Approximate Human Survey Findings?", arXiv:2605.10659). The gap between the two ranges is not a single ratio. Comparing the lower bounds, individual-level divergence runs more than 5x higher than aggregate-level divergence. Comparing the upper bounds, it runs close to 2x higher (0.090 versus 0.046). Either way, asking a simulation "what does this market do on average" and asking it "what does this one person do" are different questions with different error profiles.

Bar chart showing Jensen-Shannon divergence ranging from 0.011 to 0.046 for aggregate population predictions, versus 0.056 to 0.090 for individual-level predictions, on the same underlying survey data.
Individual-level divergence exceeds aggregate-level divergence throughout this study, from about 5x at the lower bound to about 2x at the upper bound.

That gap is a reason to ask what your study needs: a market-level read, where error is smaller, or a segment-level or individual-level read, where it is not. Neither question is answered by a single resemblance score against a past survey.

How does a randomized discrete choice experiment produce a real confidence interval?

It produces one because the interval comes from the variance in how respondents traded off attributes across the randomized choice tasks in that study, estimated with a discrete choice model. McFadden's multinomial logit, Mixed Logit, and ICLV are estimators applied to that randomized design; the causal identification comes from the randomization itself, not from the estimator. A standard multinomial logit carries the independence of irrelevant alternatives assumption, meaning it can distort preference-share and substitution estimates when alternatives are not genuinely independent; Mixed Logit relaxes that assumption by allowing preferences to vary across respondents, which matters when you are asking a substitution question rather than a simple main-effects question. Whichever estimator fits the design, the resulting confidence interval covers the effect within the population you simulated, in the study you ran. It does not extend that guarantee unconditionally to the real market, and any responsible read of the result says so.

Where do Subconscious's own validation numbers fit?

Subconscious's own validation number carries the same limitation as the others, but it is stated differently: as a ratio against a measured human ceiling, not a bare correlation to one external dataset. Subconscious's best configuration reaches 87% of the measured human ceiling on one study, a 0.832 rank correlation against the published human result, where two independent samples of real humans reach 0.959 against each other; across all 43 studies that passed design filters, the mean replication is 0.73 (Subconscious, Causal Fidelity paper). That figure is a validation result about how well the simulation reproduces known human studies. It is not a guarantee for a market you have not yet tested. Published benchmark studies can sit inside a model's training data; the replication protocol tests against studies designed to check for that, rather than assuming the problem away. The current model comparisons behind that number are tracked on the leaderboard, which updates as configurations change. The ratio against a measured human ceiling still does not substitute for a margin of error: that has to come from your own randomized design.

Resemblance score or margin of error: which one should a buyer trust for a decision?

Trust the benchmark comparison to screen a vendor, and trust the margin of error to bound the decision. The two answer different questions and neither substitutes for the other.

Resemblance score (e.g., Spearman correlation to a past survey)Margin of error (confidence interval from a randomized experiment)
What it measuresHow closely synthetic answers matched one prior human datasetThe range containing the true effect in the current study, at a stated confidence level
Derived fromA single external benchmark, often published and possibly present in training dataResponse variance inside the design you are running now
Behavior with sample sizeFixed once the benchmark study is fixed; does not shrinkShrinks predictably as sample size grows
What it can hideOrder bias, demographic flattening, thin-segment collapse (EY/Aaru: -0.38 rank correlation on inheritance planning)Bias in model or population construction, which a wider interval alone does not correct
Best for:Screening whether a vendor's simulation broadly matches known human behavior before you commitBounding the specific decision you are about to make on your own randomized experiment

For more on how validation studies are built and where their limits sit, see the methods and validation hub.

The concrete next step: before running a new study, ask any vendor two separate questions, not one. First, what is your resemblance score against a published human benchmark, and what is the denominator. Second, what confidence interval does my specific design produce, and what population does it cover. A vendor with a strong answer to the first question and no answer to the second is offering you a description of their last project, not a margin of error for yours. If you want to see how that second question gets answered on a live design, meet with the team.