Understanding the margin of error in simulations
A VP of consumer insights deciding whether to trust a synthetic panel's read on a pricing decision needs one number before signing off: the real margin of error on the study about to run, not a resemblance score from someone else's study. True margin of error comes from the response variance inside the randomized experiment you design and field this quarter, and it shrinks predictably as you add sample, landing in a confidence interval the study itself can defend. A high correlation to a prior human survey describes how well a model reproduced old data. It carries no provable bound on the choice set, segment, or question you are about to test.
- Margin of error is a property of a randomized experimental design. It comes from the variance in the study you are running now, not from how well a model matched a past one.
- A 0.90 Spearman rank correlation to a prior human survey, the most-cited case being EY and Aaru's recreation of a 3,600-respondent wealth study, is backward-looking validation. It says nothing provable about the next question you ask.
- Resemblance scores can hide bias instead of revealing it: the same EY/Aaru study posted a -0.38 rank correlation on inheritance planning even while the median across 53 questions hit 0.90.
- Predicting a single person's answer carries higher divergence than predicting the whole population's response pattern, on identical underlying data.
- Discrete choice models estimated on your own randomized design, McFadden multinomial logit, Mixed Logit, or ICLV, produce coefficient-level confidence intervals the study itself can defend, not ones borrowed from another team's data.
What does margin of error actually mean in a simulation?
It means the range around an estimated effect, at a stated confidence level, derived from response variance inside the specific randomized experiment you fielded. In a discrete choice study, that variance comes from how respondents (human or simulated) traded off attributes across the choice tasks you built. Feed more respondents into that same design and the interval narrows in a way you can compute and audit. That is what a margin of error is: a property of this study, not a compliment paid to a different one. A correlation coefficient against a past survey does not narrow with more respondents in your current study, because it is not measuring your current study at all.
Why a 0.90 correlation to last year's survey is not a margin of error
The most-cited validation case in the market right now is EY's partnership with Aaru, which recreated EY's 3,600-respondent Global Wealth Research Report and reported a median Spearman rank correlation of 0.90 across 53 questions, with an average RMSE of 7.1 percentage points (EY, "How AI simulation accelerates growth in wealth and asset management"). That number answers a real question: did the model's synthetic population resemble the 2024-vintage human sample closely enough to trust the reproduction. It does not answer the question a buyer actually has going into a new study: what is the error bound on the specific decision I am about to make. Resemblance to a fixed, published dataset does not move when you change the question, the population, or the choice set. It is fixed the moment the benchmark study is fixed.
The thin-segment failure: what resemblance scores hide
Buried inside that same 0.90 median was a -0.38 rank correlation on inheritance planning, a single topic where the simulation and the human sample diverged sharply while the aggregate score stayed reassuring (EY/Aaru). A separate study aggregating 29 real-world design-preference tests across 2,073 human participants found consistent, systematic discrepancies between LLM-simulated and real preferences, including position and order bias; the paper reports these discrepancies directionally rather than as a single effect size, so they are not directly comparable in magnitude to the correlation and divergence figures cited above ("Distorted Perspectives of LLM-Simulated Preferences," arXiv:2605.18311). Neither failure is a sample-size problem. A larger synthetic panel does not fix an order-bias artifact or a segment the model represents poorly; it just reproduces the bias with more decimal places. That distinction, bias versus variance, is the one a bare correlation number cannot make for you.
Individual-level prediction carries higher error than population-level prediction
A cross-domain benchmark comparing synthetic and human survey responses found aggregate-level Jensen-Shannon divergence of 0.011 to 0.046, versus 0.056 to 0.090 for single-answer, individual-level prediction on the same underlying data ("When Can Digital Personas Reliably Approximate Human Survey Findings?", arXiv:2605.10659). The gap between the two ranges is not a single ratio. Comparing the lower bounds, individual-level divergence runs more than 5x higher than aggregate-level divergence. Comparing the upper bounds, it runs close to 2x higher (0.090 versus 0.046). Either way, asking a simulation "what does this market do on average" and asking it "what does this one person do" are different questions with different error profiles.
That gap is a reason to ask what your study needs: a market-level read, where error is smaller, or a segment-level or individual-level read, where it is not. Neither question is answered by a single resemblance score against a past survey.
How does a randomized discrete choice experiment produce a real confidence interval?
It produces one because the interval comes from the variance in how respondents traded off attributes across the randomized choice tasks in that study, estimated with a discrete choice model. McFadden's multinomial logit, Mixed Logit, and ICLV are estimators applied to that randomized design; the causal identification comes from the randomization itself, not from the estimator. A standard multinomial logit carries the independence of irrelevant alternatives assumption, meaning it can distort preference-share and substitution estimates when alternatives are not genuinely independent; Mixed Logit relaxes that assumption by allowing preferences to vary across respondents, which matters when you are asking a substitution question rather than a simple main-effects question. Whichever estimator fits the design, the resulting confidence interval covers the effect within the population you simulated, in the study you ran. It does not extend that guarantee unconditionally to the real market, and any responsible read of the result says so.
Where do Subconscious's own validation numbers fit?
Subconscious's own validation number carries the same limitation as the others, but it is stated differently: as a ratio against a measured human ceiling, not a bare correlation to one external dataset. Subconscious's best configuration reaches 87% of the measured human ceiling on one study, a 0.832 rank correlation against the published human result, where two independent samples of real humans reach 0.959 against each other; across all 43 studies that passed design filters, the mean replication is 0.73 (Subconscious, Causal Fidelity paper). That figure is a validation result about how well the simulation reproduces known human studies. It is not a guarantee for a market you have not yet tested. Published benchmark studies can sit inside a model's training data; the replication protocol tests against studies designed to check for that, rather than assuming the problem away. The current model comparisons behind that number are tracked on the leaderboard, which updates as configurations change. The ratio against a measured human ceiling still does not substitute for a margin of error: that has to come from your own randomized design.
Resemblance score or margin of error: which one should a buyer trust for a decision?
Trust the benchmark comparison to screen a vendor, and trust the margin of error to bound the decision. The two answer different questions and neither substitutes for the other.
| Resemblance score (e.g., Spearman correlation to a past survey) | Margin of error (confidence interval from a randomized experiment) | |
|---|---|---|
| What it measures | How closely synthetic answers matched one prior human dataset | The range containing the true effect in the current study, at a stated confidence level |
| Derived from | A single external benchmark, often published and possibly present in training data | Response variance inside the design you are running now |
| Behavior with sample size | Fixed once the benchmark study is fixed; does not shrink | Shrinks predictably as sample size grows |
| What it can hide | Order bias, demographic flattening, thin-segment collapse (EY/Aaru: -0.38 rank correlation on inheritance planning) | Bias in model or population construction, which a wider interval alone does not correct |
| Best for: | Screening whether a vendor's simulation broadly matches known human behavior before you commit | Bounding the specific decision you are about to make on your own randomized experiment |
For more on how validation studies are built and where their limits sit, see the methods and validation hub.
The concrete next step: before running a new study, ask any vendor two separate questions, not one. First, what is your resemblance score against a published human benchmark, and what is the denominator. Second, what confidence interval does my specific design produce, and what population does it cover. A vendor with a strong answer to the first question and no answer to the second is offering you a description of their last project, not a margin of error for yours. If you want to see how that second question gets answered on a live design, meet with the team.