Checking a Synthetic Experiment Result Before You Act on It
The decision: trust the number, or check it first
A synthetic discrete-choice experiment just told a research or insights leader which price, message, or feature wins. The next step commits budget: a launch, a pricing change, a positioning bet. Before that spend goes out, the question is whether this specific result is close enough to how real customers actually choose, or whether it needs an independent check first.
Getting this wrong is expensive in a specific way: the decision looks well-supported at the time, and the gap only surfaces after the budget is committed and the launch has shipped.
What does checking against published research look like?
Subconscious designs synthetic experiments to be checked against independent, published choice research, not to stand on their own. A systematic review and meta-analysis of discrete choice experiments in health-related research reports agreement rates between stated-preference choices and real-world revealed-preference outcomes (PMC, National Library of Medicine), the same kind of prediction-accuracy comparison used here, not a bespoke internal metric invented for one vendor's dashboard.
"Pooled sensitivity and specificity estimates were 89% (95% CI:77–95, I 2 = 97%) and 52% (95% CI:32–72, I 2 = 95%), respectively. The area under the SROC curve (AUC) was 0.81 (95% CI:0.77–0.84)."
Zhang and colleagues, eClinicalMedicine (systematic review and meta-analysis of discrete choice experiments in health-related research) (source)
Our best configuration reaches 87% of the measured human ceiling on one study: 0.832 rank correlation against the published human result, where two independent samples of real humans reach 0.959. Across all 43 studies that pass design filters the mean is 0.73. See the causal fidelity paper.
That figure is aggregate, not a per-run guarantee: it is measured across a body of replicated studies, not certified fresh for every new experiment before a buyer sees results. A single synthetic run can still diverge from real behavior even when the underlying method replicates well on average.
When aggregate replication accuracy is not enough
Aggregate accuracy answers "does this method generally match real choices." It does not answer "does this specific result, for this specific audience and decision, match real choices." For a high-stakes call, the second question sometimes needs its own check.
Subconscious offers real-human validation as a separate service: a team can move a synthetic result into recruited real-human validation without changing the causal question being tested, so a high-stakes decision gets its own check instead of borrowing an aggregate number it was never measured against.
Real-human validation is not a usability session, a clinical trial, or automatic proof that the finding will hold in market. It confirms the choice result against recruited participants answering the same causal question; it does not certify the launch, pricing, or messaging decision that follows.
What to check before you commit budget
- Does the result rest on a method with published, independent replication evidence, or only on internal claims.
- Is the decision high-stakes enough that the aggregate replication number is not sufficient, and a dedicated real-human validation run is worth the added step.
- Does the validation step test the same causal question as the original experiment, not a different one.
Where does this fit before a launch decision?
Run the original synthetic experiment, check the method's aggregate replication track record, and add real-human validation for decisions where being wrong is expensive. See the research methodology, the case studies of past runs, and the leaderboard of replication results across studies before deciding how much independent checking a given decision needs.