Method validation
Methodology validation
Mean Spearman rank correlation across 43 blind replications of published human trials: 0.73.
Stated
How closely do simulated choices reproduce published human trial results when the model has not seen the outcomes? The comparison needs a named benchmark, a scoring method, and published results.
Revealed
Across 43 blind replications of published human trials, the mean Spearman rank correlation is 0.73. Separately, the best configuration on the Hainmueller immigration conjoint scores 0.832 against a measured human-to-human ceiling of 0.959, or 87% of that ceiling on that study. The single-study ceiling is not a denominator for the 43-study average. Replication results, including misses, are published on the open leaderboard.
Numbers
Figures with provenance.
0.832 / 0.959
best configuration vs. the measured human-to-human ceiling on the Hainmueller immigration conjoint (87% of ceiling)
Inspect provenance
0.73
mean Spearman correlation across 43 well-designed replications
Inspect provenance
300
published academic papers replicated
Inspect provenance
Model-predicted
Stated before the answer was known.
Predicted
Outcomes for published randomized controlled trials, generated blind — without sight of what each trial actually found.
Observed
The real human outcomes those trials reported.
Agreement
0.73 mean Spearman rank correlation across 43 blind replications passing the design filters.
Yashchin et al. 2024
Causality
The method reproduces causal effects, not surface correlations. A model that only matched correlations would not track the outcomes of randomized trials it was never shown.
- Intervention
- The randomized interventions the original trials ran, reproduced as designed.
- Outcome
- The behavioral outcomes those trials measured.
- Design
- Blind replication of published RCTs on McFadden discrete choice methodology, built on 1,200 hand-transcribed studies and a training set of 3.5M real humans.
Most AI generates language. Ours helps people make decisions.