Skip to content
0526.05Methodology validation

Method validation

Methodology validation

Mean Spearman rank correlation across 43 blind replications of published human trials: 0.73.

Stated

How closely do simulated choices reproduce published human trial results when the model has not seen the outcomes? The comparison needs a named benchmark, a scoring method, and published results.

Revealed

Across 43 blind replications of published human trials, the mean Spearman rank correlation is 0.73. Separately, the best configuration on the Hainmueller immigration conjoint scores 0.832 against a measured human-to-human ceiling of 0.959, or 87% of that ceiling on that study. The single-study ceiling is not a denominator for the 43-study average. Replication results, including misses, are published on the open leaderboard.

Numbers

Figures with provenance.

0.832 / 0.959

best configuration vs. the measured human-to-human ceiling on the Hainmueller immigration conjoint (87% of ceiling)

Inspect provenance
Date
2026-08-14
Method
Single-study comparison; 1000 seeded split-half runs, Spearman-Brown corrected, inclusive convention
Source
Causal fidelity paper (fidelity.subconscious.ai/papers/causal-fidelity/causal_fidelity_paper.pdf); ceiling method ditto_metastudy/human_ceiling.py, 1000 seeded split-half runs, Spearman-Brown corrected (Ditto#268)

0.73

mean Spearman correlation across 43 well-designed replications

Inspect provenance
Date
2026-05-18
Method
Blind replication of published RCTs
N
43
Source
Yashchin et al. 2024; 04-methodology.pdf (Drive, 2026-05-18)

300

published academic papers replicated

Inspect provenance
Date
2026-05-18
N
300
Source
Yashchin et al. 2024; 04-methodology.pdf (Drive, 2026-05-18)

Model-predicted

Stated before the answer was known.

Predicted

Outcomes for published randomized controlled trials, generated blind — without sight of what each trial actually found.

Observed

The real human outcomes those trials reported.

Agreement

0.73 mean Spearman rank correlation across 43 blind replications passing the design filters.

Yashchin et al. 2024

Causality

The method reproduces causal effects, not surface correlations. A model that only matched correlations would not track the outcomes of randomized trials it was never shown.

Intervention
The randomized interventions the original trials ran, reproduced as designed.
Outcome
The behavioral outcomes those trials measured.
Design
Blind replication of published RCTs on McFadden discrete choice methodology, built on 1,200 hand-transcribed studies and a training set of 3.5M real humans.
Most AI generates language. Ours helps people make decisions.
Avi Yashchin, CEO, Subconscious.ai

Method

What can be checked.

Respondents
43
Tool
McFadden discrete choice; blind replication of published RCTs
Dates
Not published
Primary source
Yashchin et al. (2024)

Open primary source

Back to case studies