Skip to content
0526.05Methodology validation

Case study

Methodology validation

Peak Spearman correlation with real human RCT outcomes in health and policy, across a blind replication of published trials.

Stated

Behavioral simulation tools claim accuracy without public, peer-reviewed validation. The category has no shared benchmark, so buyers cannot tell which method actually predicts behavior. The question a buyer should ask is not whether a vendor claims accuracy, but against what, and published where.

Revealed

Subconscious answers it against published trials it did not see the outcomes of. Blind replication of 43 well-designed studies returned a 0.73 mean Spearman correlation with real human RCT outcomes, peaking at 0.93 in health and policy. The corpus underneath is 1,200 hand-transcribed studies in economics, sociology, and psychology, and a training set of 3.5M real humans rather than model role-play.

Numbers

Figures with provenance.

0.93

peak Spearman correlation with human RCT outcomes, health and policy

Date
2026-05-18
Method
Blind replication of published RCTs
N
43
Source
Yashchin et al. 2024; 04-methodology.pdf (Drive, 2026-05-18) — “method: blind replication of published RCTs · 43 studies”

0.73

mean Spearman correlation across 43 well-designed replications

Date
2026-05-18
Method
Blind replication of published RCTs
N
43
Source
Yashchin et al. 2024; 04-methodology.pdf (Drive, 2026-05-18)

300

published academic papers replicated

Date
2026-05-18
N
300
Source
Yashchin et al. 2024; 04-methodology.pdf (Drive, 2026-05-18)

Model-predicted

Stated before the answer was known.

Predicted

Outcomes for published randomized controlled trials, generated blind — without sight of what each trial actually found.

Observed

The real human outcomes those trials reported.

Agreement

0.73 mean Spearman correlation across 43 well-designed replications, peaking at 0.93 in health and policy, across 300 published papers.

Yashchin et al. 2024

Causality

The method reproduces causal effects, not surface correlations. A model that only matched correlations would not track the outcomes of randomized trials it was never shown.

Intervention
The randomized interventions the original trials ran, reproduced as designed.
Outcome
The behavioral outcomes those trials measured.
Design
Blind replication of published RCTs on McFadden discrete choice methodology, built on 1,200 hand-transcribed studies and a training set of 3.5M real humans.

Before

A category where accuracy is asserted rather than published, with no shared benchmark a buyer can check.

With Subconscious

Blind replication of published randomized trials, on McFadden discrete choice methodology, built on 1,200 hand-transcribed studies and 3.5M real humans.

Result

0.73 mean Spearman across 43 well-designed replications, peaking at 0.93 in health and policy. 300 published papers replicated. The method is peer-reviewed and readable.

Most AI generates language. Ours helps people make decisions.
Avi Yashchin, CEO, Subconscious.ai

Method

What can be checked.

Respondents
43
Tool
McFadden discrete choice; blind replication of published RCTs
Dates
Not published
Primary source
Yashchin et al. (2024)
Back to case studies