Skip to content

We run
the experiment.

We run the experiment before budget moves. Listening tells you what happened. Prediction guesses what might happen next. Subconscious tests which action changes the outcome, validated against real human behavior.

Get a Demo↗︎
417Replicated study runs
5Research domains
93%Replication accuracy against real human behavior
(source)
OpenPublished methodology
(source)

// Differentiation

Causal,
not claimed

Experience platforms listen to what customers already said. Predictive dashboards pattern-match what came before. Consulting decks turn synthetic customers into a recommendation.

Subconscious runs randomized experiments, not scraped opinions. Change the price, message, feature, or segment inside a simulated market and measure which action moves the outcome.

That is the difference: not descriptive, not merely predictive, prescriptive with error bars.

The say-do gap is where capital gets misallocated.

// Evidence stack

The proof stack

Randomized interventions on simulated populations, reconciled against human holdouts and a public causal replication benchmark.

Every claim has a source, a number, and a limitation. The leaderboard below is the artifact, not a slogan.

Randomized interventions. Human validation. Public benchmark. Error bars.

Live rankings

ModelScoreCoverageMatchStudies
Claude Sonnet (Databricks)56.7840.0782.6319
GPT-4 (Azure OpenAI)55.6246.0278.2119
Claude Sonnet (Databricks)54.6345.7970.9523
Gemini Flash (Google)54.6323.8481.5919
GPT-4 (Azure OpenAI)53.8735.2773.222
Claude Haiku (Databricks)53.6239.9276.6019
GPT-3 Instruct (Azure OpenAI)52.7941.5075.9718
GPT-4 (Azure OpenAI)52.6734.9278.749
Gemini (Google)51.0933.4676.839
Claude Sonnet (Databricks)50.8228.5477.559
Similarity
Parameter proximity
Coverage
Confidence overlap
Match
Effect direction
Rank Corr.
Preference ordering
1 / 4

Domain breakdowns

Each table preserves the original domain-level ranking surface from the legacy leaderboard.

Public Health

8 models ranked

#public-health
ModelScoreCoverageMatchStudies
Claude Sonnet (Databricks)54.6345.7970.9523
Claude Haiku (Databricks)50.3044.9365.8323
GPT-4 (Azure OpenAI)50.2343.1267.5923
Gemini Flash (Google)49.4825.3768.9322
o1 (Azure OpenAI)48.4925.6768.7220
Gemini (Google)39.7140.4755.4223
GPT-3 Instruct (Azure OpenAI)38.9441.1252.2722
o3-mini (Azure OpenAI)37.5121.0555.2122
Similarity
Parameter proximity
Coverage
Confidence overlap

Consumer Research

8 models ranked

#consumer-research
ModelScoreCoverageMatchStudies
Claude Sonnet (Databricks)56.7840.0782.6319
GPT-4 (Azure OpenAI)55.6246.0278.2119
Gemini Flash (Google)54.6323.8481.5919
Claude Haiku (Databricks)53.6239.9276.6019
GPT-3 Instruct (Azure OpenAI)52.7941.5075.9718
o1 (Azure OpenAI)50.6525.3375.2019
Gemini (Google)50.0137.1274.8119
o3-mini (Azure OpenAI)33.1117.9156.5919
Similarity
Parameter proximity
Coverage
Confidence overlap

Economics

8 models ranked

#economics
ModelScoreCoverageMatchStudies
GPT-4 (Azure OpenAI)52.6734.9278.749
Gemini (Google)51.0933.4676.839
Claude Sonnet (Databricks)50.8228.5477.559
Gemini Flash (Google)47.4723.8177.879
GPT-3 Instruct (Azure OpenAI)45.0321.5875.799
Claude Haiku (Databricks)43.9231.1773.579
o1 (Azure OpenAI)38.4018.6765.199
o3-mini (Azure OpenAI)33.194.7763.359
Similarity
Parameter proximity
Coverage
Confidence overlap

Agricultural Sciences

4 models ranked

#agricultural-sciences
ModelScoreCoverageMatchStudies
GPT-4 (Azure OpenAI)43.1824.3053.482
Claude Sonnet (Databricks)38.5229.0847.402
Claude Haiku (Databricks)32.5628.3043.322
Gemini (Google)29.3228.9844.802
Similarity
Parameter proximity
Coverage
Confidence overlap

Business Administration

4 models ranked

#business-administration
ModelScoreCoverageMatchStudies
GPT-4 (Azure OpenAI)53.8735.2773.222
Claude Sonnet (Databricks)44.6415.6368.582
Claude Haiku (Databricks)44.0018.5162.302
Gemini (Google)43.1229.1660.102
Similarity
Parameter proximity
Coverage
Confidence overlap

What makes it different

01

What makes it causal

Prediction can be right for the wrong reason. Subconscious runs randomized controlled experiments on simulated populations - interventions, not observations - so the estimate is causal: change the price, message, or feature and see what moves.

Read the paper
02

How the evidence compares

Every accuracy number we publish is first-party: 417 replication runs across 5 domains, mid-range scores and failure modes included. Ask any alternative whether the accuracy number is theirs, whether the benchmark is standing, and whether failures are disclosed.

View benchmark
03

vs Experience management platforms

They listen: a billion signals a month of what customers already said. Root-cause dashboards pattern-match feedback; they cannot test an action before you take it. We run the experiment first, with error bars a CFO can underwrite.

04

vs Consulting synthetic customers

Strategy firms now sell synthetic customers inside seven-figure engagements: borrowed accuracy numbers, no standing benchmark, results in a deck. We publish replication runs, failures included, and ship the engine as software your team runs in minutes.

05

vs Predictive dashboards

Forecasts estimate what may happen if the world stays similar. Subconscious tests the intervention itself, so product and strategy teams can choose the action with the strongest causal lift.

Evidence pack

01

Method paper

The source of record for the validation protocol, replicated studies, and how synthetic respondent experiments are reconciled to human behavior.

Read the paper
02

Public leaderboard

A living benchmark surface for model behavior across replicated human studies and domains.

View benchmark
03

Evidence archive

A curated archive of choice modeling, causal reasoning, and experimental design references that inform the method.

Browse archive
04

Choice-model lineage

The econometric backbone behind serious decision modeling: trade-offs, utility, and structural preference recovery.

Read McFadden
05

Causal reasoning standard

Decision systems that cannot reason causally should not be trusted with consequential optimization.

Read benchmark context
06

Integrity constraints

Measurement quality, fraud, privacy, and boundary conditions are part of the evidence, not footnotes.

Read survey integrity paper

// Causal replication benchmark

What the benchmark measures

A benchmark is useful only when it checks structure, not style.

01 ::

Rank-order correlation

Whether simulated preference order matches observed human preference order across alternatives.

02 ::

Parameter sign agreement

Whether effect directions match the original human study instead of flipping the causal story.

03 ::

Coverage probability

Whether uncertainty bounds are honest enough for decision review.

04 ::

Parameter distance

How closely simulated parameter values match human estimates, not only their direction.

05 ::

Human model parity

The same statistical model used in the source study is estimated on synthetic responses before comparison.

06 ::

Transparent publication

Results include enough methodological detail for research, legal, procurement, and executive stakeholders to challenge the claim.

Frequently asked questions

// Evidence questions

For methodology review, contact partners@subconscious.ai.

01 ::

How is this different from customer listening?

Listening surfaces feedback after the fact. Subconscious tests interventions before launch, so teams can see which action is likely to change behavior.

02 ::

Is 93% a universal accuracy guarantee?

No. It is a validation-set result against real human behavior. New populations, rare behaviors, and new contexts require local validation.

03 ::

Why keep the leaderboard public?

A causal claim should be inspectable. The leaderboard shows the current causal replication benchmark instead of hiding the method behind a sales deck.

04 ::

Can enterprise teams audit the method?

Yes. We share protocols, assumptions, scoring detail, and boundary conditions with qualified research, legal, procurement, and executive stakeholders.

// Current evidence

Benchmark summary

SignalValueInterpretation
Replicated study runs417Published and reconstructed experiments currently represented in the benchmark dataset
Domains covered5Public Health, Consumer Research, Economics, Agricultural Sciences, and Business Administration
Replication accuracy93%Agreement against real human behavioral outcomes across the validation set
Publication standardFailures includedMid-range scores, boundary conditions, and methodological limits are part of the record

// From evidence to action

The operating system for decisions

The point is not a prettier research page. The point is a different operating loop: define the action, run the causal simulation, inspect uncertainty, decide, and learn from the field result.

That is how product, strategy, and growth teams stop arguing from anecdotes and start moving budget toward the action most likely to change behavior.

Get a Demo to learn how behavioral simulation can transform your decision-making.