We run
the experiment.
We run the experiment before budget moves. Listening tells you what happened. Prediction guesses what might happen next. Subconscious tests which action changes the outcome, validated against real human behavior.
Get a Demo↗︎// Differentiation
Causal,
not claimed
Experience platforms listen to what customers already said. Predictive dashboards pattern-match what came before. Consulting decks turn synthetic customers into a recommendation.
Subconscious runs randomized experiments, not scraped opinions. Change the price, message, feature, or segment inside a simulated market and measure which action moves the outcome.
That is the difference: not descriptive, not merely predictive, prescriptive with error bars.
The say-do gap is where capital gets misallocated.
// Evidence stack
The proof stack
Randomized interventions on simulated populations, reconciled against human holdouts and a public causal replication benchmark.
Every claim has a source, a number, and a limitation. The leaderboard below is the artifact, not a slogan.
Randomized interventions. Human validation. Public benchmark. Error bars.
Live rankings
| Model | Score | Coverage | Match | Studies |
|---|---|---|---|---|
| Claude Sonnet (Databricks) | 56.78 | 40.07 | 82.63 | 19 |
| GPT-4 (Azure OpenAI) | 55.62 | 46.02 | 78.21 | 19 |
| Claude Sonnet (Databricks) | 54.63 | 45.79 | 70.95 | 23 |
| Gemini Flash (Google) | 54.63 | 23.84 | 81.59 | 19 |
| GPT-4 (Azure OpenAI) | 53.87 | 35.27 | 73.22 | 2 |
| Claude Haiku (Databricks) | 53.62 | 39.92 | 76.60 | 19 |
| GPT-3 Instruct (Azure OpenAI) | 52.79 | 41.50 | 75.97 | 18 |
| GPT-4 (Azure OpenAI) | 52.67 | 34.92 | 78.74 | 9 |
| Gemini (Google) | 51.09 | 33.46 | 76.83 | 9 |
| Claude Sonnet (Databricks) | 50.82 | 28.54 | 77.55 | 9 |
Domain breakdowns
Each table preserves the original domain-level ranking surface from the legacy leaderboard.
Public Health
8 models ranked
| Model | Score | Coverage | Match | Studies |
|---|---|---|---|---|
| Claude Sonnet (Databricks) | 54.63 | 45.79 | 70.95 | 23 |
| Claude Haiku (Databricks) | 50.30 | 44.93 | 65.83 | 23 |
| GPT-4 (Azure OpenAI) | 50.23 | 43.12 | 67.59 | 23 |
| Gemini Flash (Google) | 49.48 | 25.37 | 68.93 | 22 |
| o1 (Azure OpenAI) | 48.49 | 25.67 | 68.72 | 20 |
| Gemini (Google) | 39.71 | 40.47 | 55.42 | 23 |
| GPT-3 Instruct (Azure OpenAI) | 38.94 | 41.12 | 52.27 | 22 |
| o3-mini (Azure OpenAI) | 37.51 | 21.05 | 55.21 | 22 |
Consumer Research
8 models ranked
| Model | Score | Coverage | Match | Studies |
|---|---|---|---|---|
| Claude Sonnet (Databricks) | 56.78 | 40.07 | 82.63 | 19 |
| GPT-4 (Azure OpenAI) | 55.62 | 46.02 | 78.21 | 19 |
| Gemini Flash (Google) | 54.63 | 23.84 | 81.59 | 19 |
| Claude Haiku (Databricks) | 53.62 | 39.92 | 76.60 | 19 |
| GPT-3 Instruct (Azure OpenAI) | 52.79 | 41.50 | 75.97 | 18 |
| o1 (Azure OpenAI) | 50.65 | 25.33 | 75.20 | 19 |
| Gemini (Google) | 50.01 | 37.12 | 74.81 | 19 |
| o3-mini (Azure OpenAI) | 33.11 | 17.91 | 56.59 | 19 |
Economics
8 models ranked
| Model | Score | Coverage | Match | Studies |
|---|---|---|---|---|
| GPT-4 (Azure OpenAI) | 52.67 | 34.92 | 78.74 | 9 |
| Gemini (Google) | 51.09 | 33.46 | 76.83 | 9 |
| Claude Sonnet (Databricks) | 50.82 | 28.54 | 77.55 | 9 |
| Gemini Flash (Google) | 47.47 | 23.81 | 77.87 | 9 |
| GPT-3 Instruct (Azure OpenAI) | 45.03 | 21.58 | 75.79 | 9 |
| Claude Haiku (Databricks) | 43.92 | 31.17 | 73.57 | 9 |
| o1 (Azure OpenAI) | 38.40 | 18.67 | 65.19 | 9 |
| o3-mini (Azure OpenAI) | 33.19 | 4.77 | 63.35 | 9 |
Agricultural Sciences
4 models ranked
| Model | Score | Coverage | Match | Studies |
|---|---|---|---|---|
| GPT-4 (Azure OpenAI) | 43.18 | 24.30 | 53.48 | 2 |
| Claude Sonnet (Databricks) | 38.52 | 29.08 | 47.40 | 2 |
| Claude Haiku (Databricks) | 32.56 | 28.30 | 43.32 | 2 |
| Gemini (Google) | 29.32 | 28.98 | 44.80 | 2 |
Business Administration
4 models ranked
| Model | Score | Coverage | Match | Studies |
|---|---|---|---|---|
| GPT-4 (Azure OpenAI) | 53.87 | 35.27 | 73.22 | 2 |
| Claude Sonnet (Databricks) | 44.64 | 15.63 | 68.58 | 2 |
| Claude Haiku (Databricks) | 44.00 | 18.51 | 62.30 | 2 |
| Gemini (Google) | 43.12 | 29.16 | 60.10 | 2 |
What makes it different
What makes it causal
Prediction can be right for the wrong reason. Subconscious runs randomized controlled experiments on simulated populations - interventions, not observations - so the estimate is causal: change the price, message, or feature and see what moves.
Read the paperHow the evidence compares
Every accuracy number we publish is first-party: 417 replication runs across 5 domains, mid-range scores and failure modes included. Ask any alternative whether the accuracy number is theirs, whether the benchmark is standing, and whether failures are disclosed.
View benchmarkvs Experience management platforms
They listen: a billion signals a month of what customers already said. Root-cause dashboards pattern-match feedback; they cannot test an action before you take it. We run the experiment first, with error bars a CFO can underwrite.
vs Consulting synthetic customers
Strategy firms now sell synthetic customers inside seven-figure engagements: borrowed accuracy numbers, no standing benchmark, results in a deck. We publish replication runs, failures included, and ship the engine as software your team runs in minutes.
vs Predictive dashboards
Forecasts estimate what may happen if the world stays similar. Subconscious tests the intervention itself, so product and strategy teams can choose the action with the strongest causal lift.
Evidence pack
Method paper
The source of record for the validation protocol, replicated studies, and how synthetic respondent experiments are reconciled to human behavior.
Read the paperPublic leaderboard
A living benchmark surface for model behavior across replicated human studies and domains.
View benchmarkEvidence archive
A curated archive of choice modeling, causal reasoning, and experimental design references that inform the method.
Browse archiveChoice-model lineage
The econometric backbone behind serious decision modeling: trade-offs, utility, and structural preference recovery.
Read McFaddenCausal reasoning standard
Decision systems that cannot reason causally should not be trusted with consequential optimization.
Read benchmark contextIntegrity constraints
Measurement quality, fraud, privacy, and boundary conditions are part of the evidence, not footnotes.
Read survey integrity paper// Causal replication benchmark
What the benchmark measures
A benchmark is useful only when it checks structure, not style.
Rank-order correlation
Whether simulated preference order matches observed human preference order across alternatives.
Parameter sign agreement
Whether effect directions match the original human study instead of flipping the causal story.
Coverage probability
Whether uncertainty bounds are honest enough for decision review.
Parameter distance
How closely simulated parameter values match human estimates, not only their direction.
Human model parity
The same statistical model used in the source study is estimated on synthetic responses before comparison.
Transparent publication
Results include enough methodological detail for research, legal, procurement, and executive stakeholders to challenge the claim.
Frequently asked questions
// Evidence questions
For methodology review, contact partners@subconscious.ai.
How is this different from customer listening?
Listening surfaces feedback after the fact. Subconscious tests interventions before launch, so teams can see which action is likely to change behavior.
Is 93% a universal accuracy guarantee?
No. It is a validation-set result against real human behavior. New populations, rare behaviors, and new contexts require local validation.
Why keep the leaderboard public?
A causal claim should be inspectable. The leaderboard shows the current causal replication benchmark instead of hiding the method behind a sales deck.
Can enterprise teams audit the method?
Yes. We share protocols, assumptions, scoring detail, and boundary conditions with qualified research, legal, procurement, and executive stakeholders.
// Current evidence
Benchmark summary
| Signal | Value | Interpretation |
|---|---|---|
| Replicated study runs | 417 | Published and reconstructed experiments currently represented in the benchmark dataset |
| Domains covered | 5 | Public Health, Consumer Research, Economics, Agricultural Sciences, and Business Administration |
| Replication accuracy | 93% | Agreement against real human behavioral outcomes across the validation set |
| Publication standard | Failures included | Mid-range scores, boundary conditions, and methodological limits are part of the record |
// From evidence to action
The operating system for decisions
The point is not a prettier research page. The point is a different operating loop: define the action, run the causal simulation, inspect uncertainty, decide, and learn from the field result.
That is how product, strategy, and growth teams stop arguing from anecdotes and start moving budget toward the action most likely to change behavior.
Get a Demo to learn how behavioral simulation can transform your decision-making.