Skip to content

How to compare simulated and human experimental results

Compare simulated and human experiments only when they measure the same alternatives, population, and outcome. Then test whether both experiments rank the alternatives in a similar order. A high rank correlation supports agreement for that study. It does not prove universal accuracy.

This is the validation discipline behind Subconscious. A causal experiment earns weight by reproducing a human baseline under matched conditions, not by producing a plausible response.

A four-step scale of rank correlation values from below .52 to .96, each paired with a scatter pattern from scattered points to a near-perfect diagonal line.
Rank correlation is read against the 0.52 human baseline in bands, not as a single threshold.

Define the comparison before reading the result

The unit of comparison is a matched experiment. Subconscious can use the same intervention, choice set, target population, and measured outcome for simulated and recruited-human runs.

This distinction matters because a model can match an overall average while missing the ordering of alternatives or the response of an important segment. Validation should preserve the decision a buyer will make.

Establish a human baseline

The historical method grouped responses from a published experiment with 25 levels into three subsets. Pairwise comparisons between the subsets produced Spearman rank correlations from 0.52 (p=.006) to 0.86 (p<.001). The analysis used at least 0.52 as its within-human agreement baseline.

Three pairwise comparisons between subsets of a human experiment.
Human subset comparison A and B
A second pairwise comparison between human response subsets.
Human subset comparison B and C
A third pairwise comparison between human response subsets.
Human subset comparison A and C

The source described this threshold as explaining ~25% of the variation in human decisions at a rank correlation of at least 0.52. Keep that interpretation tied to the historical experiment. It is not a general product guarantee.

Read a rank-correlation chart

Spearman rank correlation measures whether two result sets order alternatives similarly. The coefficient does not require the values to be identical.

Rank correlation .32 with p .234. The points are scattered and the chart is outlined in red.
No clear agreement

A coefficient below .52 did not meet the historical baseline.

Rank correlation .54 with p .040. The points form a wide diagonal band and the chart is outlined in green.
Agreement near the historical human baseline

The .54 result clears that study's threshold, although the wide spread still matters.

Rank correlation .76 with p .001. Points follow the diagonal with visible spread.
Stronger rank agreement

The .76 result shows stronger agreement in ordering.

Rank correlation .96 with p less than .001. The points form a clear diagonal line.
Very strong rank agreement

The .96 result shows very strong agreement for the compared experiment.

What the comparison can support

A matched result can support a narrow claim: the simulated study recovered the ordering found in the human study under the tested conditions. It cannot show that every audience, intervention, or outcome will behave the same way.

Subconscious uses this distinction to separate a tested decision from synthetic roleplay. Agreement and disagreement both become useful evidence when the intervention, population, outcome, and comparison rule remain fixed. The company publishes a replication leaderboard so a buyer can inspect the human baseline, result, and failure condition rather than accept an unqualified accuracy claim.

The Subconscious research program explains the larger validation approach. Current replication claims belong with the canonical paper, where the metric and evidence can be defined. They should not be derived from the historical 0.52 threshold on this page.

Use the counterfactual causal inference guide to see how model assumptions affect an estimated effect.

Limitations

Rank correlation does not measure calibration, individual-level fidelity, subgroup validity, or the cause of disagreement. A complete validation program should examine those questions separately and report failed replications as well as successful ones.

The charts and thresholds above describe a historical validation method. They require evidence review before publication as current Subconscious proof. The decision rule is narrow: report what replicated, what did not, and which conclusion the evidence can support.