Skip to content
Subconscious

12 Human Studies for Evaluating Simulated Choice Research

A research leader evaluating a simulated choice result needs the original human task and a matched simulation record. The twelve studies below offer different preference questions. Naming a human paper alone does not demonstrate that a later model reproduced its results.

Keep the validation tasks separate

Subconscious’s July 2026 working paper reports parameter-rank agreement across a replication corpus. That does not provide every historical case’s run record, establish observed purchases or verify that a selected set of attributes matches a complete human design. Before using an individual comparison, inspect the human and synthetic estimates, matching, estimator, model configuration, date and uncertainty. Public working paper, not peer reviewed.

Twelve human-study designs

Original human studyTask and comparison boundary
[Hainmueller and Hopkins, immigration](https://immigrationlab.org/publication/the-hidden-american-immigration-consensus-a-conjoint-analysis-of-attitudes-toward-immigrants-2/)Support for admitting hypothetical immigrants; distinguish applicant preferences from immigration volume.
[Adida, Lo and Platas, refugees](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0222504)Admission ratings and choices for Syrian refugee profiles in a historical US sample.
[Kreps and colleagues, vaccines](https://pmc.ncbi.nlm.nih.gov/articles/PMC7576409/)Hypothetical vaccine choice and stated willingness, rather than observed uptake.
[Duch and colleagues, allocation](https://pubmed.ncbi.nlm.nih.gov/34526400/)Prioritization of vaccine recipients, rather than personal vaccine acceptance.
[Skreli and colleagues, tomatoes](https://sjar.revistas.csic.es/index.php/sjar/article/download/9889/3589/)Tomato preferences in urban Albania; scope population and hypothetical WTP.
[Wu and colleagues, cars](https://irlib.pccu.edu.tw/handle/987654321/29221?locale=en-US)Ranking subcompact-car profiles among Thai respondents.
[Bechtel and colleagues, climate policy](https://www.nature.com/articles/s41467-022-33830-8)Carbon-policy support under different international cooperation conditions.
[Lüthi and Prässler, wind markets](https://www.sciencedirect.com/science/article/abs/pii/S030142151100485X)Historical onshore-developer preferences over policy conditions.
[Adam and colleagues, treatment processes](https://pmc.ncbi.nlm.nih.gov/articles/PMC6525263/)Consultation-process preferences, rather than clinical efficacy.
[Rao and colleagues, rural jobs](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0082984)Hypothetical job choices among clinician and student groups in India.
[Ares and colleagues, yogurt](https://www.sciencedirect.com/science/article/abs/pii/S0950329309001487)Tradeoffs among product type, brand, price and claim information.
[Claret and colleagues, fish](https://www.sciencedirect.com/science/article/abs/pii/S0950329312001000)Tradeoffs among origin, production method, storage and price.

This is a source map, not a claim that all twelve have public Subconscious run artifacts or belong to a particular filtered release. Individual case pages report a Subconscious score only where a matching public run record exists.

Four requirements for a replication comparison: human task, simulated run, matched estimates and documented scoring.
A human publication and a matched simulation record are separate evidence requirements.

Select a comparison from the buyer’s question

Match the outcome as well as the domain. Allocation preferences cannot validate uptake, treatment-process preferences cannot validate clinical outcomes, and hypothetical job choices cannot establish retention. Inspect individual effects and disagreements before relying on an aggregate score.

For a current decision, document how the population, alternatives and context differ from the historical study. Use a separately scoped human or observed-outcome check where the claim requires it. A positive rank correlation does not establish effect-size calibration or eliminate training-data overlap.

Read the method evidence and limitations and the case records before choosing a benchmark. A decision review can identify the closest task and the additional evidence needed for the decision.