Bechtel Carbon Tax Policy Study: Simulated vs. Published Results
Policy and public-affairs teams weighing a costly, slow fielded study on a carbon tax package need to know whether a simulated discrete-choice experiment gets close enough to trust for an early pass.
The comparison
Bechtel, Scheve, and van Lieshout fielded a discrete choice experiment across four countries, the U.S., U.K., Germany, and France, asking what carbon tax package design features people there favor, among them revenue use, exemptions, and international coordination (Improving public support for climate action through multilateralism, Nature Communications, 2022).
Subconscious ran the same discrete-choice design on a simulated U.S. population and compared the resulting preference ranking against the paper's pooled U.S. results. The two rankings correlate at rs = .6711 (p < .001) (Improving public support for climate action through multilateralism, Nature Communications, 2022).
What rs = .67 means for this decision
A Spearman correlation of .67 is a moderate, positive relationship: the simulated and fielded studies tend to agree on which carbon tax features respondents prefer more or less, but not closely enough to treat the simulated ranking as a stand-in for the fielded one on every feature. Treat it as one data point on whether the method is directionally reliable for this kind of policy-preference question, not proof that a simulated ranking will match a real vote or survey.
Limitations
This comparison covers one replication, on one country's pooled results, against one published study. It does not extend to the U.K., Germany, or France samples in the original paper, and one correlation coefficient does not establish a fixed accuracy rate for policy-preference studies generally. A moderate correlation also means feature-level orderings can still diverge even where the overall ranking pattern agrees.
Next step
For a policy or public-affairs question where getting the tradeoff wrong is expensive, use a simulated run like this one to sharpen the design and flag which features carry the most disagreement risk before committing to a fielded study. See how the underlying method works on /how-we-work, or review more replication comparisons on /case-studies.