Skip to content

Bechtel Carbon Tax Policy Study: Simulated vs. Published Results

Policy and public-affairs teams weighing a costly, slow fielded study on a carbon tax package need to know whether a simulated discrete-choice experiment gets close enough to trust for an early pass.

The comparison

Bechtel, Scheve, and van Lieshout fielded a discrete choice experiment across four countries, the U.S., U.K., Germany, and France, asking what carbon tax package design features people there favor, among them revenue use, exemptions, and international coordination (Improving public support for climate action through multilateralism, Nature Communications, 2022).

Subconscious ran the same discrete-choice design on a simulated U.S. population and compared the resulting preference ranking against the paper's pooled U.S. results. The two rankings correlate at rs = .6711 (p < .001) (Improving public support for climate action through multilateralism, Nature Communications, 2022).

Two ranked lists of carbon tax package features, fielded U.S. survey and simulated study, linked by a moderate rs = .67 correlation showing directional agreement but feature-level divergence.
A moderate rs = .67 correlation means the rankings agree on direction but can still diverge on individual carbon tax features.

What rs = .67 means for this decision

A Spearman correlation of .67 is a moderate, positive relationship: the simulated and fielded studies tend to agree on which carbon tax features respondents prefer more or less, but not closely enough to treat the simulated ranking as a stand-in for the fielded one on every feature. Treat it as one data point on whether the method is directionally reliable for this kind of policy-preference question, not proof that a simulated ranking will match a real vote or survey.

Limitations

This comparison covers one replication, on one country's pooled results, against one published study. It does not extend to the U.K., Germany, or France samples in the original paper, and one correlation coefficient does not establish a fixed accuracy rate for policy-preference studies generally. A moderate correlation also means feature-level orderings can still diverge even where the overall ranking pattern agrees.

Two-column comparison. Left, tested: U.S. pooled results. Right, not tested: U.K. sample, Germany sample, France sample, and a general accuracy rate for policy-preference studies.
This replication checked one country against one study, so it doesn't show how the method performs on the U.K., Germany, or France samples in the same paper.

Next step

For a policy or public-affairs question where getting the tradeoff wrong is expensive, use a simulated run like this one to sharpen the design and flag which features carry the most disagreement risk before committing to a fielded study. See how the underlying method works on /how-we-work, or review more replication comparisons on /case-studies.