Skip to content

What a Weak Replication Correlation Actually Tells a Product Team

The decision this replication is meant to inform

A CPG or food-brand product leader is weighing which attributes to test before a launch or reformulation: organic certification, growing region, production method, price. The practical question is whether to run a causal simulation first or go straight to a fielded human study.

Committing budget to a fielded study, or to a product change, on the strength of an untested simulation method is the expensive mistake this check is meant to prevent.

The published study used as the check

Skreli et al.'s 2014 study ran a discrete choice experiment on Albanian consumers, testing preference and willingness to pay for tomatoes across four dimensions: whether the tomato was organically grown, its growing region, whether it was hothouse grown, and its price (Spanish Journal of Agricultural Research). Subconscious ran the same attribute set through a simulated discrete choice experiment and compared the results.

Attribute testedSource of evidence
Organically grownSkreli et al. conjoint study
Growing regionSkreli et al. conjoint study
Hothouse grownSkreli et al. conjoint study
PriceSkreli et al. conjoint study
Agreement between simulation and studySpearman rs = .5213, p = .1008

What the correlation says, plainly

The reported agreement between the simulated result and the published human study was rs = .5213, p = .1008, a moderate correlation that did not clear the conventional p < .05 threshold. That is neither a strong match nor a failure to replicate: on this attribute set, the simulation tracked some of the same preference structure as the human study, without a demonstrated significant relationship.

Why reporting a weak result is the more useful signal

A vendor that only shows its best replications is not showing a product team enough to make a real decision. Publishing a case where the correlation is moderate and non-significant is the evidence that matters: it shows where a simulation-first approach tracks published human preference data and where it does not, on a specific and checkable attribute set, rather than asserting a general accuracy claim and asking a buyer to trust it.

Limitations and what this result does not establish

This single replication does not establish general accuracy for the method, or show that it reliably reproduces human choice on other attribute sets, categories, or markets. A result like this one is useful for narrowing which attributes and levels are worth testing next. It is not a substitute for a fielded study when the decision is high-stakes, such as committing to a reformulation or a launch price.

What to do with a result like this

Use a simulation pass to decide which attributes are worth testing at all, not as the final word on which one wins. When the decision is expensive to get wrong, pair the simulation with a fielded human study on the same attribute set before locking the choice. Review the research methodology behind these comparisons, browse other case studies to see where correlations are strong and where they are not, or talk to the team about designing a study that pairs a simulation pass with real-human validation.

Comparison of a simulated discrete choice experiment against a published Albanian consumer study across organic, growing region, hothouse, and price. Overall agreement is rs=.52, p=.10, below conventional significance.
A moderate, non-significant correlation across four attributes shows where the simulation tracked human preference and where it didn't, reported honestly rather than smoothed over.