Skip to content

When Your MMM Gets the Channel Ranking Backwards

A marketing mix model can rank two channels by ROAS and get the order exactly wrong, not close, inverted. Before a growth or marketing leader shifts budget on the strength of that ranking: has anything calibrated it against a real intervention, or is it running on spend and sales data alone?

Five-step chain: a hidden demand signal drives one channel's spend and sales, inflating its estimated effect and reversing the ranking, until a lift test corrects it.
A ranking never checked against a controlled intervention is a hypothesis, not a fact about which channel to cut.

The failure mode: a hidden variable inflates one channel

An MMM estimates each channel's contribution from historical spend and outcome data. That works cleanly when nothing else is pulling on both the channel and the result at the same time; it breaks when something is.

In a published simulation, one channel's spend was also driven by a demand signal the model never saw, something correlated with when people were already inclined to buy. Because that hidden driver pushed up both the channel's spend and the outcome, the model credited the channel with an effect that wasn't fully its own. The result: the model ranked that channel as more effective than a second channel, when the underlying process that generated the data showed the opposite ordering (Media Mix Model Calibration With Bayesian Priors, Google Research).

This is not a data-quality problem you fix by collecting more of the same kind of data. An unobserved confounder (a variable influencing both the channel and the outcome but absent from the model) stays invisible no matter how many additional weeks of spend and sales history get added.

Why the fix is an intervention, not a better regression

The correction in the published example didn't come from re-specifying the model with cleverer priors. It came from small controlled interventions: deliberately changing one channel's spend for a period and measuring the resulting change in sales directly, independent of the historical spend pattern. Two such tests per channel, applied as constraints on the model rather than as another input to regress on, moved both channels' estimates back toward the values implied by the true data-generating process and reversed the ranking to the correct order (worked example replicated from an experiment-calibrated simulation study of media mix models, Juan Orduz).

The general principle holds independent of any specific software implementation: an observational model answers "what happened," and only a deliberate intervention answers "what would happen if we changed this."

What this means before the next reallocation

The decision a leader actually faces is not "trust the model or don't." It's narrower: does this specific ranking, the one about to move budget, rest on anything besides spend and sales history? If the answer is no, the honest move is to treat the ranking as provisional and test the channel the model says to cut before cutting it.

This is the same logic behind running a controlled experiment rather than reading correlation off a dashboard: the question "what would sales do if we changed this channel's spend" only gets a trustworthy answer from something that changes the input and observes the result. Subconscious runs that kind of controlled test directly against the causal question at stake, and the same underlying test can move from a simulated read to real-human validation without changing the question being asked.

A decision path from a model's ranking flagging a channel to cut. It branches on whether the ranking was checked against a real intervention: no leads to testing the channel first; yes leads to acting on the ranking.
A ranking is safe to act on only after it has been checked against a real spend change, not just fit to historical data.

Where this example stops being evidence

The worked example above is a third-party technical illustration of a calibration principle, not a Subconscious case result. No accuracy figure or customer outcome from Subconscious applies to it. What it is useful for is the pattern: a plausible-looking ranking, an invisible variable, and a specific fix that required an actual intervention. See the current leaderboard for how that kind of testing gets evaluated in practice.