Why a Media Mix Model Number Needs a Controlled Test Before It Moves Budget
A media mix model can tell a VP of Marketing Analytics that paid social drove 18% of last quarter's revenue. It cannot tell them, on its own, whether cutting paid social next quarter would actually cost that revenue. That gap between a regression estimate and a causal one is the decision this article is about: should a team shift budget on the model's number alone, or hold the number until a controlled test confirms it?
What does a media mix model actually estimate?
A Bayesian media mix model (MMM) fits a regression across historical channel spend and outcomes such as revenue or conversions. It accounts for carryover effects (adstock, the fading influence of past spend) and diminishing returns (saturation, where extra spend on a channel earns less each additional dollar). The output is a contribution estimate per channel and, from that, a suggested reallocation of the next period's budget.
Publishing the assumption behind a contribution number is what lets a buyer check it before spend moves. The method's causal claim rests on an unconfoundedness assumption: that spend is not correlated with unobserved drivers of the outcome. That assumption is fragile because media spend is chosen by marketers, not randomized, and the model fits the pattern in historical spend and outcome data without running an experiment. When two channels moved together historically, for example a paid search campaign that always launched alongside a TV flight, the model can struggle to separate their individual contributions. Google's own applied research frames MMM as a causal inference method that is valid only under assumptions such as no unobserved confounding of the spend-outcome path, and treats controlled experimentation as the source of informative priors that make those assumptions more defensible (Google for Developers, "About MMM as a causal inference methodology").
What calibration methods close the gap between an MMM estimate and a measured effect?
Two experimental designs are the standard way analytics teams check an MMM estimate against real behavior:
- Geo-experiments. Split matched geographic markets into a treatment group (spend changed) and a holdout (spend unchanged), then compare outcomes. The designed geo-experiment literature describes this as a way to recover a channel's incremental effect directly from a designed intervention, rather than inferring it from historical variation (Vaver & Koehler, "Estimating Ad Effectiveness Using Geo Experiments," 2011).
- Difference-in-differences. Compare the change in outcomes for a group exposed to a spend change against the change in a comparable unexposed group over the same period, isolating the effect from other trends moving both groups together.
The reason a team runs a geo-experiment or difference-in-differences test is that a miss should show up on the record before budget moves. Both designs exist because the regression alone cannot rule out confounding. A team that skips this step and reallocates budget purely on MMM output is trusting a fit, not a measured effect.
| Method | What it estimates | What it needs | Where it falls short |
|---|---|---|---|
| Media mix model (regression) | Historical channel contribution to revenue | Time-series spend and outcome data | Cannot fully separate channels that move together; correlational |
| Geo-experiment holdout | Incremental effect of a spend change | Matched geographic markets, a live test | Requires holding real budget back during the test |
| Difference-in-differences | Incremental effect vs. a comparable unexposed group | A treatment group and a comparable control group | Depends on the control group tracking the treatment group absent the change |
Where this decision sits before spend, not just after it
The same principle, that a number needs a controlled comparison before it earns budget, applies earlier than the media plan. Before a launch, a price change, or new messaging goes live, a team can test the action itself: put alternatives in front of respondents in a controlled discrete choice experiment and estimate the causal effect of the tested attributes on stated choice within the experiment, with confidence intervals attached over the simulated population, though stated preference tends to run high relative to real-world behavior (hypothetical bias). Subconscious runs this kind of controlled experiment for pre-launch product, pricing, and messaging decisions, and a team can move from a simulated version of that test to real-human validation to confirm the same effect holds outside the simulation.
Where this does not apply
Naming what a method does not cover is what lets a buyer pick the right tool for the decision in front of them. Subconscious does not fit historical media-spend time series, produce a media mix model, or replace adstock and saturation regression on past spend data. It is not a geo-experiment platform for allocating an existing budget. For a team recalibrating what already happened across channels, MMM and geo-experiments remain the right tools. Subconscious's fit is upstream of that: testing an action or an alternative before it is committed to, not reallocating spend that has already run.
The practical rule
Treat an MMM contribution number the way a geo-experiment team already treats it: as a hypothesis worth testing, not a budget decision on its own. If the model says a channel drove 18% of revenue, the fastest way to find out whether that number is real is to hold spend back in a matched sample and watch what changes. Book time to see how the same test applies to a decision that hasn't been made yet.