Incrementality Testing: What It Proves, What It Cannot, and How to Run It Causally
A marketing leader deciding next quarter's budget needs to know one thing about incrementality testing before commissioning one: it answers the wrong question at the wrong time for that decision. Incrementality testing proves that a specific campaign, already launched with real budget, caused a measurable lift over a holdout or control group. It cannot tell a marketing leader, before that budget goes out the door, which creative, price point, or channel would have produced a larger lift, because the test has no data on the options that never ran. That gap is where budgets get spent twice: once on the campaign, once on the test that only confirms what already happened.
- Incrementality testing proves causal lift for the campaign you already ran, using a holdout, test/control split, or geo-experiment.
- It has no counterfactual for creative, price, or channel options that never received spend, so it can't compare what you did to what you didn't do.
- Platform-run lift studies (Meta, Google) are the most common version in use, and the seller grading its own media is a conflict of interest the industry now names openly.
- Independent geo-experiments, built on methods like Meta's open-sourced GeoLift, are more rigorous but require real budget, weeks of runtime, and enough markets to reach significance.
- The decision that actually moves next quarter's budget, which option to fund next, has to be answered with a randomized experiment before the campaign launches, not after.
What Does Incrementality Testing Actually Prove?
It proves that one specific campaign, already run, caused an outcome beyond what would have happened anyway, not that a different campaign would have caused more of it. A geo-experiment withholds treatment from a set of control markets and compares the difference against markets that ran the campaign; a platform lift study does the same with user-level holdouts inside a single ad account. Across 225 geo-based tests run between August 2024 and December 2025, median incremental ROAS came in at 2.31x, with 88.4 percent of tests reaching statistical significance, according to a 2026 geo-testing guide's dataset, one vendor's client base rather than an industry census. The same dataset found that incremental ROAS differed from platform-reported ROAS by 30 to 70 percent, which is the clearest evidence that the platform's own number and the causal number are not the same thing. Each measurement method also has a documented blind spot: multi-touch attribution is blind to offline conversion and prone to over-crediting, marketing mix modeling overfits without an experimental check, and incrementality tests can't economically cover every channel at once, per Measured's comparison of the three.
Why a Platform Grading Its Own Media Isn't Independent Proof
Because the company selling the media is also the company scoring whether the media worked. Meta and Google run their conversion lift studies as test/control audience splits entirely inside their own walled gardens, with no external party checking the math. That structural conflict is a large part of why 52.8 percent of US advertisers report plans to add incrementality testing this year, per an AI Digital industry review. Independent geo-experiments avoid the self-grading problem by using synthetic control methods to build a counterfactual from markets that never saw the campaign, but they cost real money and time: one 2026 testing guide puts platform lift tests at roughly $30,000 to $50,000 in minimum spend over a four-week window before noise overwhelms signal, and geo tests at twenty or more markets with four to six weeks of exposure. That's the price of an honest answer to a question you've already committed to.
Which Incrementality Method Should You Trust?
It depends on what you're using it to decide. The table below compares the three approaches a senior buyer is likely choosing between right now.
| Method | Who grades it | Timing vs. spend | Typical cost and duration | What it proves |
|---|---|---|---|---|
| Platform conversion lift study | The platform itself (Meta, Google) | After the campaign is already running | Bundled into ad spend; results in days to weeks | Lift attributable to that platform's exposure, with no external verification. **Best for:** a fast, in-platform sanity check for teams that accept the self-grading limitation. |
| Independent geo-experiment (GeoLift-style synthetic control) | An independent analyst or agency | After launch, before a scale-up decision | Twenty or more markets, 4 to 6 weeks of exposure; platform lift variants need roughly $30,000 to $50,000 minimum spend | Causal lift of the exact campaign that ran; median 2.31x iROAS across one guide's 225-test client dataset, not an industry census. **Best for:** proving an already-running campaign's true lift to a skeptical finance stakeholder. |
| Pre-spend randomized experiment on a simulated population | The experiment design itself, checked against human-study replication | Before any budget commits | No media spend required; results in days | Which creative, price, or channel option a population would choose, with a confidence interval over the simulated population. **Best for:** the actual decision, which option to fund next, made before a dollar leaves the building. |
What Can't Incrementality Testing Tell You?
It can't tell you what would have happened with a creative, price, or channel you never launched, because every geo-experiment or lift study is built around the specific intervention it measures. GeoLift, Meta's open-source geo-testing package, constructs its counterfactual with synthetic control: it builds a weighted combination of untreated "donor" markets to estimate what the treated markets would have done without the campaign, per its published methodology. That counterfactual is for the campaign that ran, not for the campaign you didn't run. If the real question is "would a different price or a different hero creative have moved more volume," a geo-test has no answer, because it never observed those alternatives in the market at all.
Where Does the Budget Decision Actually Get Made?
It gets made with a randomized experiment run before any money is spent, on a simulated population, analyzed with discrete choice models that estimate which option people would actually choose. Discrete choice estimators, McFadden's foundational discrete choice model, Mixed Logit, and ICLV, are not themselves causal methods; the causal identification comes from randomizing which product, price, or message a respondent sees within the experiment design. That's why the framing is randomized experiments analyzed with discrete choice models, not "causal methods like DCE." A plain multinomial logit also carries the independence of irrelevant alternatives assumption, meaning it treats a substitution between any two options as unaffected by a third; Mixed Logit and ICLV relax that assumption when preference share and substitution questions matter. In validation testing, a simulated study reproduced the direction and outcome of the original human study 93 percent of the time, a validation-set result, not a guarantee for a new market, and not a claim that runs around the fact that published studies can sit in a model's training data (go.subconscious.ai/paper). Any confidence interval from that kind of experiment covers the effect within the simulated population, not the real market unconditionally. The current results and how they're scored are public on the leaderboard.
Does Incrementality Testing Replace MMM and MTA?
No, and 2026 practitioner consensus has settled on treating attribution, marketing mix modeling, and incrementality testing as complementary layers rather than competing ones. In practice, few teams run all three together: only 39 percent of organizations use attribution, incrementality testing, and MMM in combination, despite the methods being designed to cover each other's blind spots, per IAB/BWG Global's State of Data 2026 findings cited by House of Martech. That unified stack is still a post-spend stack: MMM runs quarterly on historical spend, MTA runs daily on historical touches, and incrementality testing runs on a campaign that already launched. None of the three answers which option to fund before the money moves. That's a separate layer that has to run first. More on how these methods fit together is in the methods and validation hub.
How Should You Sequence These Tests?
Run the pre-spend causal experiment first, to pick the creative, price, or channel most likely to win, then commit budget to that option, then use a platform lift study or an independent geo-experiment to prove the committed campaign actually worked, then feed that result into the MMM baseline for ongoing allocation. Reversing that order, spending first and asking which option was best afterward, is how a 30 to 70 percent gap between platform-reported ROAS and true incremental ROAS becomes the first time anyone finds out the campaign wasn't the strongest option available.
Before your next campaign locks a creative, price, or channel, run the pre-spend comparison as a randomized experiment rather than a guess, then hold the winner to a geo-test to prove it out. If you want a second opinion on where that experiment should sit in your measurement stack, talk to us.