Best Incrementality Testing Tools in 2026, Compared by the Decision They Serve
The best incrementality testing tool for 2026 depends on which question you're asking. Every tool in the category (Meta Conversion Lift, Google's Meridian GeoX, Measured, Haus, Sellforte, Recast/GeoLift) measures spend that is already live in market. A randomized discrete-choice experiment measures a price, message, or feature before spend commits, which none of the lift-test tools can do. A VP of growth comparing vendor renewal quotes against Google's free geo tool is really choosing between those two questions, not between vendors.
- Platform-native lift tests (Meta Conversion Lift, Meridian GeoX) are free and strictly retrospective.
- Vendor platforms (Sellforte, Measured, Haus, Recast) differ mainly in test-type coverage and channel reconciliation, not in whether they can test something before launch.
- DIY geo-holdouts are cheapest and most underpowered: their sample size is markets, not people.
- No tool in the category can evaluate a price, message, or product that has not shipped yet; that requires a randomized experiment run before spend commits, not a lift test run after.
- A framing now common among measurement vendors treats incrementality results as priors that calibrate a marketing mix model (MMM), not as a standalone source of truth.
What does incrementality testing actually measure in 2026?
It measures what would have happened to sales or conversions if a specific, already-live ad spend had not occurred. Measured's own comparison draws the line cleanly: MMM shows what spend correlates with sales across history; incrementality testing shows what would have happened absent the spend you already committed. Both questions look backward. Neither tells a buyer what will happen to a campaign, price, or product that hasn't launched, because the method needs a treatment and control group already in market, and you can't randomize people into a product that doesn't exist.
The 2026 field: platform-native tools, vendor platforms, and DIY geo-holdouts
Three categories cover the market this year. Platform-native tools live inside one walled garden or one MMM: Meta Conversion Lift measures Meta spend, and Google's Meridian GeoX (announced at Google Marketing Live 2026) runs holdback, go-dark, and heavy-up geo designs that feed the open-source Meridian MMM. Vendor platforms (Sellforte, Measured, Haus, Recast) sell cross-channel geo and holdout testing as a managed service, differentiated by test-type coverage and integrations. DIY frameworks, built on open-source packages like GeoLift, give a team full control over a synthetic difference-in-differences design at the cost of building and powering it themselves.
Comparing the tools by the decision they serve
The honest comparison is which question each tool answers, before or after the money moves.
| Tool | Category | What it measures | Decision it serves | Best for |
|---|---|---|---|---|
| Meta Conversion Lift | Platform-native | Incremental conversions from Meta spend already running | Was this platform's live spend incremental | Brands validating spend inside one platform's walled garden |
| Meridian GeoX | Platform-native, open-source | Geo holdback, go-dark, or heavy-up lift, feeding Google's Meridian MMM | Was live spend incremental, and what prior should it set for MMM | Teams already running Meridian MMM who want a free, publisher-agnostic geo test |
| Sellforte | Vendor platform | Cross-channel geo and holdout testing across multiple test types | Was cross-channel spend incremental, across several test designs | Teams that want broad test-type coverage from one vendor |
| Measured | Vendor platform | Cross-channel incrementality with MMM and MTA framing built in | Was spend incremental, reconciled against MMM and MTA outputs | Teams that already run MMM and MTA and want one vendor to reconcile all three |
| Haus | Vendor platform | Geo and holdout incrementality, managed and lighter-weight | Was spend incremental in a managed, lighter-weight setup | Leaner teams that want a managed test without a full enterprise contract |
| Recast / GeoLift | Vendor + open-source | Synthetic difference-in-differences geo lift, vendor-managed or self-built | Was geo-level spend incremental | Teams with in-house data science who want to own the model |
| DIY synthetic diff-in-diff | Open-source / self-built | Geo lift, powered by number of markets, not households | Was spend incremental, at whatever precision your market count allows | Teams with data science capacity and no budget for a vendor |
| Controlled discrete-choice experiment (Subconscious) | Randomized pre-launch experiment | Which price, message, or feature a market prefers, before any of it ships | Which action to take, before spend commits | Teams deciding what to launch, not auditing what already ran |
Why are geo lift tests so hard to power correctly?
Because the sample size in a geo test is the number of markets you have, not the number of people in them, which makes statistical power scarce by design. A geo experiment analyzed with synthetic difference-in-differences draws its power from how many designated market areas you can split into treatment and holdout, not from the much larger number of individual customers inside those markets (power analysis walkthrough). That same walkthrough found that with a handful of DMAs, brands often need a minimum detectable effect in the 5 to 10 percent range sustained over four to eight weeks just to see a signal above noise. That range moves with market count, baseline variance, and category, so treat it as a planning heuristic from one analysis, not a guarantee for your test.
The calibration loop: how incrementality, MMM, and MTA fit together now
Measurement vendors increasingly frame the three methods as a loop rather than competitors. Incrementality tests set the priors an MMM starts from, the MMM decides which channels are worth testing next, and MTA optimizes spend inside channels the MMM has already validated (Liftlab's breakdown of the calibration loop). The loop improves on treating the three as competing verdicts, but every input to it comes from spend that is already live. It calibrates what happened; it cannot say what to launch next.
What none of these tools can do: deciding before you spend
Every tool above shares one precondition: the spend has to exist in market first. None can be pointed at a price you haven't set, a message you haven't written, or a feature you haven't built, because there's nothing live to hold out against. For a buyer whose actual decision this quarter is which of three prices to launch, or which of two messages to run, the entire incrementality category is the wrong tool, not because it measures poorly, but because it measures too late to inform the decision at hand.
Where a randomized discrete-choice experiment fits instead
This is the gap a randomized, pre-launch experiment closes. Subconscious runs randomized experiments on a simulation of a market, validated against real human behavior, before the action is live. The experiments are analyzed with discrete choice models (McFadden discrete choice, Mixed Logit, ICLV); these are estimators, and the causal read comes from the randomization in the experiment design, the same logic that makes an A/B test causal. A basic logit assumes independence of irrelevant alternatives; Mixed Logit and ICLV relax that assumption when substitution patterns matter.
For pricing, one caveat is non-negotiable: stated willingness to pay runs high relative to real purchases unless the design ties responses to real stakes. A known direction to correct for, like a geo test's power constraint.
The validation figure is 93 percent replication accuracy: how often a simulated study reproduces the direction and outcome of the original human study, per the validation paper. It is a validation-set result, not a guarantee for an untested market, and the protocol accounts for the risk that published studies sit in a model's training data. Per-market performance is on the leaderboard; estimator details are in methods and validation.
Which tool should you actually pick?
It depends on which question you're actually asking. If the question is "did the spend I already committed work," pick from the table above by budget and channel: Meridian GeoX if you're already on Meridian MMM and want a free option, Sellforte if you need broad test-type coverage, Haus if you want a lighter managed contract, GeoLift if you'd rather own the model in-house. If the question is "which price, message, or feature should I launch," none of those tools can answer it, because all of them require the thing you're deciding on to already be live. That decision needs a randomized experiment run before spend commits, not a lift test run after. Some teams need both: a discrete-choice experiment to decide what to launch, then a geo or platform lift test to confirm it worked once it's live. Recent applications of this sequencing are documented in the case studies archive.
Write down which of the two questions your next budget cycle is asking, then check the leaderboard for replication performance in a market like yours. If you want to talk through where a pre-launch experiment fits next to the incrementality stack you already run, the team is at /meet.