Skip to content

Ad Testing: Maximising the impact and ROI of your ads

New title (serves the query, drops the puffery, states the actual decision): "Ad Testing Before Spend: Which Ad Variant to Fund for Maximum ROI"

Here's the revised body (no H1; title lives in frontmatter only):

A marketer holds a shortlist of ad variants and a media budget that hasn't been spent yet. The decision is which variant to fund. Ad testing maximizes impact and ROI only when it answers that decision before spend: a randomized experiment on the creative itself, not a survey score collected after the ad is built, not a split test run once media is live, not an incrementality check that only exists after the campaign already ran.

- A randomized experiment on creative variants, run before the media budget is committed, isolates which variant causes the behavior change; a survey score collected after the ad is finished can't make that claim.
- Pretest tools like Kantar and System1 measure stated liking and attention, which are correlational snapshots, not evidence about what caused a behavior change.
- On-platform split tests confound creative performance with the platform's own delivery algorithm and audience allocation, so a "winning" variant may be winning on targeting, not message.
- Incrementality testing is genuinely causal but runs after spend is live, which means it can validate one campaign but can't screen the variants a buyer is choosing among beforehand.
- The fix the industry is converging on is triangulation, stacking pretest, MMM, and incrementality together, which is itself an admission that none of the three proves causation alone.

## What "testing before launch" actually means

Most of what gets called ad testing happens before launch, but after the creative decision is already made. Pretest scores are collected on a finished ad. Split tests compare finished variants once media is live. Kantar's own analysis says a single pretest score doesn't reliably predict in-market sales impact; it needs triangulation with sales-linked models before anyone acts on it ([Kantar](https://www.kantar.com/north-america/inspiration/advertising-media/can-ad-testing-really-predict-sales-impact)). EMARKETER frames that triangulation, pretest plus MMM plus incrementality, as the defensible 2026 measurement framework, with incrementality as causal ground truth and platform attribution as a tactical signal only ([EMARKETER](https://www.emarketer.com/content/mmm--incrementality--other-measurement-trends-that-will-define-2026)). Stacking three methods to cover for the fact that none proves causation alone is a reasonable workaround. It is not the same as running one method that identifies causation directly, on the decision a buyer is actually making, before the creative locks.

## Can ad testing really predict sales impact?

Not reliably, if "testing" means a stand-alone pretest score. Kantar's own conclusion: pretest diagnostics need sales-linked models to say anything about downstream impact, because a likability or breakthrough score isn't built to answer a causal question ([Kantar](https://www.kantar.com/north-america/inspiration/advertising-media/can-ad-testing-really-predict-sales-impact)). A survey question asking whether someone likes an ad captures a stated reaction, not a causal test of what caused a change in behavior. The two correlate in aggregate, which is why pretest scores still work as a screen. Correlated is not causal, and that gap is where budget gets wasted on the wrong variant.

## Why pretest scores measure liking, not causation

System1 shipped an AI layer for its Test Your Ad tool in April 2026, trained on its emotional-norms database, to predict likely emotional and distinctiveness response before an ad launches ([PPC Land](https://ppc.land/system1-adds-ai-layer-to-test-your-ad-as-creative-measurement-race-heats-up/)). That's a real gain in speed over manual panel testing, but it scores the same construct faster: how an audience reacts to a finished ad, not what specific element caused a behavior change. This is a version of the say-do gap: what someone reports feeling about an ad and what they actually do when it competes for their attention in market are two different measurements. A working paper on discrete choice modeling documents this divergence: stated preference and observed choice behavior don't move together as reliably as survey-based methods assume ([arXiv](https://arxiv.org/pdf/2307.13966)).

## Why on-platform split tests confound creative with delivery

Once a campaign goes live on Meta or TikTok, a split test compares creative variants inside a system that is simultaneously optimizing delivery, bidding, and audience allocation for each variant. A variant that "wins" may be winning because the algorithm found it a cheaper audience, not because the message performed better. Nothing about a live split test isolates creative from delivery, so the result answers a different question than the one the buyer asked before spend: which message works, holding delivery constant. That isolation requires random assignment of the creative variable alone, independent of platform delivery logic, which a split test does not do.

## Why incrementality testing can't screen creative before the buy

Incrementality testing is the one method in this stack that is genuinely causal: it compares exposed and randomized holdout groups to measure real lift. 52 percent of US brand and agency marketers now use incrementality testing, per a July 2025 TransUnion/EMARKETER survey of 196 marketing professionals, because platform-reported ROAS runs 20 to 60 percent above measured incremental lift ([TransUnion](https://newsroom.transunion.com/new-transunion-research-reveals-marketers-confidence-in-measurement-has-stalled/)). That range is an aggregate across the surveyed categories and channels, not a fixed number for any one brand, but the direction is consistent: platforms overstate their own impact. Incrementality testing is the correct fix for that gap, but it fixes a different problem than the one a buyer faces before the media buy. It validates a single campaign once spend is already running. It cannot screen eight creative variants, four value propositions, or three price framings against each other before any of them are funded, because that would mean running all of them live first, at full media cost, to find out which one to keep.

![A left-to-right timeline showing creative variants drafted, then a randomized pre-spend experiment, then the winning variant getting locked, then budget committed and the campaign going live, followed by on-platform split testing and incrementality testing, both of which happen after launch.](/images/authority/ad-testing-maximising-impact-roi-ads.svg "Only the randomized pre-spend experiment runs before the creative is locked and the budget is committed; every other method runs after that decision is already made.")

## How does a randomized experiment on creative work before spend?

It works by randomly assigning creative variants to a simulated population, validated against real human choice behavior before being trusted on a new market (more below), and measuring which variant causes a change in choice, before any variant is finalized or funded. McFadden discrete choice, Mixed Logit, and ICLV are the estimators used to read the results. They are not what makes the result causal; identification comes from the randomized manipulation in the experiment design, not the modeling technique applied afterward. A flat multinomial logit assumes independence of irrelevant alternatives: relative preference between two variants shouldn't shift just because a third is added. When two variants share a visual hook or compete for the same attention, that assumption breaks down, which is why Mixed Logit models correlated substitution instead of forcing independence. ICLV goes further, linking latent constructs like perceived trust or distinctiveness to the observed choice, connecting the "why" behind a preference to the outcome.

In Subconscious's own validation set, simulated studies reproduced the direction and outcome of the original human study at 93 percent replication accuracy, published at [go.subconscious.ai/paper](https://go.subconscious.ai/paper). That's a validation-set result on studies used to build and check the replication protocol, not a guarantee for a market the model hasn't been tested against. Published human studies can in principle overlap with training data; the replication protocol is built to score against that risk rather than assume it away. The gap between simulated and human results, study by study, is published on the [leaderboard](/leaderboard), so it can be checked directly rather than taken on faith.

## Choosing a method: pretest, split test, incrementality, or randomized pre-spend experiment

| Method | What it actually measures | When it runs | Best for |
|---|---|---|---|
| Pretest tools (Kantar, System1) | Stated emotional response and attention to a finished ad | Before spend, after the creative is already built | Best for: screening out obvious creative failures, like a confusing message or low attention, before production is finalized. |
| On-platform split test (Meta, TikTok) | Platform-reported performance of finished variants, entangled with the platform's own delivery and bidding | After budget is committed and the campaign is live | Best for: comparing delivery-adjusted media performance once the creative direction is already chosen. |
| Incrementality test | Causal lift of the live campaign versus a randomized holdout of real users | After launch, with real spend running | Best for: proving how much a live campaign actually moved sales, for renewal and budget-scaling decisions. |
| Randomized pre-spend experiment (discrete choice models) | The causal effect of specific creative elements on choice, isolated from platform delivery | Before the creative is locked, before any spend | Best for: choosing which of several creative variants, value propositions, or price framings to fund before the media buy. |

## What to do before your next media buy

Run the shortlist, not just the winner, through a randomized experiment before any variant is locked or funded. That's the only method here that compares variants without paying to run each one live. Treat the on-platform split test as a delivery check on the variant already chosen, not as evidence about which message works. Keep the incrementality test where it belongs, after launch, to confirm at scale what the pre-spend experiment already flagged. For the estimators and validation protocol behind this, see [methods and validation](/blog/methods-and-validation) and the [case studies](/case-studies).

The concrete next step: before your next creative sprint locks a shortlist, check the published study-by-study accuracy on the [leaderboard](/leaderboard) and decide whether that bar is high enough for the decision you're making. If you want to see this run against your own shortlist, [the team](/meet) can walk through it.

Every listed problem addressed: H1 removed from body, title now names "before spend" and "which variant to fund," the unsupported "grown fast" trend claim cut to a level statement, Zappi dropped (uncited) from both the bullet and the table, the flat unsourced brand-growth assertion removed, bullet one reframed around the mechanism instead of a definitional claim, and "simulated population" now carries an immediate validation anchor before the estimator discussion. Nothing else was restructured.