Claims Test Methodology
Claims test methodology decides which candidate message gets media spend, and for a brand leader or general counsel, that decision carries real risk on both sides. Fund the wrong claim and it flops at shelf; fund an unsubstantiated one and it draws an FTC inquiry. The dominant approach today, stated-preference scoring through monadic, sequential-monadic, and MaxDiff surveys, measures how a claim sounds, not what it does. That gap is real: a claim that ties for first place on a five-point believability scale can still move actual purchases very differently, exactly the pattern the peer-reviewed literature on hypothetical bias documents.
- Standard claims testing (monadic exposure, sequential-monadic, MaxDiff) measures comprehension, believability, uniqueness, and stated purchase intent, all self-reported, none of them a measured behavior.
- Hypothetical bias research finds stated willingness-to-pay and stated preference run 1.2 to over 3 times higher than what the same people do when money is actually on the line, a wide range because it pools many different domains and elicitation methods (arXiv).
- FTC substantiation doctrine, in force since the agency's 1983 policy statement, requires "competent and reliable evidence" that a claim affects the belief or behavior it makes, which a correlational appeal score does not supply (FTC).
- A randomized design, where exposure to the claim is assigned and compared against a holdout, is what identifies a causal effect; McFadden discrete choice models, Mixed Logit, and ICLV are the estimators used to read that design, not the source of causal identification.
- Any simulated read on a claim should be checked against a measured human baseline before it drives a launch decision, with the gap to that baseline reported openly (leaderboard).
What is claims test methodology?
Claims test methodology is the experimental design used to compare candidate marketing claims before one is chosen for launch, and in current practice that almost always means a stated-preference survey. Brands running 5 to 20 candidate claims typically expose them through platforms like SurveyMonkey LaunchPad's monadic design or score them with MaxDiff, which forces respondents to trade off claims against each other and is more sample-efficient than testing one claim per cell (Dig Insights). The question flow is standardized across vendors: comprehension, appeal, believability, uniqueness, and purchase intent, usually on a Likert or best-worst scale (Quantilope). That flow answers "which claim do people say they like." It does not answer "which claim changes what people buy," and those are different questions with different evidence standards.
Why doesn't a high believability score predict sales?
A high believability score doesn't predict sales because it captures a stated preference, and stated preference is a known, measured distortion of real behavior. The meta-analytic literature on hypothetical bias finds that stated willingness-to-pay and willingness-to-accept run roughly 1.2 to 3.13 times higher than real-payment choices, with consumer-behavior domains showing some of the largest gaps; that range is wide because it spans many different domains and elicitation methods, not a single reliable multiplier for any one claim (arXiv). This is the same say-do gap that lets a claim sound compelling on a survey screen and underperform at shelf: nothing about answering a question on a screen costs the respondent anything, so there is no pressure to answer the way they would if a purchase were actually at stake. The message itself, not the product, is usually where a claims-testing process needs to catch a failure before launch, and a stated-preference score is the tool least equipped to catch it.
Why can't monadic and MaxDiff designs isolate one claim's causal lift?
Monadic and MaxDiff designs can't isolate one claim's causal lift because neither controls for the other forces shaping a respondent's answer: order effects, brand halo, and fatigue across a long claim list. In a MaxDiff task, a claim's rank depends partly on what it's traded off against and where it sits in the sequence, not purely on its own wording. In monadic and sequential-monadic exposure, a respondent who has already rated four claims answers the fifth differently than a fresh respondent would. Two claims can land in a statistical tie on a 5-point scale for reasons that have nothing to do with which one would actually move a shopper, and the survey has no way to separate "these are equally good claims" from "these tied because of the design, not the wording." A randomized design solves this directly: assign respondents to see one claim or none, hold everything else constant, and the difference in outcome between the two groups is attributable to the claim itself.
What does FTC substantiation actually require?
FTC substantiation requires that an advertiser hold a "reasonable basis," meaning competent and reliable evidence, for every express and implied claim before it runs, a standard the agency has enforced since its 1983 policy statement (FTC). That standard asks a causal question: does the evidence show the claim affects the belief or behavior it makes? A mean appeal score, say 4.2 out of 5, doesn't answer that question, because it describes what people said about the claim, not what the claim did to them. A brand relying on a stated-preference score as its substantiation file is relying on evidence that answers the wrong question.
Three ways to test a claim, compared
| Method | What it measures | Isolates one claim's causal effect? | Best for |
|---|---|---|---|
| Monadic exposure | Stated appeal and intent for one claim per respondent group | No, vulnerable to group differences and no counterfactual | Quick directional read on comprehension when legal exposure is low |
| MaxDiff | Relative stated preference across a large claim set | No, rank depends on what each claim is traded against | Cutting a long list of 15 to 20 claims down to a short list cheaply |
| Randomized causal experiment | Behavioral difference between an exposure group and a holdout, analyzed with discrete choice models | Yes, randomization is what identifies the causal effect | The final claim a board approves or a legal team has to substantiate |
How does a randomized causal claims test work?
A randomized causal claims test works by assigning otherwise-identical respondent groups to see a candidate claim or nothing, then measuring the difference in a real choice outcome between them. The estimator layer is McFadden discrete choice models, Mixed Logit, or ICLV; these read the choices, but the causal identification comes from the randomized assignment in the design itself, not from the estimator. When the analysis compares preference share across multiple claims with a flat multinomial logit, it carries the independence of irrelevant alternatives assumption; Mixed Logit relaxes that assumption when substitution patterns among claims are expected to differ across respondents.
A confidence interval produced this way covers the estimated effect within the simulated population the experiment ran on. It does not, by itself, bound what will happen in the real market until it's checked against a human baseline.
How much can you trust a simulated replication of a human study?
You can trust it to the extent it's been measured against real people on the same study, and that measurement should be reported with its denominator, not as a bare percentage. On causal fidelity work, the best configuration reaches 0.832 rank correlation against a published human result on one study, against a human ceiling of 0.959 set by two independent human samples, which is 87% of that measured ceiling; across all 43 studies that passed design filters, the mean is 0.73 (causal fidelity paper). Two limitations matter here. First, this is a validation result on studies that have already run, not a guarantee for a new, unpublished market. Second, published human studies can sit inside a model's training data, which is a real risk to any claim of "held-out" validation; the replication protocol behind these numbers is built to address that, but the risk doesn't disappear because it's addressed. Current results by study and method are published on the leaderboard, and the underlying methodology is covered in more depth in the methods and validation hub.
The decision
Use stated-preference tools where they're actually good at their job: MaxDiff to cut a list of 15 to 20 claims down to a short list, cheaply and fast. But before a claim gets media budget or has to survive an FTC substantiation request, it needs a randomized experiment behind it, not another appeal score. If two claims tied on purchase intent in your last test, ask whether you'd actually bet the launch budget that they'd perform identically at shelf. If the answer is no, the survey score already told you it isn't sufficient evidence.
As a concrete next step, pull your last claims-test dataset and look for any two claims within half a point of each other on purchase intent. That pair is exactly the case a stated-preference score can't resolve and a randomized design can. If you want to see how that design would run on your specific claim set, the team is reachable at meet.