Claims Test Methodology
Claims test methodology decides which candidate message gets media spend, and for a brand leader or general counsel, that decision carries two separate risks. Fund the wrong claim and it flops at shelf. Run a claim without a reasonable basis for its truth and it can draw an FTC inquiry. Those are different problems with different evidence. The dominant approach to the first, stated-preference scoring through monadic, sequential-monadic, and MaxDiff surveys, measures how a claim sounds. A claim that ties for first place on a five-point believability scale can still move purchases very differently from another, and the literature on hypothetical bias shows why a stated score cannot settle it alone.
- Standard claims testing measures comprehension, believability, uniqueness, and stated purchase intent, all self-reported, none of them measured behavior.
- Hypothetical bias research finds mixed evidence on how large the stated-versus-real gap is, and it depends heavily on context and measurement (Haghani et al., arXiv). Believability alone does not establish sales lift.
- FTC substantiation concerns whether an advertiser has a reasonable basis for the truth of objective product claims before it runs them. It is a separate question from whether a claim persuades (FTC policy statement, November 23, 1984).
- A randomized design, in which exposure to a claim is assigned and compared against a control, identifies the effect of the claim on a defined endpoint. Monadic exposure can be built this way. McFadden discrete choice models, Mixed Logit, and ICLV are analysis models for the choices, not the source of identification.
- Randomization does not turn a stated intent into an observed sale. Any simulated read on a claim should be checked against a measured human result before it drives a launch decision.
What is claims test methodology?
Claims test methodology is the experimental design used to compare candidate marketing claims before one is chosen for launch. In current practice that usually means a stated-preference survey. Monadic testing, offered for example in SurveyMonkey's message testing, has each respondent evaluate one message in isolation, which the vendor says reduces comparison bias (SurveyMonkey, checked October 2, 2026). Sequential-monadic designs have each respondent evaluate several concepts in turn. MaxDiff forces respondents to trade off claims against each other and returns a rank order..
A survey can evaluate relevance, believability, clarity, memorability, uniqueness and purchase intent. That flow answers "which claim do people say they like." It does not answer "which claim changes what people buy," and those are different questions with different evidence standards.
Why doesn't a high believability score establish sales lift?
A believability score captures a stated response, and a stated response is not a measured purchase. The literature on hypothetical bias in stated choice experiments is mixed on how large the gap is. A review by Haghani and colleagues finds that health-related choice experiments often show negligible bias, while experiments in consumer behaviour and transport suggest significant bias is common, with "considerable context and measurement dependency" (Haghani, Bliemer, Rose, Oppewal and Lancsar, arXiv).
No single multiplier converts a stated score into expected sales, and the direction and size depend on the task. The practical limit is that a hypothetical question need not impose the consequences of an actual purchase. Believability can still be a useful diagnostic, because a claim nobody believes is unlikely to help. It does not establish that a believed claim will change purchases. Predictive value for a given category is a calibration question that needs its own evidence.
What can monadic, sequential-monadic and MaxDiff designs identify?
It depends on assignment, control and endpoint, more than on question format.
- Monadic exposure. If respondents are randomly assigned to one claim each, with a control group that sees no claim or a baseline, the difference between groups estimates that claim's effect on the measured endpoint. Without a control, or with groups that differ, it gives a read on each claim in isolation.
- Sequential-monadic exposure. A respondent who has already rated four claims answers the fifth in light of the first four, so order effects, halo and fatigue become part of the answer. Randomizing order helps but does not remove them.
- MaxDiff. A claim's rank depends partly on what it is traded off against and where it sits in the sequence. It supports prioritizing a long list. It does not estimate the effect of exposure to one claim against none.
In all three, the measured endpoint is still a survey response. A randomized design supports a causal claim about that response. It does not turn hypothetical intent into observed sales.
What does FTC substantiation actually require?
The FTC's policy statement of November 23, 1984 says advertisers and agencies need a reasonable basis for advertising claims before they are disseminated. It concerns objective claims, express and implied, that make factual assertions about products or services. What counts as a reasonable basis depends on the type of claim and product, the consequences of a false claim, the benefits of a truthful one, the cost of developing substantiation and what experts in the field consider reasonable (FTC).
That standard is about whether the claim is true, for example whether a product delivers the benefit it asserts. It does not generally ask whether the advertisement changes belief or behavior. Message-effect testing and substantiation answer separate questions. Counsel decides what evidence a given claim needs. A preference or message-performance study is evidence about persuasion, and it should not be presented as a legal substantiation file for the truth of a product claim.
Three ways to test a claim, compared
| Method | What it measures | Identifies one claim's effect on the endpoint? | Best for |
|---|---|---|---|
| Monadic exposure | Stated appeal and intent for one claim per respondent group | Only if assignment is random and a control group exists | Directional read on comprehension and appeal when legal exposure is low |
| MaxDiff | Relative stated preference across a large claim set | No, rank depends on what each claim is traded against | Cutting a long list of 15 to 20 claims to a short list |
| Randomized claims experiment | Difference in a stated or observed choice between an exposure group and a control, analyzed with a discrete choice model | Yes, for the tested endpoint, because assignment is random | The claim a team is deciding to fund, with a defined endpoint |
How does a randomized claims test work?
A randomized claims test assigns otherwise-comparable respondent groups to see a candidate claim or nothing, then measures the difference in a choice outcome between them. The analysis layer is McFadden discrete choice models, Mixed Logit, or ICLV. These read the choices, and the identification comes from the randomized assignment in the design. When the analysis compares preference share across multiple claims with a flat multinomial logit, it carries the independence of irrelevant alternatives assumption. Mixed Logit relaxes that assumption when substitution patterns among claims are expected to differ across respondents.
A confidence interval produced this way covers the estimated effect within the population the experiment ran on. It does not automatically bound the real market response. A matched human comparison is separate evidence, and a comparison on human stated choices still does not measure purchases.
How much can you trust a simulated replication of a human study?
To the extent it has been measured against real people, and the measurement should state what was compared. The July 2026 causal fidelity working paper, which is not peer reviewed, reports Spearman rank correlations between parameters estimated from simulated and human choices. The mean is 0.73 across the 43 studies that passed its design filters, and 0.55 across roughly 300 replications. These are parameter-rank correlations, not accuracy percentages and not agreement on effect sizes. They describe studies that have already run, not a new, unpublished market. Published human studies can also sit inside a model's training data, a risk to any claim of held-out validation. The replication protocol is built to address that, and the risk does not disappear because it is addressed. Published results are on the research page, and the method is covered further in the methods and validation hub.
The decision
Use stated-preference tools where they fit: MaxDiff to prioritize candidates from a long list. Before a claim gets media budget, test the shortlist in a randomized design with a defined endpoint, and keep any legal substantiation of product claims as a separate workstream with counsel. If two claims tied on purchase intent in your last test, ask whether you would bet the launch budget that they would perform identically at shelf.
As a concrete next step, pull your last claims-test dataset and compare the uncertainty around the leading claims. If the scores do not distinguish the candidates, define a controlled contrast and endpoint before commissioning another test. To see how that design would run on your claim set, book a decision review.