Skip to content
Subconscious

Claims Test Methodology

Claims test methodology decides which candidate message gets media spend, and for a brand leader or general counsel, that decision carries two separate risks. Fund the wrong claim and it flops at shelf. Run a claim without a reasonable basis for its truth and it can draw an FTC inquiry. Those are different problems with different evidence. The dominant approach to the first, stated-preference scoring through monadic, sequential-monadic, and MaxDiff surveys, measures how a claim sounds. A claim that ties for first place on a five-point believability scale can still move purchases very differently from another, and the literature on hypothetical bias shows why a stated score cannot settle it alone.

What is claims test methodology?

Claims test methodology is the experimental design used to compare candidate marketing claims before one is chosen for launch. In current practice that usually means a stated-preference survey. Monadic testing, offered for example in SurveyMonkey's message testing, has each respondent evaluate one message in isolation, which the vendor says reduces comparison bias (SurveyMonkey, checked October 2, 2026). Sequential-monadic designs have each respondent evaluate several concepts in turn. MaxDiff forces respondents to trade off claims against each other and returns a rank order..

A survey can evaluate relevance, believability, clarity, memorability, uniqueness and purchase intent. That flow answers "which claim do people say they like." It does not answer "which claim changes what people buy," and those are different questions with different evidence standards.

Why doesn't a high believability score establish sales lift?

A believability score captures a stated response, and a stated response is not a measured purchase. The literature on hypothetical bias in stated choice experiments is mixed on how large the gap is. A review by Haghani and colleagues finds that health-related choice experiments often show negligible bias, while experiments in consumer behaviour and transport suggest significant bias is common, with "considerable context and measurement dependency" (Haghani, Bliemer, Rose, Oppewal and Lancsar, arXiv).

No single multiplier converts a stated score into expected sales, and the direction and size depend on the task. The practical limit is that a hypothetical question need not impose the consequences of an actual purchase. Believability can still be a useful diagnostic, because a claim nobody believes is unlikely to help. It does not establish that a believed claim will change purchases. Predictive value for a given category is a calibration question that needs its own evidence.

What can monadic, sequential-monadic and MaxDiff designs identify?

It depends on assignment, control and endpoint, more than on question format.

In all three, the measured endpoint is still a survey response. A randomized design supports a causal claim about that response. It does not turn hypothetical intent into observed sales.

What does FTC substantiation actually require?

The FTC's policy statement of November 23, 1984 says advertisers and agencies need a reasonable basis for advertising claims before they are disseminated. It concerns objective claims, express and implied, that make factual assertions about products or services. What counts as a reasonable basis depends on the type of claim and product, the consequences of a false claim, the benefits of a truthful one, the cost of developing substantiation and what experts in the field consider reasonable (FTC).

That standard is about whether the claim is true, for example whether a product delivers the benefit it asserts. It does not generally ask whether the advertisement changes belief or behavior. Message-effect testing and substantiation answer separate questions. Counsel decides what evidence a given claim needs. A preference or message-performance study is evidence about persuasion, and it should not be presented as a legal substantiation file for the truth of a product claim.

Three ways to test a claim, compared

MethodWhat it measuresIdentifies one claim's effect on the endpoint?Best for
Monadic exposureStated appeal and intent for one claim per respondent groupOnly if assignment is random and a control group existsDirectional read on comprehension and appeal when legal exposure is low
MaxDiffRelative stated preference across a large claim setNo, rank depends on what each claim is traded againstCutting a long list of 15 to 20 claims to a short list
Randomized claims experimentDifference in a stated or observed choice between an exposure group and a control, analyzed with a discrete choice modelYes, for the tested endpoint, because assignment is randomThe claim a team is deciding to fund, with a defined endpoint

How does a randomized claims test work?

A randomized claims test assigns otherwise-comparable respondent groups to see a candidate claim or nothing, then measures the difference in a choice outcome between them. The analysis layer is McFadden discrete choice models, Mixed Logit, or ICLV. These read the choices, and the identification comes from the randomized assignment in the design. When the analysis compares preference share across multiple claims with a flat multinomial logit, it carries the independence of irrelevant alternatives assumption. Mixed Logit relaxes that assumption when substitution patterns among claims are expected to differ across respondents.

A five-step diagram showing a candidate claim randomly assigned to an exposure group or a holdout group, a choice measured in both groups, and the difference between them estimated as the claim's effect with a discrete choice model.
The estimated effect is the gap between a randomized exposure group and a holdout that never saw the claim, not one group's average score.

A confidence interval produced this way covers the estimated effect within the population the experiment ran on. It does not automatically bound the real market response. A matched human comparison is separate evidence, and a comparison on human stated choices still does not measure purchases.

How much can you trust a simulated replication of a human study?

To the extent it has been measured against real people, and the measurement should state what was compared. The July 2026 causal fidelity working paper, which is not peer reviewed, reports Spearman rank correlations between parameters estimated from simulated and human choices. The mean is 0.73 across the 43 studies that passed its design filters, and 0.55 across roughly 300 replications. These are parameter-rank correlations, not accuracy percentages and not agreement on effect sizes. They describe studies that have already run, not a new, unpublished market. Published human studies can also sit inside a model's training data, a risk to any claim of held-out validation. The replication protocol is built to address that, and the risk does not disappear because it is addressed. Published results are on the research page, and the method is covered further in the methods and validation hub.

The decision

Use stated-preference tools where they fit: MaxDiff to prioritize candidates from a long list. Before a claim gets media budget, test the shortlist in a randomized design with a defined endpoint, and keep any legal substantiation of product claims as a separate workstream with counsel. If two claims tied on purchase intent in your last test, ask whether you would bet the launch budget that they would perform identically at shelf.

As a concrete next step, pull your last claims-test dataset and compare the uncertainty around the leading claims. If the scores do not distinguish the candidates, define a controlled contrast and endpoint before commissioning another test. To see how that design would run on your claim set, book a decision review.