Skip to content

Comparing different approaches to claims' diagnostics

Title picked: "Comparing Claims Diagnostics: What Each Method Actually Answers"

Article body below, per brief.

A marketing or legal lead deciding how to test a claim before launch is really deciding which question they need answered. Stated-preference research such as MaxDiff, monadic testing, and TURF analysis answers only one of them: which claim people say they would pick. A randomized, controlled comparison with confidence intervals answers the other one, and that is the question FTC substantiation law and go-to-market risk actually turn on.

- Stated-preference methods (MaxDiff, monadic testing, TURF analysis) rank which claim respondents say they like best. They do not measure what a claim causes someone to do.
- FTC claims substantiation requires "competent and reliable scientific evidence," and for health-related claims the agency expects randomized, controlled human clinical trials, not surveys or focus groups.
- AI and synthetic-respondent panels are a newer third option. Vendor accuracy figures are self-reported against undisclosed benchmarks, and no accepted framework exists yet for evaluating them.
- Randomized experiments analyzed with discrete choice models reach 87% of the measured human ceiling on one study's best configuration, with a mean of 73% across the 43 studies passing design filters. The ratio and its source are covered below.
- The decision that matters isn't speed versus budget. It's matching the method to what the claim needs to prove: stated appeal, legal substantiation, or a causal effect on behavior.

## What are the three ways to test a marketing claim right now?

Claims diagnostics splits into three tracks that rarely talk to each other. Market research shops (Conjointly, Toluna, Quantilope, Dig Insights, IMS/G&R, SIS International) run monadic tests, MaxDiff/best-worst scaling, and TURF analysis to rank which of a set of claims resonates most with respondents. It's fast, cheap, stated-preference work, and practitioners typically sequence it as a funnel: MaxDiff to shortlist candidates, then monadic or TURF testing to check reach and standalone performance of the finalists, with no causal comparison step built in ([Medium](https://medium.com/@kunalmra/whats-the-order-of-these-methods-maxdiff-turf-and-monadic-testing-c61f6965a8b0)). Survey design itself constrains this track further: claims questionnaires have to stay short and attribute-specific, because later statements risk bias from earlier ones, and preference-format questions can't be combined with sequential monadic ratings in the same instrument ([G&R primer](https://www.gandrllc.com/on-our-minds/claims-substantiation-research/)).

On the second track, FTC substantiation law requires "competent and reliable scientific evidence," research conducted and evaluated objectively by experts using methods generally accepted in the field, and for health-related claims the agency explicitly expects randomized, controlled human clinical trials. Anecdotal evidence and testimonials are excluded outright ([Covington](https://www.cov.com/en/news-and-insights/insights/2023/01/ftc-issues-new-guidance-on-health-related-claims-to-replace-the-dietary-supplements-advertising-guide)).

2025 and 2026 added a third track: AI and synthetic respondents. Toluna markets synthetic personas for claims testing with accuracy claims up to 90% against human validation tests ([Toluna](https://www.globenewswire.com/news-release/2026/04/14/3273025/0/en/toluna-harnesses-ai-to-transform-the-speed-and-scale-of-claims-testing.html)), and other vendors cite figures in a similar range. But Dig Insights and independent reviewers note that no accepted framework exists for evaluating synthetic-respondent accuracy, that vendor percentages often lack transparent methodology or peer review, and that synthetic panels show sycophantic bias and underrepresented variance against real populations ([Dig Insights](https://diginsights.com/resources/truth-about-synthetic-data/)). Buyers are left choosing between speed, legal defensibility, or something in between.

## Is picking a method just a speed-versus-rigor tradeoff?

That's the standard advice, and it skips the one thing that matters: what the claim needs to prove. The typical claims-testing article treats method choice as a tradeoff: use MaxDiff or monadic surveys for quick directional ranking of many claims, escalate to an RCT-style substantiation study only when legal risk demands it, and increasingly treat synthetic respondents as a faster stand-in for either. The implicit advice is to pick a tool based on budget and timeline, not on what the claim needs to prove.

Stated-preference methods answer the first question: MaxDiff, monadic testing, and TURF analysis show which claim respondents say they like best. They don't answer the second one, which is what a claim causes someone to do, purchase, switch, or pay more. The FTC's own standard is explicit here (see above): preference and liking scores are not evidence of a claim's truth or its causal effect on behavior. Only a controlled, randomized comparison against a baseline, a claim versus no claim or a claim versus an alternative claim, isolates that effect. Claims substantiation research also requires a defensible design tied to the specific claim language, distinct from the general concept or message testing used for marketing optimization ([Decision Analyst](https://www.decisionanalyst.com/blog/derivingmarketingclaims/)).

Synthetic-respondent testing inherits the same gap and adds a second one: its correlation figures are self-reported against undisclosed benchmarks, with no peer-reviewed validation and documented sycophancy and variance-collapse problems (see above). A buyer choosing between monadic surveys and synthetic panels is choosing between two flavors of correlation, and neither answers the causal question, isolated from confounds, that substantiation and go-to-market decisions actually require.

| Method | What it actually measures | Source and limitation | Best for |
|---|---|---|---|
| MaxDiff, monadic testing, TURF | Stated preference: which claim respondents say they'd pick | Preference-format items can't combine with sequential monadic ratings in one survey ([G&R primer](https://www.gandrllc.com/on-our-minds/claims-substantiation-research/)) | Best for: fast, cheap ranking of many claim candidates before legal exposure exists |
| AI / synthetic respondent panels | Vendor-reported correlation with human survey answers, against undisclosed benchmarks | No accepted accuracy framework exists; documented sycophancy and variance collapse (see above) | Best for: directional screening only, not for any claim carrying legal risk |
| RCT-style human substantiation trials | Causal effect of the claim, isolated against a no-claim or alternative-claim baseline | Required by the FTC for health-related claims (see above); anecdote and testimonial evidence excluded | Best for: claims that must survive an FTC substantiation review |
| Randomized experiments analyzed with discrete choice models (McFadden DCE, Mixed Logit, ICLV) | Causal effect of a claim, estimated from a randomized comparison, not a bare preference score | Validated at 87% of the measured human ceiling on one study, mean 73% across 43 studies passing design filters; see below | Best for: teams that need causal evidence faster and cheaper than a full human RCT |

## How accurate are AI and synthetic panels for claims testing?

There's no accepted framework yet for measuring it, so a vendor's percentage isn't comparable to anything else in this market. Toluna's claim of up to 90% correlation and similar figures from other vendors are self-reported against benchmarks the vendors themselves define, without peer review (see above). Dig Insights and other independent reviewers have also documented sycophantic response bias and compressed variance in synthetic panels relative to real populations. None of that establishes what "accuracy" would even mean for a claims-diagnostics tool, since correlation with stated survey answers still isn't a measure of causal effect on behavior.

## What does an 87%-of-ceiling replication rate actually mean?

It means our best configuration's rank correlation against a published human result is 0.832, and two independent samples of real humans reach 0.959 with each other on that same study, so 0.832 is 87% of that measured ceiling. Across all 43 studies passing design filters, the mean is 0.73, or 73% of the comparable ceiling. That ratio, measured against real human data rather than a self-reported benchmark, is what we mean by replication accuracy ([causal fidelity paper](https://fidelity.subconscious.ai/papers/causal-fidelity/causal_fidelity_paper.pdf)). It is a validation result on studies run to date, not a guarantee for a new market.

![Bar chart showing two ratios against a human ceiling: a best-configuration rank correlation of 0.832 against a human ceiling of 0.959, which is 87%, and a mean of 0.73 across 43 studies, which is 73%.](/images/authority/comparing-different-approaches-claims-diagnostics.svg "The replication rate is a ratio against a measured human ceiling: 87% of it on the best configuration in one study, and 73% on average across 43 studies passing design filters.")

Published human studies can sit in a model's training data. The replication protocol holds out studies and rank-correlates against independently collected human responses to limit that risk, but the risk doesn't disappear because a protocol exists. A confidence interval from one of these randomized experiments covers the effect within the simulated population that experiment ran on. It does not by itself bound what happens once a claim ships in the live market. For the underlying method mechanics, see [methods and validation](/blog/methods-and-validation); current results across studies are on the public [leaderboard](/leaderboard).

## Which method should you actually pick?

Match the method to what the claim needs to prove. If you're ranking many claim candidates before any legal exposure exists, MaxDiff or monadic testing is the right cheap first pass. If a claim is heading toward an FTC substantiation review, a stated-preference score or a synthetic panel's correlation figure won't hold up, and only a randomized, controlled comparison against a baseline answers the question the agency asks. Randomized experiments analyzed with discrete choice models sit between a full human RCT and a stated-preference survey: causal by design, validated against a measured human ceiling, faster and cheaper than a human trial.

As a next step, pull the specific claim language you're testing and ask which question it actually needs answered: preference, or causal effect on behavior. If it's the second one, check the [leaderboard](/leaderboard) for current study results before choosing a vendor. If you want a second opinion on your specific claim and market, [talk to us](/meet).