How to Test Patient Messaging in a Regulated Environment
Patient messaging under FDA oversight rarely gets tested where it matters. A brand team choosing which patient-directed message to build a campaign on needs the comparison to run before the message goes into legal and medical review, not after. Testing after review only ranks whichever executions already survived that process, so the result tells the team which survivor scored best - not which message, from the full set the team debated, deserves the media budget and can carry its required risk information.
- Randomize the full candidate set before review, not the shortlist review already narrowed, so the evidence sets the option space instead of the review queue.
- Keep the required risk information present in every version tested, since a message that can't carry its major statement can't ship regardless of how well it scores.
- A monadic scale (appeal, believability, intent) measures how a message reads, not whether it changes what happens in a clinical encounter.
- An in-market split arrives after the spend and after a prescriber has mediated the outcome, which is too late to inform the choice.
- Check any validation number against its denominator: the strongest published result here is 0.832 against a measured human ceiling of 0.959, one study, best configuration, with a mean of 0.73 across the 43 studies that pass design filters.
Why review order decides more than the test does
The conventional workflow drafts several executions, sends them through legal and medical review, and only tests whatever comes out the other side. By the time a monadic test with appeal, believability, and intent scales runs, the option space has already been set by which executions cleared review, not by which ones would perform best. The test then answers a narrower question than the one the brand team actually faces: instead of "which message should we build the campaign on," it answers "which of the messages that already survived review scores highest." Those are different decisions, and only one of them is open when the test runs.
Why doesn't a rating scale predict what a message does in a clinical encounter?
A rating scale predicts how a message reads to the person taking the survey, not what a patient does with it in front of a prescriber. Appeal, believability, and intent are self-report constructs: a respondent tells you how convincing a message seemed, not whether it changed the question they asked their doctor or the request they made at the counter. Patient-directed messaging earns its media budget by moving behavior inside a clinical encounter the survey never observes. A message can score well on all three scales and still fail to move that behavior, because the scales were never built to measure it.
Why does an in-market split arrive too late to guide the choice?
An in-market split test only reports its result after the media has already been bought and aired. By then the budget decision the test exists to inform has already been made. The outcome it measures - a click, a recall lift, a site visit - is also mediated by a prescriber, since the patient's next step almost always runs through a clinical conversation the ad itself doesn't control. A click-through rate is several steps removed from the decision the campaign is trying to influence, and the delay means the evidence lands after the money is spent rather than before.
What changes when the message is randomized before submission
Randomizing the message before it goes to review changes what the test can answer. Instead of testing whichever executions survived review, the team varies the specific elements under debate - claim framing, tone, the placement and weight of risk information - and holds the required risk information present in every arm, so what gets tested is what could actually be submitted. Each element's effect is estimated with a confidence interval rather than a single score, and that interval describes the effect within the tested population, not a guarantee about the real market.
Two things follow from moving the test earlier. The choice happens before submission, so review time gets spent on the message the evidence already supports rather than on adjudicating between finalists nobody tested against the full field. And the process leaves a stated design and a set of effect estimates, which is the form of support a team can hand to its own legal and medical reviewers - not proof of compliance, but a documented basis for the choice they're being asked to approve.
What the major statement rule requires of every arm you test
FDA's rule on the major statement in broadcast advertisements requires that the statement be presented in a clear, conspicuous, and neutral manner - not buried in pacing, visuals, or competing audio that undercuts it (Federal Register vol. 88, no. 223, Nov. 21, 2023, p. 80958: govinfo.gov). The rule took effect May 20, 2024, with a compliance date of November 20, 2024. If a message tested for appeal or intent doesn't carry its risk information the same way it would need to on air, the test result describes a version of the message that can't legally run. Holding the risk information constant and present in every arm during the test, rather than adding it back in after a winner is chosen, is what keeps the test measuring something that can actually ship. Whether a specific execution meets the clear-conspicuous-neutral standard is a determination for the team's own legal and regulatory reviewers, not a claim any test output can make on its own.
How reliable is a simulated randomized comparison?
The strongest published result for this approach is a rank correlation of 0.832 against a measured human-to-human ceiling of 0.959 on the Hainmueller immigration conjoint study - one study, the best configuration reported, which works out to roughly 87% of that measured ceiling (causal fidelity paper, PDF). Across the full set of 43 published randomized studies that pass design filters, the mean rank correlation is 0.73, and that broader number is the more honest baseline to plan against, not the single-study best case. A limitation worth stating plainly: published human studies can sit inside a model's training data, which would let it pattern-match a known result rather than predict a new one; the replication protocol behind these figures screens for that risk, but the general problem of testing against material a model may have seen before doesn't disappear just because a protocol exists.
The estimators behind the comparison - McFadden discrete choice, Mixed Logit, and ICLV - are statistical models for choice data, not causal methods on their own. Causal identification comes from randomizing the message elements in the experiment design; the estimator's job is to turn the resulting choices into effect sizes. Mixed Logit is used specifically because it relaxes the independence-of-irrelevant-alternatives assumption a flat multinomial logit imposes, which matters when patient messages are close substitutes for each other. Current results across studies and methods are tracked on the leaderboard, and the underlying design questions are covered in more depth on the methods and validation hub.
Three ways to test a patient message
| Monadic test on survivors | In-market split | Randomized comparison before submission | |
|---|---|---|---|
| What it measures | Self-reported appeal, believability, intent on messages already cleared by review | Actual click or recall lift, in market, after review and after spend | Estimated effect of each message element, with risk information present in every arm |
| When it runs | After legal and medical review narrows the field | After production and media spend | Before the message is submitted to review |
| What sets the option space | The review queue | The review queue plus the media plan | The full candidate set the brand team is actually debating |
| What it leaves behind | A score on a scale | An outcome, after the money is spent | A stated design and effect estimates the team's own reviewers can examine |
| Best for | Teams that already trust their review-narrowed shortlist and only need a tiebreaker | Teams that need a real-world read after launch and can absorb the cost of a wrong first choice | Teams choosing which message to build the campaign on, before the review cycle starts |
What to do next
Take the messages currently competing for the brand team's decision and list, for each one, whether the required risk information is present in the version being discussed - not a placeholder version, the one that would actually run. Any message where that's not true isn't ready to be tested yet, regardless of how well it might score. Once the candidate set carries its risk information consistently, that's the set worth comparing, before any of them goes to review. For a closer look at how that comparison is designed and validated, meet with the team.