Skip to content

How to Test Patient Messaging in a Regulated Environment

Patient messaging under FDA oversight rarely gets tested where it matters. A brand team choosing which patient-directed message to build a campaign on needs the comparison to run before the message goes into legal and medical review, not after. Testing after review only ranks whichever executions already survived that process, so the result tells the team which survivor scored best - not which message, from the full set the team debated, deserves the media budget and can carry its required risk information.

Why review order decides more than the test does

The conventional workflow drafts several executions, sends them through legal and medical review, and only tests whatever comes out the other side. By the time a monadic test with appeal, believability, and intent scales runs, the option space has already been set by which executions cleared review, not by which ones would perform best. The test then answers a narrower question than the one the brand team actually faces: instead of "which message should we build the campaign on," it answers "which of the messages that already survived review scores highest." Those are different decisions, and only one of them is open when the test runs.

Why doesn't a rating scale predict what a message does in a clinical encounter?

A rating scale predicts how a message reads to the person taking the survey, not what a patient does with it in front of a prescriber. Appeal, believability, and intent are self-report constructs: a respondent tells you how convincing a message seemed, not whether it changed the question they asked their doctor or the request they made at the counter. Patient-directed messaging earns its media budget by moving behavior inside a clinical encounter the survey never observes. A message can score well on all three scales and still fail to move that behavior, because the scales were never built to measure it.

Why does an in-market split arrive too late to guide the choice?

An in-market split test only reports its result after the media has already been bought and aired. By then the budget decision the test exists to inform has already been made. The outcome it measures - a click, a recall lift, a site visit - is also mediated by a prescriber, since the patient's next step almost always runs through a clinical conversation the ad itself doesn't control. A click-through rate is several steps removed from the decision the campaign is trying to influence, and the delay means the evidence lands after the money is spent rather than before.

What changes when the message is randomized before submission

Randomizing the message before it goes to review changes what the test can answer. Instead of testing whichever executions survived review, the team varies the specific elements under debate - claim framing, tone, the placement and weight of risk information - and holds the required risk information present in every arm, so what gets tested is what could actually be submitted. Each element's effect is estimated with a confidence interval rather than a single score, and that interval describes the effect within the tested population, not a guarantee about the real market.

A four-step flow showing draft, then review narrowing options, then testing only the narrowed set, then production, illustrating that the conventional test never sees the full candidate pool.
Testing after review can only rank the messages review already allowed through; it cannot tell the team what to submit.

Two things follow from moving the test earlier. The choice happens before submission, so review time gets spent on the message the evidence already supports rather than on adjudicating between finalists nobody tested against the full field. And the process leaves a stated design and a set of effect estimates, which is the form of support a team can hand to its own legal and medical reviewers - not proof of compliance, but a documented basis for the choice they're being asked to approve.

What the major statement rule requires of every arm you test

FDA's rule on the major statement in broadcast advertisements requires that the statement be presented in a clear, conspicuous, and neutral manner - not buried in pacing, visuals, or competing audio that undercuts it (Federal Register vol. 88, no. 223, Nov. 21, 2023, p. 80958: govinfo.gov). The rule took effect May 20, 2024, with a compliance date of November 20, 2024. If a message tested for appeal or intent doesn't carry its risk information the same way it would need to on air, the test result describes a version of the message that can't legally run. Holding the risk information constant and present in every arm during the test, rather than adding it back in after a winner is chosen, is what keeps the test measuring something that can actually ship. Whether a specific execution meets the clear-conspicuous-neutral standard is a determination for the team's own legal and regulatory reviewers, not a claim any test output can make on its own.

How reliable is a simulated randomized comparison?

The strongest published result for this approach is a rank correlation of 0.832 against a measured human-to-human ceiling of 0.959 on the Hainmueller immigration conjoint study - one study, the best configuration reported, which works out to roughly 87% of that measured ceiling (causal fidelity paper, PDF). Across the full set of 43 published randomized studies that pass design filters, the mean rank correlation is 0.73, and that broader number is the more honest baseline to plan against, not the single-study best case. A limitation worth stating plainly: published human studies can sit inside a model's training data, which would let it pattern-match a known result rather than predict a new one; the replication protocol behind these figures screens for that risk, but the general problem of testing against material a model may have seen before doesn't disappear just because a protocol exists.

The estimators behind the comparison - McFadden discrete choice, Mixed Logit, and ICLV - are statistical models for choice data, not causal methods on their own. Causal identification comes from randomizing the message elements in the experiment design; the estimator's job is to turn the resulting choices into effect sizes. Mixed Logit is used specifically because it relaxes the independence-of-irrelevant-alternatives assumption a flat multinomial logit imposes, which matters when patient messages are close substitutes for each other. Current results across studies and methods are tracked on the leaderboard, and the underlying design questions are covered in more depth on the methods and validation hub.

Three ways to test a patient message

Monadic test on survivorsIn-market splitRandomized comparison before submission
What it measuresSelf-reported appeal, believability, intent on messages already cleared by reviewActual click or recall lift, in market, after review and after spendEstimated effect of each message element, with risk information present in every arm
When it runsAfter legal and medical review narrows the fieldAfter production and media spendBefore the message is submitted to review
What sets the option spaceThe review queueThe review queue plus the media planThe full candidate set the brand team is actually debating
What it leaves behindA score on a scaleAn outcome, after the money is spentA stated design and effect estimates the team's own reviewers can examine
Best forTeams that already trust their review-narrowed shortlist and only need a tiebreakerTeams that need a real-world read after launch and can absorb the cost of a wrong first choiceTeams choosing which message to build the campaign on, before the review cycle starts

What to do next

Take the messages currently competing for the brand team's decision and list, for each one, whether the required risk information is present in the version being discussed - not a placeholder version, the one that would actually run. Any message where that's not true isn't ready to be tested yet, regardless of how well it might score. Once the candidate set carries its risk information consistently, that's the set worth comparing, before any of them goes to review. For a closer look at how that comparison is designed and validated, meet with the team.