Payer Perception Research When You Cannot A/B Test
A payer perception study replaces the A/B test it cannot run by randomizing the dossier elements a team controls before the meeting, then measuring which ones move the coverage decision. ---
A market access lead locking a value story before the first payer meeting can still get a causal read on it, by randomizing the dossier instead of the plan. Each plan is approached once, its outcome negotiated rather than observed, and the pool of decision makers for a single product is too small for a randomized field trial at any budget. A randomized choice experiment moves the randomization onto the value story itself, estimating the effect of each dossier element on the coverage decision before that meeting happens.
- A formulary decision cannot be A/B tested because each plan is approached once and negotiated, not observed repeatedly.
- The workaround randomizes the dossier itself: endpoint emphasis, comparator, budget impact framing, rebate structure, restriction criteria.
- Choice experiments analyzed with McFadden, Mixed Logit, or ICLV produce an estimated effect with a confidence interval; a themes summary does not.
- On one published conjoint the best configuration reaches 87% of a measured human-to-human ceiling (0.832 against 0.959); across the 43 studies passing design filters the mean is 0.73 (Causal Fidelity paper).
- Advisory boards keep a role after the experiment runs, on the questions it could not settle.
Why can't a payer decision be A/B tested?
A payer decision can't be A/B tested because there is one plan and one negotiation, and the outcome forms once, with no second arm to compare it against. A market access team does not get to submit two versions of a dossier to the same P&T committee and see which one clears review faster. The committee meets once, forms a position, and that position becomes the outcome. There is no untreated control group sitting next to it having reached a different conclusion under identical conditions.
The population problem compounds this. The medical and pharmacy directors who will rule on a given product are few, against a marketing experiment that can split thousands of users into arms unnoticed, and no verified public figure exists for that population's size in the United States. A randomized trial on real committees at real formularies is not a designable study: the sample is too small, the stakes per observation too high, and the negotiation changes the moment a company instruments it.
The causal question stays answerable once the object being randomized changes from the plan to the dossier. A randomized choice experiment assigns different value stories to different simulated or surveyed respondents and measures which elements move the coverage decision, with an interval around the estimate.
What advisory boards get right, and where they stop
Advisory boards surface objections a team did not anticipate. Ten or fifteen structured conversations with payer medical directors turn up concerns a slide deck missed, and no simulated experiment replaces the judgment of someone who sits on these committees. What they cannot produce is an effect: a summary of themes from twelve conversations carries no interval, so a team cannot say whether the gap between two value stories is large, small, or noise. A handful of interviews was never built to estimate an attribute-level effect. What a payer says in that room is also a stated position inside a bargaining context, with nothing at stake for the speaker, so a themes summary has no way to correct for a say-do gap.
The decision in front of you: value story, price, contracting posture
The senior buyer here is a market access lead who has to lock a value story, a price, and a contracting posture ahead of formulary and pathway discussions, with no second attempt in the room. That decision breaks into a handful of concrete calls:
- Which clinical endpoint leads the narrative, and which becomes supporting evidence.
- Whether the dossier argues budget impact or cost offset as the primary economic frame.
- Which restriction criteria get conceded before the meeting and which get defended.
- How the rebate and contracting structure gets positioned relative to list price.
Every one of these choices gets made once, in front of the actual committee. The sequencing of concessions in particular cannot be learned by trying both orders, because the meeting happens once.
Randomize the dossier, not the plan
The plan cannot be randomized. The dossier can. A randomized choice experiment varies the clinical benefit size, the comparator, the budget impact framing, the rebate structure, and the restriction criteria across choice tasks presented to a payer-relevant sample, and estimates the effect of each element on the simulated coverage decision, with a confidence interval around each estimate. That interval covers the effect within the population sampled or simulated for the experiment; it does not, on its own, bound what a specific committee will do in a specific room.
Rebate structure and price positioning deserve a separate caveat. A choice task is a stated-preference design, not an incentive-aligned one - nobody in the experiment is spending a real budget or signing a real contract. Estimated price and rebate sensitivity from this kind of design tends to run high relative to what a payer will actually concede once real money is on the table. Treat that estimate as a directional input for sequencing concessions, not as a negotiating floor.
Which estimator: McFadden discrete choice, mixed logit, or ICLV?
Each turns randomized choice data into an effect; the causal claim comes from the randomization in the design. A McFadden discrete choice model estimates attribute-level effects and assumes independence of irrelevant alternatives, which can flatten substitution between formulary tiers. Mixed Logit relaxes that assumption by letting preferences vary across respondents, which matters when a medical director and a pharmacy director weight the same restriction criterion differently. ICLV models a latent construct between the dossier and the choice, a belief about clinical credibility or budget risk, useful when the real driver is not observable in the attribute list. Which estimator fits is a design question to settle before data collection.
How close does a simulated payer population get to real behavior?
A simulated population's accuracy means something only next to the measured human ceiling it was validated against. On the Hainmueller immigration conjoint, Subconscious's best configuration reaches 0.832 rank correlation against the published human result, where two independent samples of real humans reach 0.959 against each other, so 87% of that measured ceiling on one study. Across the 43 published studies passing design filters the mean is 0.73, the more representative number (Causal Fidelity paper).
| Comparison | Rank correlation | Basis |
|---|---|---|
| Two independent human samples | 0.959 | Measured human ceiling on the Hainmueller conjoint |
| Best simulated configuration | 0.832 | Same single study, 87% of that ceiling |
| Mean across studies passing design filters | 0.73 | 43 published studies, the planning number |
Two caveats belong next to this number. Published studies can sit inside a language model's training data, so a replication protocol has to be built to catch that; the methodology and per-study breakdown are on the leaderboard and in the methods and validation archive. This is also a validation result on studies already run, so a new therapeutic area or payer type is exactly where the mean across 43 studies is the honest number to plan against. None of it is validated for a reimbursement submission or accepted before a health technology assessment body; it supports the value story and the meeting prep.
Numbers submitted into a reimbursement process do get checked against reality. Kossmeier, Themanns, Hatapoglu and colleagues compared company-submitted sales forecasts with actual post-launch sales for 102 products applying for reimbursement in Austria (Frontiers in Pharmacology, 2021, full study). That study measures forecast accuracy rather than payer perception method accuracy, and Subconscious publishes its misses against real human choice data on the same basis.
What a defensible design documents
A defensible dossier experiment documents its attributes and levels, choice-task construction, sample composition, and estimation model well enough for a health economics group to audit unaided. The field has a consensus checklist for this: Bridges et al., "Conjoint Analysis Applications in Health - a Checklist: A Report of the ISPOR Good Research Practices for Conjoint Analysis Task Force," Value in Health, 2011 (DOI: 10.1016/j.jval.2010.11.013). A design that skips attribute-level justification, cannot state how choice tasks were generated, or reports a topline number without the underlying model is not auditable, however polished the output looks.
Where the real advisory board still earns its budget
The advisory board earns its budget on questions a randomized dossier experiment cannot settle: a committee's politics, a regulatory nuance the model never saw, or language too new to have comparable published data. Once the experiment has ranked which elements move the coverage decision, the advisory conversation becomes a focused check on what remains uncertain, which is a better use of a payer medical director's hour.
Before the next meeting, take the last dossier this team built and mark which of the five elements above were ever varied and tested against each other, and which were asserted. The ones never varied are where a randomized choice experiment replaces a guess with an estimate. To walk through what that experiment looks like for this decision, talk to the team.