Skip to content

Research Methods for Hard to Reach Prescribers and Rare Disease Populations

Research Methods for Hard to Reach Prescribers and Rare Disease Populations

A commercial or medical affairs lead choosing a positioning, a service model, or a target profile for a rare disease or a narrow specialty can't wait for a discrete choice experiment sized the way the textbook requires: the frame is too small no matter what the honorarium reaches. Change what gets sampled: prove the substitute against published choice experiments already run in the same clinical area, using the replication record that shows where it holds and where it doesn't.

The arithmetic problem in rare and narrow-specialty research

A discrete choice experiment isn't sized by intuition. Its sample size follows from the number of choice tasks, attributes, and attribute levels in the design, and health economics has a stated rule of thumb (Johnson and Orme) for the minimum respondent count that design implies. When de Bekker-Grob, Donkers, Jonker, and Stolk reviewed published health discrete choice experiments against that rule, they found 32% used fewer than 100 respondents, 41% used 100 to 300, and 71% did not clearly report how they determined sample size (The Patient, 2015). That's the published literature before anyone tries to study a rare indication. In a specialty with a few hundred prescribing physicians nationally, or a disease with a diagnosed population in the low thousands, the recruitable frame sits below the design requirement before response rates are even considered.

Response rates compound the problem. Physician surveys, including web-based ones, draw response rates that vary by specialty and by contact channel, and the physicians who do respond are not a random draw from the frame (Cunningham, Quan, Hemmelgarn, and colleagues, BMC Medical Research Methodology 15:32, 2015). Recruiting further through key opinion leaders and advocacy groups adds another layer of self-selection on top of that.

Why do qualitative interviews fail as a substitute for a discrete choice experiment?

Qualitative interviews fail as a substitute because they cannot produce an effect with a confidence interval, and a positioning or service-model decision needs exactly that. Eight or twelve interviews can surface themes, language, and objections. They cannot tell a commercial lead how much one attribute moves choice relative to another, or how tight that estimate is. A simulated arm can supply that effect and interval. The interval describes the effect within the simulated population, conditioned on the clinical profile; it doesn't extend automatically to the real, unmeasured population. When a qualitative deck is the only artifact behind a launch decision worth tens of millions, the decision rests on anecdote, not an estimate.

What does the published evidence say about how DCE sample size is set?

The published evidence says most health discrete choice experiments are already undersized or unjustified relative to the field's own rule of thumb, before anyone applies that design to a rare population. The three figures below come from the same review and describe different things: how many respondents a study reached, and whether it explained how it chose that number. Read them separately, not as slices of one pie.

Bar chart showing that 32 percent of published health discrete choice experiments used fewer than 100 respondents, 41 percent used 100 to 300, and 71 percent did not clearly report how they determined sample size, per de Bekker-Grob et al. 2015.
Most published health discrete choice experiments are already undersized or unjustified before a rare population makes recruitment harder.

How does a simulated arm change what's recruitable?

A simulated arm replaces the constraint of who will answer a survey with a population built from the clinical and demographic profile the study needs, run through the same randomized choice design a fieldable study would use. The randomization does the causal work: a discrete choice model like McFadden's, a Mixed Logit, or an ICLV specification is an estimator applied after the fact, analyzing the randomized manipulation of attributes and levels inside the choice task. When the decision depends on substitution patterns as much as on which option wins - whether prescribers who reject option A shift to option B or to no treatment - a flat multinomial logit's independence-of-irrelevant-alternatives assumption can distort the answer; Mixed Logit and ICLV relax that assumption when substitution among a small set of options is part of the question.

How do you validate a simulated arm before trusting it for a launch decision?

You validate a simulated arm by replicating published randomized discrete choice experiments with it and reporting the correlation against their measured effects, including the misses. Across the 43 published randomized studies that pass design filters, the mean rank correlation against the original published result is 0.73. The best configuration reaches 87% of a measured human-to-human ceiling on one study: 0.832 rank correlation against the published result on the Hainmueller immigration conjoint, where two independent samples of real humans reach 0.959 with each other (Causal Fidelity paper). Across a wider set of roughly 300 replications, many built from studies with weaker designs or thinner reporting, the mean drops to 0.55 (Causal Fidelity paper). These are validation results on past studies. A new, untested indication has to clear that same bar on its own record before the simulated arm earns weight in the decision.

One caveat this protocol has to carry rather than hide: published studies used for replication can sit inside a model's training data, so a strong correlation on a study that was public before training proves less than the same correlation on a study published afterward. The replication set and the current numbers by therapeutic area are public on the leaderboard; check whether a study close to your indication is in the filtered set before weighting the simulated arm heavily in a decision.

What the small recruited sample does that the simulation can't

Recruited human sample (n in the single or low double digits)Simulated arm (clinically conditioned)
What it establishesFace validity: whether the attributes, levels, and wording make sense to real prescribers or patientsChoice shares and effect sizes, with confidence intervals scoped to the simulated population, across the full randomized design at the sample size the design requires
What it can't doProduce an effect size with a confidence interval at this nStand in for recruitment on its own, or substitute for evidence on the specific untested population without replication in that clinical area
Where it comes fromDirect outreach, KOL networks, advocacy groupsClinical and demographic conditioning, run through the same randomized choice design a fieldable study would use
Best for:Fixing the instrument before fielding it at any scaleCovering the part of the design the recruitable frame cannot reach, once replication in the same therapeutic area holds up

Where willingness-to-pay and preference-share questions need extra care

Willingness-to-pay questions need extra care because stated WTP from a discrete choice experiment runs high relative to what an incentive-aligned design would produce, unless the design ties choices to a real consequence. That bias runs in one direction, toward overstatement, so a positioning decision built on stated WTP from a rare-population study should treat the number as an upper bound on willingness to pay. Preference-share questions carry the IIA caveat noted above: a flat logit spreads share to a new or removed option in a fixed proportion, which rarely matches how prescribers or patients actually substitute in a narrow specialty with only a handful of real alternatives.

What to do with this before the next planning cycle

Start by checking whether a published randomized discrete choice experiment already exists in your therapeutic area on the leaderboard, since that's the study a simulated arm would need to replicate before it earns any weight in your decision. Use whatever recruitable sample you have, even eight or twelve prescribers or patients, to pressure-test the instrument's attributes and wording rather than to produce a topline number. Read the methods and validation coverage on how replication protocols are scored before deciding how much confidence a given correlation deserves. If you want to talk through how this applies to your specific indication or specialty, the team is a reasonable next stop once you've done the above.