Mini-lecture: Conjointly's guide to sample selection
A research director deciding whether to follow Conjointly's sample-selection guidance is really deciding one thing: how many respondents a study needs, and whether that respondent pool will tell the truth. Conjointly's guide, like the rest of the category, answers only the first half. It sizes a sample against a target confidence interval; it says nothing about whether the people filling out that survey exist, pay attention, or would act the same way once money or a real choice is on the line. A study can hit its target N exactly and still produce a causal effect that isn't real.
- Conjointly's sample-size guidance sits in the same lineage as Sawtooth's Johnson rule of thumb (n·t·a/c ≥ 500), with 300 respondents as the common default and roughly 200 per subgroup when segments are reported separately.
- Academic discrete choice experiments show no real consensus on N: a systematic review of healthcare DCEs found sample sizes from 10 to 3,727, median 294.
- The bigger risk now sits upstream of sample size: opt-in online panels carry roughly double the error of probability-based panels, and AI agents defeat the attention checks panels use to certify a respondent as human.
- A precisely sized sample of unreal respondents still produces a precise, unreal effect. Sample-size math and respondent-reality checks are two different questions, and most guides only answer the first.
- The check that closes the gap is replication against a real human study, not a larger N.
What Conjointly's sample-size guidance recommends
Conjointly's guidance, like most in the category, treats sample selection as a statistical power problem. The reference point most of the field converges on is Sawtooth's Johnson rule of thumb for choice-based conjoint studies: n·t·a/c ≥ 500, where n is respondents, t is tasks per respondent, a is alternatives per task, and c is the largest number of levels in any attribute. In practice this collapses to a familiar default: roughly 300 respondents, with about 200 per subgroup if the plan calls for reporting results by segment (Sawtooth Software). This is useful math. It tells a buyer how much precision a given N buys, and how that precision degrades as tasks, alternatives, or attribute levels increase. What it does not do is say anything about who fills out those 300 or 200 surveys.
Why does academic practice show such a wide range of "right" sample sizes?
Because there isn't one right sample size, only a right size for a given effect and a given population, and researchers disagree on both constantly. A systematic review of discrete choice experiments in healthcare found published sample sizes ranging from 10 to 3,727 respondents, with a median of just 294 (PMC). That spread is not evidence of sloppiness. It reflects that the "correct" N depends on the effect size researchers expect to detect, the design's efficiency, and how finely they plan to slice the results. What the spread does undercut is the idea that a single formula, applied mechanically, settles the sample-selection decision. Johnson's rule and the 300-respondent default are reasonable starting heuristics, not proof that a study's findings are real once the target N is hit.
Why does sample size say nothing about whether your respondents are real?
Because N measures how many observations you have, not what those observations are worth, and the value of each observation has been dropping. Pew Research compared six online panels across 29,937 U.S. adults and 28 benchmark variables and found that opt-in samples carry roughly twice the average error of probability-based panels (Pew Research Center). That gap exists before anyone touches a sample-size formula. Layered on top of it: CloudResearch's 2025 analysis found that AI agents now pass thousands of standard attention checks with near-perfect accuracy, defeating the exact screens researchers rely on to certify a respondent pool as clean (CloudResearch). A sample-size formula has no term for either of these problems. It assumes the respondents behind N are what they claim to be.
Sample size versus sample reality
These are two separate audits, and a buyer needs both before trusting a result.
| Sample-size question (Johnson's rule, 300 default) | Sample-reality question (panel and validation audit) | |
|---|---|---|
| What it checks | Whether N is large enough to hit a target confidence interval given tasks, alternatives, and attribute levels | Whether respondents are real, attentive, and representative, and whether the resulting effect matches real human behavior |
| Where the gap shows up | Underpowered subgroup cuts, unstable estimates on rare attribute levels | AI agents passing attention checks, opt-in panels running roughly double the error of probability panels, stated preference diverging from actual behavior |
| What a bigger N fixes | Standard error, confidence interval width | Nothing; a larger sample of unreal respondents just makes a false result more precise |
| Best for | Deciding how many tasks and respondents a design can support | Deciding whether the study's causal effect is worth acting on |
What does causal validation check that a sample-size formula can't?
It checks whether the effect a study finds actually reproduces what real people do, not just whether the sample was big enough to estimate it precisely. Subconscious runs randomized experiments on a simulated population and analyzes them with discrete choice models, McFadden discrete choice, Mixed Logit, and ICLV among them. These are estimators, not causal methods in themselves; the causal identification comes from the randomized manipulation built into the experiment design, not from the choice of model. Mixed Logit and ICLV also relax the independence-of-irrelevant-alternatives assumption baked into a flat logit, which matters directly if the question is about substitution between options rather than a single average effect.
The validation number that matters here is replication accuracy: how often a simulated study reproduces the direction and outcome of the original human study, measured against a set of held-out human studies. Our best configuration reaches 87% of the measured human ceiling on one study: 0.832 rank correlation against the published human result, where two independent samples of real humans reach 0.959. Across all 43 studies that pass design filters the mean is 0.73 (the causal fidelity paper). Two caveats belong next to that number, not after it. First, it is a validation-set result, not a guarantee for a new, unstudied market; every new study still needs its own check. Second, published human studies can sit in a model's training data, which would inflate a replication score if left unaddressed; the protocol behind the 0.832 and 0.73 figures holds out studies and populations to control for this, but the concern does not fully disappear just because a protocol exists. A confidence interval computed from a simulated experiment covers the effect within that simulated population; it does not, on its own, bound what the real market will do. The leaderboard publishes results across methods and studies so a buyer can see how replication accuracy holds up outside a single case.
A decision checklist for a sample-selection call
- Start with Johnson's rule or Conjointly's 300-respondent default to size the study; it's a reasonable floor, not a finish line.
- Ask which panel the respondents come from and whether it's probability-based or opt-in; opt-in carries roughly double the error, before any other issue shows up.
- Ask what stops an AI agent from completing the survey as a human; standard attention checks no longer do this reliably.
- Ask whether the resulting preference estimates have been checked against real human behavior, not just internal statistical consistency.
- If the study reports willingness-to-pay, treat stated figures as directionally high unless the design is incentive-aligned; hypothetical bias runs in a known direction.
Where does this leave a buyer choosing between guides?
It leaves the sample-size guide as necessary but incomplete, and the reality check as the part most guides skip. Conjointly's mini-lecture, Sawtooth's rule of thumb, and the healthcare DCE literature all answer "how many." None of them answer "are these respondents, and their stated preferences, real." A study can be correctly sized and still be measuring something that doesn't exist outside the survey tool. The methods and validation hub covers how to check the second half of that question once the first half is settled.
Next step: before running a new choice study, pull the panel's own quality documentation and ask two questions directly, what share of respondents failed an attention check in the last quarter, and how the panel screens for AI-generated responses. If you want a second opinion on a specific market or design, meet the team.