Skip to content

Mini-lecture: Conjointly's guide to sample selection

A research director deciding whether to follow Conjointly's sample-selection guidance is really deciding one thing: how many respondents a study needs, and whether that respondent pool will tell the truth. Conjointly's guide, like the rest of the category, answers only the first half. It sizes a sample against a target confidence interval; it says nothing about whether the people filling out that survey exist, pay attention, or would act the same way once money or a real choice is on the line. A study can hit its target N exactly and still produce a causal effect that isn't real.

What Conjointly's sample-size guidance recommends

Conjointly's guidance, like most in the category, treats sample selection as a statistical power problem. The reference point most of the field converges on is Sawtooth's Johnson rule of thumb for choice-based conjoint studies: n·t·a/c ≥ 500, where n is respondents, t is tasks per respondent, a is alternatives per task, and c is the largest number of levels in any attribute. In practice this collapses to a familiar default: roughly 300 respondents, with about 200 per subgroup if the plan calls for reporting results by segment (Sawtooth Software). This is useful math. It tells a buyer how much precision a given N buys, and how that precision degrades as tasks, alternatives, or attribute levels increase. What it does not do is say anything about who fills out those 300 or 200 surveys.

Why does academic practice show such a wide range of "right" sample sizes?

Because there isn't one right sample size, only a right size for a given effect and a given population, and researchers disagree on both constantly. A systematic review of discrete choice experiments in healthcare found published sample sizes ranging from 10 to 3,727 respondents, with a median of just 294 (PMC). That spread is not evidence of sloppiness. It reflects that the "correct" N depends on the effect size researchers expect to detect, the design's efficiency, and how finely they plan to slice the results. What the spread does undercut is the idea that a single formula, applied mechanically, settles the sample-selection decision. Johnson's rule and the 300-respondent default are reasonable starting heuristics, not proof that a study's findings are real once the target N is hit.

Why does sample size say nothing about whether your respondents are real?

Because N measures how many observations you have, not what those observations are worth, and the value of each observation has been dropping. Pew Research compared six online panels across 29,937 U.S. adults and 28 benchmark variables and found that opt-in samples carry roughly twice the average error of probability-based panels (Pew Research Center). That gap exists before anyone touches a sample-size formula. Layered on top of it: CloudResearch's 2025 analysis found that AI agents now pass thousands of standard attention checks with near-perfect accuracy, defeating the exact screens researchers rely on to certify a respondent pool as clean (CloudResearch). A sample-size formula has no term for either of these problems. It assumes the respondents behind N are what they claim to be.

Bar chart showing opt-in panel error at roughly double the error of probability-based panels.
Opt-in panels run about twice the average error of probability-based panels across Pew's 28 benchmark variables, a gap sample-size math cannot see or correct.

Sample size versus sample reality

These are two separate audits, and a buyer needs both before trusting a result.

Sample-size question (Johnson's rule, 300 default)Sample-reality question (panel and validation audit)
What it checksWhether N is large enough to hit a target confidence interval given tasks, alternatives, and attribute levelsWhether respondents are real, attentive, and representative, and whether the resulting effect matches real human behavior
Where the gap shows upUnderpowered subgroup cuts, unstable estimates on rare attribute levelsAI agents passing attention checks, opt-in panels running roughly double the error of probability panels, stated preference diverging from actual behavior
What a bigger N fixesStandard error, confidence interval widthNothing; a larger sample of unreal respondents just makes a false result more precise
Best forDeciding how many tasks and respondents a design can supportDeciding whether the study's causal effect is worth acting on

What does causal validation check that a sample-size formula can't?

It checks whether the effect a study finds actually reproduces what real people do, not just whether the sample was big enough to estimate it precisely. Subconscious runs randomized experiments on a simulated population and analyzes them with discrete choice models, McFadden discrete choice, Mixed Logit, and ICLV among them. These are estimators, not causal methods in themselves; the causal identification comes from the randomized manipulation built into the experiment design, not from the choice of model. Mixed Logit and ICLV also relax the independence-of-irrelevant-alternatives assumption baked into a flat logit, which matters directly if the question is about substitution between options rather than a single average effect.

The validation number that matters here is replication accuracy: how often a simulated study reproduces the direction and outcome of the original human study, measured against a set of held-out human studies. Our best configuration reaches 87% of the measured human ceiling on one study: 0.832 rank correlation against the published human result, where two independent samples of real humans reach 0.959. Across all 43 studies that pass design filters the mean is 0.73 (the causal fidelity paper). Two caveats belong next to that number, not after it. First, it is a validation-set result, not a guarantee for a new, unstudied market; every new study still needs its own check. Second, published human studies can sit in a model's training data, which would inflate a replication score if left unaddressed; the protocol behind the 0.832 and 0.73 figures holds out studies and populations to control for this, but the concern does not fully disappear just because a protocol exists. A confidence interval computed from a simulated experiment covers the effect within that simulated population; it does not, on its own, bound what the real market will do. The leaderboard publishes results across methods and studies so a buyer can see how replication accuracy holds up outside a single case.

A decision checklist for a sample-selection call

Where does this leave a buyer choosing between guides?

It leaves the sample-size guide as necessary but incomplete, and the reality check as the part most guides skip. Conjointly's mini-lecture, Sawtooth's rule of thumb, and the healthcare DCE literature all answer "how many." None of them answer "are these respondents, and their stated preferences, real." A study can be correctly sized and still be measuring something that doesn't exist outside the survey tool. The methods and validation hub covers how to check the second half of that question once the first half is settled.

Next step: before running a new choice study, pull the panel's own quality documentation and ask two questions directly, what share of respondents failed an attention check in the last quarter, and how the panel screens for AI-generated responses. If you want a second opinion on a specific market or design, meet the team.