Skip to content

Sample Size Calculator and Guide to Survey Sample Size

A research director sizing a discrete choice or conjoint study is really deciding how much to trust the number that comes out the other end, and that decision starts before the study is even fielded. A standard sample size calculator, the kind built into SurveyMonkey or Qualtrics, answers only one question: how tight will the confidence interval be around a single population proportion. It says nothing about whether a choice-based study's preference estimates reflect what people actually do. Getting the sample size right and getting the decision right are two separate problems, and a buyer who conflates them ships a precisely measured, confidently wrong result.

What a sample size calculator actually tells you

A sample size calculator tells you how confidently wrong you can afford to be, not whether the thing you measured is true. The Cochran formula behind SurveyMonkey's tool and surveysystem.com's long-standing calculator takes three inputs, confidence level, margin of error, and population size, and returns the number of respondents needed to estimate one population proportion within that margin. Plug in 95% confidence, a ±5% margin, and an unbounded population, and both tools return roughly the same answer: n≈384. That's the correct math for "what share of our customers prefer X," a single descriptive number. It was never built to size a study that estimates several attribute-level utilities at once, which is what a conjoint or discrete choice experiment does.

Why conjoint and discrete choice studies need a different formula

A choice-based study needs more respondents, tasks, or alternatives than a descriptive survey because it's estimating a set of relative utilities, not one proportion. Sawtooth Software, the field standard for CBC design, publishes Johnson and Orme's heuristic: n·t·a/c ≥ 500, where n is respondents, t is tasks per respondent, a is alternatives per task, and c is the largest number of levels in any attribute. Sawtooth also recommends a practical floor of 300 respondents per study and 200 per subgroup you intend to analyze separately. A more formal, power-based alternative comes from the PMC guide to DCE sample sizing in healthcare, which derives sample requirements from statistical power rather than a rule of thumb, and academic work continues to refine Orme's heuristic because it can underestimate the sample a given design actually needs.

How many respondents do you need for a discrete choice experiment?

Start with Johnson and Orme's n·t·a/c ≥ 500 and treat 300 respondents as a floor, not a target. If your design has 4 tasks, 3 alternatives per task, and a maximum of 5 levels on any one attribute, the formula wants n ≥ 500 × 5 / (4 × 3), or roughly 209 respondents, but Sawtooth's practical floor still puts you at 300 for the full study and 200 for any subgroup you plan to cut the data by. For anything higher stakes than a rough directional read, especially where you need defensible statistical power rather than a rule of thumb, the PMC guide's power-based approach is the more rigorous path. Either way, this number answers "will my part-worth estimates be stable," not "are those part-worths real."

Three ways to size a study, compared

ApproachWhat it estimatesBasisWhat it can't tell you
Generic MOE calculator (SurveyMonkey, surveysystem.com)One population proportion, within a stated margin of errorCochran margin-of-error formula: confidence level, margin, population sizeAnything about multi-attribute choice, or whether the proportion reflects real behavior
Sawtooth rule of thumb / PMC power-based guidanceStability of part-worth utilities across attributes and levelsJohnson and Orme's n·t·a/c ≥ 500, or formal power calculationsWhether the preferences the model recovers are causal or just internally consistent
Causal design validated against a human baselineWhether the design reproduces the direction and outcome of real human choicesRandomized experiments analyzed with discrete choice models (McFadden, Mixed Logit, ICLV), checked against a holdout of human resultsGuaranteed accuracy on a market that hasn't been studied before; validation-set performance is not a guarantee

Best for: pick the generic calculator when you're estimating one proportion off a descriptive survey. Pick Sawtooth's rule of thumb or the PMC power method when you're sizing a conjoint or DCE and need the study to hold up statistically. Pick a causal, baseline-validated design when the decision riding on the result is expensive enough that "statistically stable" isn't the same bar as "actually true."

What a bigger sample size can't fix

A larger n narrows the confidence interval around your estimate; it does not move that estimate closer to what people actually do. This matters most for willingness-to-pay questions, where stated preference research consistently runs high. Hypothetical bias research reports stated WTP running roughly 2-3x higher than revealed WTP on average, with documented bias magnitudes ranging from 25% to 300% depending on the study. Doubling your sample size in that situation buys you a tighter interval around a number that's still wrong in the same direction, by the same margin. Any willingness-to-pay estimate your team relies on should carry that caveat explicitly, because the bias runs one way: stated numbers overstate real willingness to pay, and a bigger sample makes the overstatement more precise, not less real.

Does a tighter confidence interval mean the preference is real?

No. A confidence interval tells you how much sampling noise surrounds your estimate; it says nothing about whether the underlying preference would hold up if people actually had to choose. This is where the sample-size question and the causal question split. A confidence interval from a simulated or fielded experiment covers the estimated effect within that specific study population, not the real market unconditionally, and a flat logit model carries the independence-of-irrelevant-alternatives (IIA) assumption, which can distort preference-share and substitution estimates when alternatives aren't truly independent, a limitation Mixed Logit is specifically designed to relax. None of that shows up in a sample size calculator's output. It shows up when you check the design against what real people actually did.

A branching diagram: a single-proportion survey uses the Cochran margin-of-error formula; a choice-based or DCE study uses Johnson and Orme's rule of thumb or power-based DCE guidance; both paths still require checking the design against real human behavior before trusting the result.
Sizing a sample and validating a preference are two different checks, and only one of them tells you if the number is real.

How causal validation checks a design against real behavior

Validation means running the design as a randomized experiment, analyzing it with discrete choice models, and checking the result against a holdout of real human data, not against a wider or narrower confidence interval. Subconscious's replication protocol shows our best configuration reaches 87% of the measured human ceiling on one study: 0.832 rank correlation against the published human result, where two independent samples of real humans reach 0.959. Across all 43 studies that pass design filters the mean is 0.73, detailed in the causal fidelity paper. That figure is a validation-set result, not a guarantee for a market that hasn't been studied yet, and it comes with a caveat worth stating plainly: some published human studies used for validation could sit inside a model's training data, which is exactly the kind of contamination risk the replication protocol is built to catch, not a problem it pretends doesn't exist. Method choice matters too. McFadden's discrete choice model, Mixed Logit, and ICLV are estimators, not causal methods on their own; causal identification comes from the randomized manipulation in the experiment design itself, which is why the accurate description is randomized experiments analyzed with discrete choice models. Study-by-study replication performance, broken out by method and category, is published on the public leaderboard. More on how these methods compare in practice is in the methods and validation hub.

Before you field your next choice-based study, run it through Johnson and Orme's rule of thumb yourself, confirm you're above Sawtooth's 300-respondent floor, and then ask the harder question a calculator can't answer: has this design, or one like it, been checked against real human behavior. The leaderboard is a reasonable place to look for that second answer. If you want a second read on your specific design, meet the team.