Skip to content
Subconscious

Mini-lecture: Conjointly's guide to sample selection

A research director needs two checks: enough information for the intended estimate and a respondent pool appropriate to the population. A sample-size heuristic does not establish power, respondent identity, attention, or behavioral transport. Current vendor quality procedures should be evaluated alongside the design’s statistical requirements.

What does the sample-size heuristic establish?

Sawtooth’s Johnson heuristic is n ≥ 500c/(ta): respondents n, tasks t, alternatives a excluding none, and maximum attribute levels c for main effects. More tasks or alternatives generally increase level exposure at fixed n; more levels reduce exposure per level. The heuristic does not calculate a target interval or design-specific power. Suggested starting counts such as 300 overall or 200 per subgroup need adjustment for contrasts, interactions, response quality, and precision requirements.

Why does academic practice show such a wide range of "right" sample sizes?

Because there isn't one right sample size, only a right size for a given effect and a given population, and researchers disagree on both constantly. A systematic review of discrete choice experiments in healthcare found published sample sizes ranging from 10 to 3,727 respondents, with a median of just 294 (PMC). That spread is not evidence of sloppiness. It reflects that the "correct" N depends on the effect size researchers expect to detect, the design's efficiency, and how finely they plan to slice the results. What the spread does undercut is the idea that a single formula, applied mechanically, settles the sample-selection decision. Johnson's rule and the 300-respondent default are reasonable starting heuristics, not proof that a study's findings are real once the target N is hit.

Why does sample size say nothing about whether your respondents are real?

Pew’s 2023 report compares six samples fielded in 2021, totaling 29,937 U.S. adults, against 28 population benchmarks. Opt-in samples had roughly twice the average absolute benchmark error of probability samples in that comparison; this is not a universal panel multiplier. Westwood's PNAS survey-agent study reported attention-check evasion in 99.8% of 6,000 tested agent trials. That result concerns the agents and checks tested; it does not measure fraud prevalence in every panel. A panel audit should examine layered screening rather than infer human identity from attention checks alone.

Four checks before sizing a study: level exposure, contrast precision, sampling frame and screening, and matched behavioral evidence.
Statistical information and respondent quality require separate checks; the Johnson heuristic establishes neither power nor panel validity.

Sample size versus sample reality

These are two separate audits, and a buyer needs both before trusting a result.

Sample-size question (Johnson's rule, 300 default)Sample-reality question (panel and validation audit)
What it checksRough exposure of attribute levels; power needs a separate design-specific calculationIdentity, attention, sampling, and endpoint-specific human or behavioral agreement
Where the gap shows upSparse contrasts, interactions, and subgroup estimatesFraud, nonresponse, coverage gaps, and hypothetical behavior
What a bigger N fixesCan improve precision under the actual sampling/design assumptionsDoes not automatically remove systematic bias or fraud
Best forDeciding how many tasks and respondents a design can supportDeciding whether the study's causal effect is worth acting on

What does causal validation check that a sample-size formula can't?

For a simulated choice study, randomization identifies a contrast within the simulated response process under design assumptions. A matched human or behavioral test then examines transport. Mixed logit represents heterogeneous preferences; ICLV is a model specification and does not automatically relax IIA.

For sample selection, request a design-specific uncertainty or power calculation and a matched quality audit. A rank-correlation benchmark is a different measure from directional agreement, correct outcomes, or interval coverage. Evaluation holdout does not itself exclude foundation-model training exposure. Inspect the research and the actual study’s prospective or otherwise held-out comparison.

A decision checklist for a sample-selection call

Where does this leave a buyer choosing between guides?

Conjointly’s current panel-quality documentation describes screening, soft launches, automated fraud checks, and manual review. Sawtooth also discusses representativeness and poor data. These procedures matter, but none alone proves behavioral transport for the buyer’s endpoint. Ask for both the quality record and the study-specific validation. See the methods hub.

Next step: before running a new choice study, pull the panel's own quality documentation and ask two questions directly, what share of respondents failed an attention check in the last quarter, and how the panel screens for AI-generated responses. If you want a second opinion on a specific market or design, meet the team.