Skip to content

How to perform smart sampling and data checking?

A research director who owns a sampling and data-checking protocol is choosing between two different guarantees: that respondents are real, and that the resulting choice model is causally right. Bot checks, attention checks, and rule-of-thumb sample sizes deliver the first guarantee. Only replication against a real human behavioral benchmark delivers the second.

What a data quality check is actually testing

Every fraud screen, speeder flag, and captcha answers one question: is this response coming from an attentive human. That's necessary, not sufficient. The NORC and CloudResearch figures above describe the scale of that one question: how many respondents are fake, and how well the standard test still catches them. Some vendors pitch layered fraud detection across recruitment, registration, in-survey behavior, and post-survey review. All of that answers one question: was the respondent real. It doesn't answer whether the study's design and estimator recover a real causal effect.

Does passing a bot check mean your data is valid?

No. A bot check confirms the respondent was human and attentive. It says nothing about whether their stated choices, run through a discrete choice model, predict what people actually do. A study can fail the second test while passing the first cleanly.

Two-column comparison showing a bot check verifies respondent authenticity while replication verifies the model reproduces real behavioral outcomes.
A bot check clears the respondent. Only replication against real behavior clears the model.

The evidence for that gap comes from fully human, fraud-free panels, not fraudulent ones. Pew Research's 2016 benchmark, still the most recent large-scale study of its kind, compared nine nonprobability samples against known population values and found average bias exceeding 10 points, including 15.1 points on Hispanic respondents and 11.3 on Black respondents (Pew Research). Every one of those respondents would have passed a bot check. AAPOR's 2013 task force report goes further: there's no generally accepted theoretical basis for treating nonprobability online panel results as projectable to the general population at all (AAPOR). Cleaning the panel doesn't fix a design that was never validated against real behavior.

Why fraud-free panels still produce biased answers

The bias Pew and AAPOR document is structural, not contamination. It comes from who opts into online panels and how they answer hypothetical questions, not from bots slipping through. More screening can't fix that: screening removes fake respondents, not the gap between what a real, attentive respondent says they'd do and what they actually do when money or effort is on the line. Vendors selling layered fraud detection solve the NORC problem. They don't solve the Pew and AAPOR problem, and a buyer evaluating a sampling protocol needs to know which one is on the table.

What do DCE sample-size rules of thumb leave out?

They tell you a coefficient will be statistically precise. They don't tell you it's correct. Practitioners commonly reach for Orme's 300-respondent minimum or Hensher, Rose, and Green's 50-per-alternative rule and treat either as a sign the data is ready. Both are heuristics for a coefficient's standard error, not tests of whether the design recovers a true causal preference. A well-powered Mixed Logit estimated on biased or unvalidated stated-preference data still produces a precise wrong answer: a tight confidence interval around a number that doesn't reflect real behavior.

ApproachWhat it tells youWhat it misses
Rule of thumb (n=300, or 50 per alternative)Enough respondents for a computable, roughly stable coefficientWhether that coefficient reflects a real causal preference or a biased panel's guess. Best for: a quick sanity check before scoping a study, never the sole validity claim.
Formal power calculationThe sample size needed to detect a specified effect size at a chosen confidence levelPanel representativeness or hypothetical bias in stated choices. Best for: teams that already trust their panel and need to size precision for a known effect.
Replication against a human baselineWhether the model's direction and outcome match a real, incentive-relevant studyA guarantee that results transfer to a market that's never been benchmarked. Best for: any decision where the cost of a precise wrong answer exceeds the cost of a validation check.

What actually proves a choice model is causally right?

Replication: rerun the design against a real human study's outcome and check whether the model reproduces the direction and result, not just a plausible-looking coefficient. On a validation set of past studies, this reaches 87% of the measured human ceiling (0.832 against a 0.959 human-to-human ceiling; mean 0.73 across the 43 studies passing design filters), a validation-set result, not a guarantee for a brand-new market (the causal fidelity paper). That caveat matters twice over: some published studies used for validation could overlap with a model's training data, and the public leaderboard exists to control for exactly that risk, ranking approaches against held-out human studies selected to avoid training-data overlap rather than pretending the risk isn't there. That's the place to compare vendors on this claim, not on fraud-detection marketing.

McFadden discrete choice models, Mixed Logit, and ICLV are estimators, not causal methods on their own. Causal identification comes from the randomized manipulation built into the experiment design; the estimator fits a model to the choices that randomization produced. A confidence interval from a simulated experiment covers the effect within the simulated population the experiment ran on. It doesn't bound the real market unconditionally; that link has to be established through replication against the population the study is meant to represent.

How do you build a sampling and checking protocol that catches both problems?

Run fraud screening and causal validation as two separate steps. They catch two different failure modes.

  1. Screen for bots and inattentive respondents with layered detection, not a single attention check. Standard attention checks no longer catch AI-generated responses reliably (CloudResearch).
  2. Size the study with a formal power calculation tied to the effect size that matters for the decision, not a rule of thumb like n=300.
  3. Before trusting the output, check whether the design and estimator have been replicated against a real, held-out human study, and ask for the replication rate, not just the fraud rate.
  4. Treat any willingness-to-pay estimate as directionally inflated unless the design is incentive-aligned; stated WTP runs high when nothing real is at stake.
  5. For preference-share or substitution questions modeled with a flat logit, check the independence of irrelevant alternatives (IIA) assumption before trusting the result. A flat logit pulls share from every competitor proportionally, an assumption that breaks when two options are close substitutes. Mixed Logit and ICLV relax this by letting preferences vary across the simulated population.

More on how the estimators and validation protocol fit together: methods and validation.

Where willingness-to-pay estimates still mislead

A WTP figure from a stated-preference discrete choice study runs high, in a known direction, unless the design pays out real money or otherwise makes the choice consequential. That's a separate failure from panel fraud or panel bias: even a perfectly sampled, perfectly human, direction-validated panel will overstate what people would actually pay when the question is hypothetical. A protocol that stops at fraud detection and rule-of-thumb sample size will never catch this, because WTP inflation shows up in the size of the effect, not in whether the respondent was real or the model was precise.

Pull the last study your team shipped and check one thing: does the write-up compare the model's output to a real human behavioral benchmark, or does it stop at the fraud rate and the sample size. If it stops there, that's the gap to close before the next wave. Want the replication protocol behind that 87% figure walked through for your market? Book time with the team.