Skip to content

How to Get Participants For Your Study

---

A research director who needs three hundred completes by next Friday has a sourcing decision to make, but it is not "which panel has the lowest bot rate." The decision is whether the study design itself, panel aside, produces an answer about why people choose rather than just what they claim they'd choose. A perfectly screened sample still only reports stated preference, and stated preference has never been the same thing as the behavior that drives a purchase, a signup, or a churn.

Panel hygiene fixes fraud, not the stated-preference gap

Panel providers have spent the last two years competing on fraud metrics because they had to. Click-farm and bot infiltration pushed usable response rates in some pipelines from roughly 75% down to about 10%, which is a real operational crisis and a real reason to screen harder (CloudResearch). That figure comes from a vendor blog with a fraud-detection product to sell, not a peer-reviewed source, so treat it as a directional signal of a real trend rather than a precise industry rate. But fraud screening answers a narrower question than buyers assume. It tells you the person on the other end of the survey is a human who read the question. It does not tell you that the answer they gave predicts what they will actually do. Those are separate problems, and a checklist that stops at "screen for attention checks, pick a reputable panel, set fair incentives" treats the second problem as solved by solving the first.

Which panel should a buyer choose for a clean sample?

Pick the panel whose attention-check pass rate and cost per quality respondent fit the budget, then treat the resulting data as stated preference, not causal proof. A cleaner panel produces more attentive answers, not more predictive ones. In the largest recent comparison of major panels, only 59% of MTurk respondents passed both attention checks, versus 87% on Prolific and 85% on CloudResearch, and cost per quality respondent ran $1.90 on Prolific against $4.36 on MTurk and $8.17 on Qualtrics (PLOS One, Douglas, Ewell, Brauer, 2023). That is a meaningful difference in data hygiene. It says nothing about whether the underlying survey question was structured to isolate a causal driver of choice versus a stated opinion shaped by social desirability or professional-panelist fatigue. A panel with a 95% attention-check pass rate can still return a stated-preference number that has no relationship to what happens when the product actually ships.

ProviderAttention-check pass rateCost per quality respondentBest for
Prolific87%$1.90Best for buyers who need the cleanest stated-preference sample at the lowest cost per quality respondent.
CloudResearch85%Not reported in this studyBest for buyers already using the platform's panel management tools who still need strong screening; confirm cost per quality respondent independently, since this study didn't measure it.
MTurk59%$4.36Best for high-volume, low-cost pilot runs where the design, not sample cleanliness, carries the validity.
Qualtrics PanelsNot directly compared on attention checks in this study$8.17Best for buyers who need enterprise panel management and compliance and can absorb the higher cost; pair with independent attention-check screening, since this study didn't measure Qualtrics on that metric.

(All figures from PLOS One, 2023.) This table answers a procurement question. It does not answer the validity question, which sits one level up.

Can synthetic respondents replace a human panel entirely?

No, not as a like-for-like substitute, and the evidence on this is specific. A peer-reviewed comparison found that 48% of regression coefficients estimated from ChatGPT-simulated survey responses differed significantly from real ANES human data, and the sign of the effect flipped in 32% of those mismatched cases (Political Analysis, Cambridge). That comparison covers one dataset (ANES), one model, and one ChatGPT vintage, not a general claim about every LLM or every survey domain, but the direction of the finding is the part that matters for a sourcing decision: coefficients moved, and some flipped sign. A flipped sign is not noise. It means a synthetic-respondent study can point a pricing or messaging decision in the wrong direction while still returning a clean, complete, well-formatted dataset. The underlying caution holds regardless of how many practitioners voice it: an LLM trained on internet text is not a randomized experiment, and asking it to role-play a respondent does not create one.

Where fraud filtering and causal design solve different problems

The reason buyers conflate these two fixes is that both happen inside the same "collect responses" step of a study, which makes it easy to assume that fixing one fixes the other.

Two parallel tracks. The top track shows panel hygiene steps ending in a clean stated-preference dataset. The bottom track shows randomized design steps ending in a causal effect with a confidence interval.
A clean panel and a causal design fix two separate failure points, and a study can pass the first while still failing the second.

A study can pass every fraud check on the top track and still fail the bottom one, because nothing in attention-check screening tests whether the survey isolates what actually drives the choice.

What actually closes the say-do gap in a stated-preference study?

A randomized manipulation, not a screening pass, closes the say-do gap, because it is the randomization that lets you attribute a change in choice to a specific intervention rather than to whoever happened to answer. Methods like McFadden discrete choice, Mixed Logit, and ICLV are estimators that fit a model to choice data; they are not themselves the source of causal identification. The precise description is a randomized experiment analyzed with a discrete choice model, not a "causal method like DCE." A standard multinomial logit also carries the IIA assumption, meaning it assumes adding or removing an option doesn't change the relative odds between the others, which is one reason Mixed Logit is preferred when substitution patterns matter. None of this depends on which panel supplied the respondents. It depends on whether the study randomized something and measured the resulting choice.

How does a buyer validate a sourcing decision against real behavior?

A buyer checks the design's output against a held-out human baseline before trusting it. The same discipline that applies to panel selection applies to any simulated or synthetic-respondent design. On causal fidelity testing, the best configuration reaches 87% of the measured human ceiling on one study: a 0.832 rank correlation against the published human result, where two independent samples of real humans reach 0.959 correlation with each other. Across all 43 studies that passed design filters, the mean is 0.73, well below the single-study best case (causal fidelity paper). That number is a validation result on studies already run, not a guarantee for a market you haven't tested yet, and it carries the same caveat that applies to any comparison against a published human study: the published result may have been in a model's training data, which is exactly why a replication protocol and a held-out baseline matter more than a single correlation figure. A confidence interval built from a simulated experiment covers the effect within that simulated population; it is not a claim about the real market until it has been checked against holdout human data. Current comparisons across methods and providers are tracked on the leaderboard, and the underlying validation approach is described on the methods and validation hub.

The one decision procurement checklists skip

The checklist answers "did I get a clean sample." The decision that actually determines whether a study is worth running is "does this design produce a causal, replicable answer, regardless of which panel supplied the respondents." A senior buyer facing a participant-sourcing decision this week should pick the panel that fits the budget and screening bar, using the PLOS One figures above as a starting point, then check whether the design itself randomizes the variable under test, the one thing panel hygiene cannot fix. More panel comparisons are collected on /blog/comparisons, and current method-by-method fidelity results are on the leaderboard. When the design needs a held-out human baseline before it ships, /meet is the next step.