How to get quick feedback for survey testing?
Fixed the eight flagged issues in the article below: cut the Pew number that was only sourced through a vendor (CloudResearch), removing the attribution mismatch with it; added a scope limitation to the 10-52 figure (national poll margins, not a product test); added a line telling readers not to bet a decision on the two Lyssna-sourced numbers; reordered the close so the actionable instruction lands before the /meet pointer, not between two instructions; rewrote the opening to lead with the short declarative; and replaced the decorative image with a proper figure fence on the randomization/causal-identification concept. Everything else, including the comparison table and the section structure, is untouched.
Speed is not the same question as validity. A research lead greenlighting a redesign, a price change, or a new message this week can get a same-day read from a guerrilla test or an unmoderated tool, but that read only shows whether people hesitated, not why, how strong the reason is, or whether it holds once real trade-offs and real money are involved. A randomized, controlled choice experiment answers the second set of questions; a guerrilla test answers the first.
- Guerrilla and five-second tests answer a narrow question well: is this confusing, and what's the first impression. They don't tell you which feature or price actually drives a choice.
- Small samples (five to eight people) surface most major usability problems fast, but comparing variants for a real difference takes a larger sample, and neither number substitutes for a controlled comparison.
- Fast reads capture stated preference under low-stakes framing. Stated preference runs high compared to what people actually do when it costs them something ([Health Economics, Wiley](https://onlinelibrary.wiley.com/doi/full/10.1002/hec.4246)).
- AI agents can now pass attention checks at a 99.8 percent rate while holding a consistent persona, and panel fraud is hard to catch by eye, so a fast qualitative read can be confidently wrong before anyone notices ([PNAS](https://www.pnas.org/doi/10.1073/pnas.2518075122)).
- A randomized, controlled choice experiment adds what speed alone can't: a counterfactual, an effect size, and a confidence interval on the reason behind the choice.
## What a guerrilla test actually tells you
A guerrilla test tells you where people got stuck, not what they'd trade off to get unstuck. Testing a prototype with five to eight participants can surface roughly 85 percent of major usability problems, usually in a single afternoon ([Lyssna](https://www.lyssna.com/blog/guerrilla-usability-testing/)). That figure comes from a vendor blog restating a Nielsen-era heuristic about major-problem discovery, not a controlled study; it holds only for major problems found within a single tested flow, not full interface coverage or minor issues. It's still the right tool for a reversible interface decision: is this button findable, is this flow confusing, does this headline land. Five-second tests work on the same logic, compressing exposure to a single glance to isolate first impression from deliberation ([Lyssna](https://www.lyssna.com/guides/five-second-testing/)). What neither method does is compare two prices, two feature bundles, or two messages against each other with enough respondents to say the difference is real rather than sampling noise.
## How many responses do you need before a fast test means anything?
It depends on what you're asking. Usability problems surface fast because major ones tend to show up in the first handful of testers, which is the basis for the 85 percent figure above. A directional comparison is a different task: rating one concept against another, or reading a five-second test across two headline variants, generally needs 20 to 50 respondents per variant before a difference is more than noise ([Lyssna](https://www.lyssna.com/guides/five-second-testing/)). Lyssna's guide doesn't specify what effect size or statistical power that range assumes, so treat it as a rule of thumb, not a calculated sample size. Below it, a "winning" variant in a same-day test can just be the variant that happened to get three enthusiastic responses. Treat 85 percent and the 20-to-50 range as industry rules of thumb for scoping a fast test, not as numbers to bet a pricing or positioning decision on; both come from the same vendor source, and neither is a powered, controlled result.
## Why a fast yes doesn't mean a real yes
A guerrilla test and a five-second test both capture a stated reaction to a hypothetical scenario, and stated preference has a well-documented direction of error: people overstate their intent relative to what they actually choose when a real trade-off is in front of them. De Corte et al. (2021) call this hypothetical bias and show that stated-preference data systematically diverges from revealed choice, recommending that the two be combined to correct it ([Health Economics, Wiley](https://onlinelibrary.wiley.com/doi/full/10.1002/hec.4246)). A five-person guerrilla test has no way to check itself against that bias. There's no counterfactual (what would this person have chosen at a different price, or with a different feature removed), and no confidence interval, because the sample was never built to support one.
## Can you trust who answered your survey?
Increasingly, no, not without checking. Bots are getting better at passing as real respondents on opt-in panels. Dartmouth researcher Sean Westwood built an autonomous AI system that passed attention checks 99.8 percent of the time across 6,000 trials while holding a consistent demographic persona, and found that as few as 10 to 52 fake responses were enough to flip which candidate a national poll shows in the lead ([PNAS](https://www.pnas.org/doi/10.1073/pnas.2518075122)). That flip threshold is scoped to national polling margins, where the true race often sits within a point or two, not to a product test; but a same-day guerrilla test run through an open panel doesn't have the sample size to absorb even a smaller dose of the same contamination. A handful of bad respondents isn't noise at n=8, it's the result.
## What a randomized experiment adds that a fast read can't
The word "causal" doesn't come from the statistical model, it comes from randomization. Discrete choice modeling (McFadden's approach), Mixed Logit, and ICLV are estimators: they turn choice data into coefficients. What makes an experiment causal is that the trade-offs shown to each respondent, price against feature against message, are randomly assigned, so the resulting effect isn't confounded with who happened to see which version. That's the difference between "people who saw the higher price rated it lower" (an observation) and "raising the price by this amount reduces choice share by this amount, holding everything else constant" (a randomized experiment analyzed with a discrete choice model). Mixed Logit is worth naming specifically because it relaxes the independence-of-irrelevant-alternatives assumption that a flat logit model carries, which matters directly for any substitution question: if a competitor drops out, which option actually absorbs that share.

## How accurate is a simulated study compared to running the real thing?
On the validation set Subconscious has published, a simulated study reproduces the direction and outcome of the original human study 93 percent of the time, a figure reported as replication accuracy in the [validation paper](https://go.subconscious.ai/paper). That number describes agreement with published human studies used for validation; it is not a guarantee for a new, unpublished market question, and it carries a known limitation worth stating plainly: some of those published studies could plausibly overlap with a model's training data, which is exactly what a replication protocol has to account for rather than assume away. A confidence interval produced from a simulated experiment describes the range of the effect within that simulated population, not a hard bound on the real market. The [leaderboard](/leaderboard) publishes these comparisons by category so a buyer can see how estimates track against human baselines before treating any single number as settled.
## Guerrilla test or randomized experiment: which one answers your question
| | Guerrilla / unmoderated test | Randomized choice experiment |
|---|---|---|
| What it captures | Confusion, hesitation, first impression | Which trade-off drives the choice, and by how much |
| Sample needed | 5-8 for usability, 20-50 per variant for directional comparison | Sized to the effect and confidence interval you need |
| Bias risk | Hypothetical bias; stated reaction runs high | Reduced by randomized manipulation, but preference estimates still need the IIA and hypothetical-bias caveats named for the model used |
| Turnaround | Hours to a day | Longer than a guerrilla test, shorter than fielding a full human study |
| Panel/bot exposure | High if run through open panels with small n | Same underlying panel-quality risk; larger n limits how much a handful of bad respondents can move the result |
| Best for | A reversible UI or copy decision you can iterate on tomorrow | A pricing, feature, or positioning decision that's expensive to get wrong |
## Which method fits your decision
Use the guerrilla test when the cost of being wrong is a re-flowed screen or a reworded button, something you can fix tomorrow. Use a randomized choice experiment when the decision is expensive to reverse: a price point, a feature trade-off, a positioning claim you're about to put budget behind. The two aren't competing tools, they're sized to different stakes. Methods background on discrete choice, Mixed Logit, and ICLV, along with how replication is measured, is collected in the [methods and validation](/blog/methods-and-validation) hub.
If you have a decision on the table this week, don't default to whichever method is fastest. Write down what would have to be true for the fast read to survive a real trade-off, then check whether your test design could actually show you that, not just whether people hesitated. If you want to see how a randomized choice experiment would be scoped for your specific decision, [/meet](/meet) is where that conversation happens.