Skip to content

Conversations, Not Checkboxes: A Smarter Way to Run Surveys

A research lead deciding whether to move survey budget from static forms to AI-moderated conversation is really deciding whether the answer they get was ever tested against a real decision. Conversational interviewing fixes inattention: people write more, and write it more carefully, than they do clicking through a grid. It does not fix hypothetical bias: a more articulate transcript is still someone describing what they'd do, not doing it. The smarter run is a randomized discrete choice experiment validated against real behavior, not a friendlier interface.

What's wrong with checkbox surveys, and does conversation actually fix it?

Checkbox surveys suffer from satisficing: straightlining, speeding, low-effort clicking that quietly corrupts a dataset. Conversation genuinely fixes this part. AI-moderated interviewing is now the fastest-growing corner of the market research industry. Vendors including Listen Labs and Outset pair LLM-moderated conversation with follow-up probing, Listen Labs alongside quant formats like Likert scales and MaxDiff inside a single session, and both are extending "conversation" from one-off studies toward ongoing tracking of customer experience.

The mechanism behind the growth is real. A 1,800-participant web-survey field experiment found that LLM chatbot probing produced more detailed, informative open-ended answers than standard forms (arXiv 2504.13908). People engage longer with a conversation than with a grid, and they say more when something follows up on what they just said. That is a genuine data quality gain over a checkbox form. It is also not the whole story: the same study found that chatbot probing inflated false-positive coding errors, driven by respondent acquiescence bias. More detail didn't just mean more truth. Some of it was agreement dressed up as insight.

Why a more articulate answer still isn't a causal one

Inattention and hypothetical bias are different diseases, and conversation only treats one of them. Satisficing happens when someone hasn't engaged with the question at all. Hypothetical bias happens after full engagement: the respondent has thought carefully, answered sincerely, and still described what they'd do rather than done it. No amount of follow-up probing turns a stated intention into a revealed choice, because the person was never given a real tradeoff with a real cost.

This gap is documented, not speculative. A peer-reviewed comparison in Health Economics found that stated-preference survey responses systematically diverge from revealed, actual behavior, and treated closing that gap as a methods problem requiring bias-reduction techniques, not an interface problem (de Corte et al., 2021). A separate natural experiment using a 2008 US tax rebate compared what people said they'd do with the money against what they actually spent it on, and found the same divergence when the stakes were real money rather than a hypothetical scenario (CEPR/VoxEU). Neither paper is about survey UX. Both are about what happens when you ask someone to narrate a decision instead of putting them through one.

Two failure modes that look identical from the outside

A two-column comparison. Left column, inattention: caused by low engagement, fixed by conversational probing and follow-up questions. Right column, hypothetical bias: caused by a fully sincere but hypothetical answer with no real cost attached, fixed only by a randomized experiment checked against real behavior.
Conversation reduces low-effort answering; it does not close the gap between what people say and what they would actually do.

Conversational interviewing versus a randomized discrete choice experiment

DimensionConversational AI interviewingRandomized discrete choice experiment
What it measuresHow someone narrates a decision when probed for detailHow someone actually chooses when an attribute is manipulated
Data quality gainMore detailed open-text answers than static forms ([arXiv 2504.13908](https://arxiv.org/abs/2504.13908))A measured effect with a confidence interval that covers the simulated population, not a guarantee for the real market
Known failure modeInflated false-positive coding errors from respondent acquiescence bias ([arXiv 2504.13908](https://arxiv.org/abs/2504.13908))Requires a defined attribute set, and the estimate is only as good as the design
Checked against real behaviorRarely; validation usually means engagement or transcript richness, not outcome accuracy93 percent replication accuracy, how often the simulated study reproduces the direction and outcome of the original human study, a validation-set result, not a guarantee for a new market (go.subconscious.ai/paper)
Best forEarly discovery, open-ended exploration, generating hypotheses worth testingDecisions with real budget behind them: pricing, messaging, launch calls, feature tradeoffs

How does a randomized experiment answer the question conversation can't?

A randomized experiment answers it by manipulating an attribute across respondents and measuring whether the response moves the way a real choice would, instead of asking someone to describe their reasoning. McFadden discrete choice, Mixed Logit, and ICLV are estimators, not causal methods on their own; they're the statistical models used to read the results. The causal identification comes from the randomization in the design, not from which model fits the output afterward. Mixed Logit is worth naming specifically because it relaxes the independence-of-irrelevant-alternatives assumption that a flat multinomial logit carries, which matters directly for any preference-share or substitution question: if you're asking which action drives a shift in share between options, that IIA assumption determines whether the answer is trustworthy.

The result of running that design against a simulated population, then checking it against a real human study, is the 93 percent replication accuracy figure: how often the simulated study reproduces the direction and outcome of the original human study, measured on a validation set (go.subconscious.ai/paper). It is not a claim that any new market will replicate at that rate; it's a track record on the studies used to build it, published on the leaderboard rather than asserted. One honest limitation belongs alongside it: some of those held-out human studies are published research, and published research can sit inside a model's training data. The replication protocol is built to guard against that contamination risk. It doesn't make the risk disappear, and no article claiming validation against a human baseline should pretend otherwise.

Where does this leave a buyer deciding between them?

It leaves the decision where it actually sits: conversation for finding the question, a randomized experiment for answering it. If the study is exploratory, mapping what customers are even thinking about before a category exists in your data, a conversational interface is the right tool and the engagement gains are real. If the study is meant to inform a decision with money behind it, which price holds, which message drives the outcome, which feature change moves the metric, a transcript full of well-articulated reasoning is not evidence that the reasoning would survive a real choice. That's the difference between a stated preference and a randomized intervention, and it doesn't close with a better chat interface.

Methods and validation background for readers building this into a research stack is collected on the methods and validation blog. As a next step without talking to anyone: pull one recent decision your team validated with a conversational study, check whether it involved a randomized manipulation of an attribute with a measured effect size, and if it didn't, that's the study worth rerunning as a discrete choice experiment before the next budget cycle. If you want a second opinion on a specific design, the team is reachable through meet.