Skip to content

How to get the most out of open-ended questions

Fixed the em dashes, the citation year mismatch, the hedged training-data caveat, the unqualified Displayr claim, and the two table cells that were floating free of the body's definitions and limitations. Structure, headings, and everything else are unchanged.

An insights director staring down three thousand open-ended verbatims before a launch decision needs one thing: to know which comments describe the real reason people chose what they chose, and which are filler. The direct answer is that coding the text better, even with a well-built AI pipeline, only cleans up self-report. To find out which claimed reason actually moved the choice, you need a randomized discrete-choice experiment on the same population, with a confidence interval attached to the answer. Verbatims tell you what respondents say moved them; only a randomized manipulation tells you what did.

- Open-ended text captures stated reasons filtered through recall bias, social desirability, and satisficing. AI coding organizes that text; it does not remove the bias underneath it.
- 95 percent of market researchers now use AI tools regularly or experimentally for tasks like verbatim coding, and 66 percent rely on AI built into their research software, per Qualtrics' 2026 Market Research Trends report, based on a Q3 2025 survey of 3,000+ researchers.
- At least nine distinct AI coding methods are in active commercial use, with no single validated standard among them, per Displayr's 2026 catalog.
- The fix is not a better codebook. It is a randomized experiment on the same population, analyzed with Mixed Logit or ICLV, that tests which of the stated reasons actually carried causal weight.
- A validated simulation protocol reproduces the direction and outcome of published human studies 93 percent of the time on its validation set (go.subconscious.ai/paper); current studies and their scores are public on the [leaderboard](/leaderboard).

## The open-end problem no coding tool fixes

Panel-scale data collection has buried analysts in verbatims faster than any team can read them. That's the honest starting point: most of that text was never going to be signal. A meaningful share of any open-ended corpus is satisficer noise, typed answers like "good," "fine," or a copy-pasted sentence from three questions earlier, produced by respondents trying to finish the survey, not answer the question. That noise exists before any AI model touches the data. Coding it faster does not make it truer.

## What's the fastest way to make sense of thousands of open-ended verbatims?

Right now, the fastest way is AI-assisted coding, and the market has already standardized on using it, just not on how. Qualtrics' 2026 Market Research Trends report, based on a Q3 2025 survey of more than 3,000 market research professionals across 14 countries, found 95 percent now use AI tools regularly or experimentally, with 66 percent relying on AI embedded directly in their research software, up from 62 percent in 2024 ([Qualtrics, 2026 Market Research Trends](https://www.qualtrics.com/articles/research-teams-not-using-ai-are-four-times-more-likely-lose-organizational-influence/)). Tools from Qualtrics Text iQ, Ascribe/Forsta, and others now run sentiment scoring, automated theme extraction, and human-in-the-loop code assignment. Displayr's own catalog counts nine separate AI methods in active use for coding open-ends, a vendor listing rather than an independent survey of practice ([Displayr, 2026](https://www.displayr.com/9-ai-methods-to-code-open-ended-survey-responses-in-2026/)). Nine competing approaches, even by a vendor's own count, is evidence of a fragmented practice, not a validated one.

## Where AI-assisted coding still breaks

It still breaks at the same place human coding always broke: turning what someone typed into what someone meant. A JMIR AI study directly compared LLM-generated thematic summaries against human-coded ones in qualitative health care research and found LLMs can approximate human themes but do not reliably replicate them without independent validation ([JMIR AI, 2025](https://ai.jmir.org/2025/1/e64447)). A 2025 review of LLM-assisted thematic analysis goes further, warning that heavy reliance on automation risks "premature closure," where the model settles on a plausible-sounding theme structure before the actual variation in the data has been explored ([arXiv, 2511.14528](https://arxiv.org/html/2511.14528)). Hallucination and prompt sensitivity compound the problem: change the prompt, get a different codebook, from the same verbatims.

## The gap between what people say and what moves them

None of that fixes the deeper issue: a theme's prevalence in a codebook is not the same measurement as its causal weight in a decision. A verbatim that fifty respondents typed and a discrete-choice effect that shifted fifty respondents' actual choices both arrive as a number, which makes them easy to mistake for the same kind of evidence. They are not, because they were produced by different processes.

![Two parallel process chains. The top chain runs from respondent recall to typed reason to AI-coded theme to a prevalence percentage. The bottom chain runs from random assignment to discrete choice to a Mixed Logit or ICLV estimate to a confidence interval.](/images/authority/get-most-out-open-ended-questions.svg "A verbatim's prevalence and a discrete-choice model's effect size both look like proof, but only one traces back to a randomized manipulation.")

## Can you get a causal answer straight from what people typed?

No. Self-report, however well coded, has no random assignment in it, and without random assignment there's no way to separate what caused a choice from what a respondent believes or claims caused it. Discrete choice models, McFadden's original conditional logit, Mixed Logit, and ICLV, are estimators, not causal methods on their own. The causal identification comes from the randomized manipulation built into the experiment design: respondents are randomly shown different combinations of attributes and prices, and the model estimates which attribute actually shifted the choice. (A standard multinomial logit also carries the independence-of-irrelevant-alternatives assumption, which is why substitution patterns are usually checked against a Mixed Logit before they're trusted.) That's the structural difference between "why respondents say they chose it" and "why they chose it."

## What the 93 percent replication number does and doesn't cover

It means the simulation reproduced the direction and outcome of the original human study 93 percent of the time, on a validation set of published studies (go.subconscious.ai/paper). It is not a guarantee for a new, unpublished market question, and it comes with a caveat that has to be stated plainly: some of those published studies can sit in a model's training data, which is exactly why the replication protocol tests against held-out studies rather than resting on a single comparison. Current studies and their individual scores are tracked in the open on the [leaderboard](/leaderboard), so a buyer can check the number against a live record rather than a claim in a deck.

## Open-ended coding vs a randomized experiment: what each actually proves

| | Open-ended coding (AI-assisted) | Randomized discrete-choice experiment |
|---|---|---|
| What it measures | What respondents say moved them, after recall bias and social desirability | Which claimed reason actually shifts the choice, under random assignment |
| Output | Theme prevalence, sentiment score | Effect size with a confidence interval, under random assignment (McFadden DCE, Mixed Logit, ICLV) |
| Main risk | Satisficer noise, hallucinated themes, prompt-sensitive coding | Model misspecification, IIA violations if only a flat logit is used |
| Validated against | Human coder agreement, unevenly, across nine competing methods | 93 percent replication accuracy on a validation set of published studies, training-data caveat applies (go.subconscious.ai/paper) |
| Best for | Generating the list of candidate reasons worth testing | Deciding which candidate reason to bet a launch on |

## How to combine both, and what to do next

Use the open-ends for what they're good at: generating the list of candidate reasons, in respondents' own language, before you spend anything on testing. Then take that list into a randomized experiment on the same population and let the discrete-choice model tell you which one actually moved the choice, with a confidence interval that covers the effect within that simulated population, not a bound on how the real market will behave in every future condition. Treat the verbatims as hypotheses, not conclusions, and treat the experiment's confidence interval, not the theme's word count, as the number you brief the launch decision on. For a worked example of this pairing in practice, see the [methods and validation](/blog/methods-and-validation) archive and a related [case study](/case-studies).

Concrete next step: pull last quarter's top three open-ended themes, write each as a testable claim, and run them as attributes in a randomized discrete-choice experiment before the launch decision locks. If you want a second set of eyes on the design, [meet the team](/meet).