AI Customer Conversations vs. Controlled Experiments: What Each One Can Prove
Before a launch, a pricing change, or a new message ships, most teams face the same choice: talk to a chatbot that stays in character as a customer, or run a controlled test that measures what a defined population would actually choose. The two produce different kinds of evidence. A conversation gives you hypotheses to test. A randomized comparison gives you an estimated effect, with uncertainty, that you can examine in a launch review, as long as the review also asks what that estimate is and is not evidence for.
The appeal of an open-ended AI conversation
An AI-driven conversation lets you type questions to a chat interface configured as a customer and get answers back in real time. You can ask follow-ups, probe an unexpected answer, and get a transcript that reads like a real interview. That immediacy is why teams reach for it: two weeks out from a launch with no time to recruit interviewees, testing several positioning directions before committing research budget to one, rehearsing an interview guide before running it with real customers, or getting oriented on a buyer segment the team has never sold to.
None of that is wrong as a way to move faster. For example, a team that might realistically fit 5 to 10 real interviews into a research cycle could run 20 or 50 such conversations, across as many as 10 different customer types, in the same window. That ratio describes conversation throughput, not a validated or current Subconscious benchmark. The problem shows up when the output of any single one of those conversations gets treated as measured signal rather than a fluent guess.
What a fluent transcript cannot tell you
A chat-style AI conversation that stays in character produces answers that sound plausible and internally consistent. Consistency is not the same as accuracy. Three published studies show why a team should test a synthetic respondent against people before relying on it:
- Spread of answers. Bisbee and colleagues compared ChatGPT-generated survey answers with a human survey on political attitudes. The generated answers had far smaller standard deviations than the human answers, were sensitive to prompt wording, and changed between April and July 2023 under the same prompt (Political Analysis, 2024).
- Demographic groups. Wang, Morgenstern and Dickerson tested four language models against 3,200 human participants across 16 identities. They report that the models misportray and flatten identity groups, and that prompting mitigations reduce but do not remove the problem (arXiv, 2024).
- Purchase intent. Maier and colleagues (including authors from PyMC Labs) tested a Semantic Similarity Rating method on 57 personal-care product surveys with 9,300 human participants. They report 90% of human test-retest reliability for that task and that the method works under specific elicitation, not by default (arXiv, October 2025). The result covers personal-care surveys, not every product category.
A fourth paper is about something different. Sfeir and colleagues tested whether language models can help an analyst specify and estimate logit choice models. They found that proprietary models can produce "valid and behaviourally sound utility specifications," especially with structured prompts (Journal of Choice Modelling, 2026). That is evidence about modeling assistance, not about how simulated respondents behave.
None of that makes the transcript useless. It means a single open-ended conversation, run once, with no comparison group and no uncertainty estimate, is a hypothesis, not a validated finding. A team that ships a launch, price, or message decision on that hypothesis alone trades the cost of waiting for real signal against the cost of being wrong in a way nothing in the conversation flags.
A different question: comparison, not conversation
The underlying need in most of the scenarios above is not "have a conversation" but "find out which of these options a defined population would actually choose." That is a comparison question, and it calls for a controlled, randomized experiment across a defined population rather than a single scripted character. Subconscious.ai tests product, pricing, and messaging actions with randomized discrete-choice experiments on a simulated population. Instead of asking one configured character what it thinks, the study randomizes an action across many simulated respondents and estimates the effect on their simulated stated choice.
| Open-ended AI conversation | Randomized simulation | |
|---|---|---|
| Primary output | A transcript from one configured respondent | An estimated effect on simulated stated choice across a specified population, with uncertainty |
| What varies | Your follow-up questions | The action being tested (price, message, feature), randomized across simulated respondents |
| How you judge it | You judge the transcript's plausibility | You judge the design, the population definition and the uncertainty, then compare against a human study |
| Best use | Fast orientation, drafting interview questions, early exploration | Choosing among finalists before you commit launch or roadmap budget |
The simulated population has to match the decision. A study needs a defined, relevant population (for example, "US grocery shoppers who bought a private-label product in the last month") and a stated way of building it. Ask any provider where its population data comes from, what it covers, and how it was validated for a segment like yours. This article makes no claim about audience-data reach.
Where a conversation with a real person is still the right tool
Neither a scripted AI conversation nor a randomized experiment replaces talking to an individual customer. Real interviews remain the right method when the goal is specific, idiosyncratic detail from one person, a genuinely novel reaction that a model has no basis to predict, or the final validation step before a high-stakes decision ships. The instinct to run a fast, exploratory step before committing real interview time is sound; the difference is what that step should measure. A conversation that stays in character surfaces questions worth asking. A controlled experiment gives you a comparison to act on.
Moving from a measured comparison to human validation
When the stakes justify it, replicate the same causal question with real participants. That means a matched-question replication, not the same file run on people. The team has to check four things:
- Instrument: people see the same attributes, levels and task wording, adapted for a human survey, with any change recorded.
- Sample: recruited respondents match the population the simulation described, and a quota or screening rule says how.
- Outcome: a stated choice among people is compared with a simulated stated choice, so agreement supports a claim about stated choice and nothing more.
- Behavior: a claim about purchases needs an incentive-compatible purchase, a transaction holdout or a bounded live test.
This check preserves what an open-ended conversation cannot offer on its own: a way to see where the answer fails. Who recruits and fields the human study, and on what timeline, is scoped in a decision review. This article does not describe a standard service.
For a launch, pricing, or messaging decision, start by defining the comparison you need answered, then decide whether a quick conversation or a randomized study is the right way to get it. The replication leaderboard shows the published fidelity evidence (rank agreement on choice parameters across studies), How it works explains the study design, and case studies show customer examples. To scope a study for your decision, book a decision review.