What Is a Simulated Buyer, and When Should You Trust One?
A simulated buyer is a model of a real audience member, built from demographic, behavioral, and prior-response data, that answers a research question the way that audience actually would. It is not a chatbot improvising an opinion. It is the unit a controlled behavioral experiment runs against before a team commits budget to a launch, price, or message.
The decision this article answers: which research questions can move through simulation, and which ones still need real people. Route the wrong question to the wrong method and a team either ships on a guess it can't defend, or spends a field budget answering a directional question that simulation could have settled.
What makes a simulation trustworthy rather than a guess
Three layers determine whether a simulated buyer's answer means anything:
A frontier reasoning model. GPT-class, Claude-class, Gemini-class systems supply the general language and reasoning ability underneath every answer.
Audience conditioning. Demographic and behavioral inputs, such as age, geography, household income, occupation, attitudes, and prior brand exposure, bind the model to a specific audience segment rather than a generic voice. Conditioning on real prior response data for that audience is what separates a calibrated simulation from a costume.
A response protocol. Rules for how the model answers: question format, scale, and follow-up handling. A protocol that accepts every answer at face value produces a chatbot. A protocol that scores disagreement, flags low-confidence answers, and reproduces similar results on a repeat run produces something you can audit.
Four properties separate an audited simulation from a thin prompt wrapper:
- Fidelity to a real audience. The model is calibrated against real prior data for that segment, not just a job title and an age typed into a prompt.
- Disagreement and pushback. A real respondent says "I would not buy this," misreads the question, or changes position under a follow-up. A model that agrees with everything is answering as a chatbot, not as the audience.
- A confidence signal. Every answer should carry an estimate of how reliable it is, so a team can flag the low-confidence ones instead of treating every output as settled.
- Reproducibility. Run the same setup against the same stimulus again and the result should land in the same range, not swing wildly. Subconscious tests this property against real-human studies rather than assuming it.
Questions that reward simulation, and ones that don't
The dividing line is whether the question rewards general reasoning about preferences or demands unique lived experience the model cannot have had.
Simulation handles these well:
- "Would you buy any of these product concepts?"
- "What's off-putting about this messaging?"
- "Describe your process for evaluating a vendor switch."
- "What would push you from your current vendor to a competitor?"
- "Is this ad creative confusing in any way?"
Simulation handles these poorly:
- "Describe the specific moment last summer when you switched insurance providers."
Each calls for reasoning about preference, reaction, or evaluation criteria, something a well-conditioned model can approximate. The insurance-switch question instead demands invented autobiographical detail, and a model will fabricate specifics rather than admit it has none.
Simulated testing versus real-human fielding
| Dimension | Simulated testing | Real-human fielding |
|---|---|---|
| Historical study-schedule example | Minutes to hours in one inherited setup | A 3 to 6 calendar-week range in one inherited planning example |
| Historical field-cost example | Amortized across a subscription or platform cost in one inherited setup | Thousands to tens of thousands per field in one inherited planning example |
| Iteration design | Repeat the same protocol to check stability | Treat each additional field as a new study decision |
| Hard-to-reach audiences | Straightforward to configure | Often impractical |
| Statistical validation | Directional signal | Defensible population estimates with probability sampling; model-dependent for opt-in panels |
| Novel-category prediction | Unreliable outside training data | Genuine signal |
| Sensory or emotional response | Limited: the model can reason about a reaction, not feel one | Full |
The timing and cost figures are inherited historical planning examples. They are not current Subconscious or vendor prices, delivery estimates, guarantees, or service levels.
The pattern that holds up: route concept screening, message iteration, segment exploration, and multi-market comparison to simulation first. Reserve real-human fielding for population-level validation, hero claims, and any number that has to survive regulatory or PR scrutiny. Treating that split as roughly 80 percent simulation-first work and 20 percent real-fielding work is a reasonable planning heuristic for a research queue, not a fixed ratio every program will hit.
What a simulated study group looks like
Most teams run simulated buyers in groups rather than one at a time:
- 50 to 500 modeled respondents (a nominal count, not an effective sample size: responses share a common model error source rather than being independent draws)
- Stratified across the demographic and behavioral dimensions that matter to the decision
- Calibrated against real prior data for that audience when it's available
- Run against a defined instrument: a concept test, an ad pretest, or a structured comparison
- Output as structured comparison data alongside open-ended qualitative response
Where simulation is the wrong tool
Three situations call for real fielding instead:
Statistically validated population claims. Anything a team needs to defend as "X percent of the target population thinks Y" requires a real study designed to produce that estimate.
Genuinely novel categories. Products or events with no analog the model has seen before produce plausible-sounding guesses with no signal behind them.
Sensory or emotional response. Reactions to a physical product, a package design, or a video ad require real human perception. A model can reason about the likely reaction; it cannot feel one.
Moving from a simulated result to a validated one
The useful capability is testing the same causal question, does this concept, price, or message change buyer behavior, through simulation first, then real-human validation when the decision calls for it, without changing what the study measures. Subconscious can test or validate a study with real human participants, so a team that started with a simulated screen can move to real-human validation on the same causal question rather than starting over.
Subconscious's method is a controlled causal experiment on a simulated market, benchmarked for replication accuracy against real human studies.
Subconscious tests this against roughly 300 replicated human studies across 9 domains. Our best configuration reaches 87% of the measured human ceiling on one study: 0.832 rank correlation against the published human result, where two independent samples of real humans reach 0.959. Across all 43 studies that pass design filters the mean is 0.73. It is a validation result, not a guarantee for a new market. See the causal fidelity paper.
Where the term comes from, and its limits
The academic backbone is Argyle and colleagues' 2023 paper on conditioning a frontier model on a real respondent's demographic background to produce opinion distributions that match benchmark surveys, an approach the literature calls silicon sampling (Cambridge University Press, Political Analysis). The commercial category built products around that idea afterward.
The boundary that matters for a buyer is not whether the underlying model is impressive. A study design that tracks that boundary, and that can hand off a directional finding to a validated one without redefining the question, is what separates a research program from a stack of plausible-sounding guesses.
Next step: check the leaderboard for how simulated results have tracked against real-human studies, or see how Subconscious runs a study end to end.