Skip to content

What Is a Simulated Buyer, and When Should You Trust One?

A simulated buyer is a model of a real audience member, built from demographic, behavioral, and prior-response data, that answers a research question the way that audience actually would. It is not a chatbot improvising an opinion. It is the unit a controlled behavioral experiment runs against before a team commits budget to a launch, price, or message.

The decision this article answers: which research questions can move through simulation, and which ones still need real people. Route the wrong question to the wrong method and a team either ships on a guess it can't defend, or spends a field budget answering a directional question that simulation could have settled.

Decision path from "research question," splitting on whether it needs a defensible population estimate: concept/message/segment questions go to a simulated buyer; population-estimate questions go to real-human fielding.
The question's demands, not its topic, decide whether simulation or real-human fielding is the trustworthy method.

What makes a simulation trustworthy rather than a guess

Three layers determine whether a simulated buyer's answer means anything:

A frontier reasoning model. GPT-class, Claude-class, Gemini-class systems supply the general language and reasoning ability underneath every answer.

Audience conditioning. Demographic and behavioral inputs, such as age, geography, household income, occupation, attitudes, and prior brand exposure, bind the model to a specific audience segment rather than a generic voice. Conditioning on real prior response data for that audience is what separates a calibrated simulation from a costume.

A response protocol. Rules for how the model answers: question format, scale, and follow-up handling. A protocol that accepts every answer at face value produces a chatbot. A protocol that scores disagreement, flags low-confidence answers, and reproduces similar results on a repeat run produces something you can audit.

Four properties separate an audited simulation from a thin prompt wrapper:

Questions that reward simulation, and ones that don't

The dividing line is whether the question rewards general reasoning about preferences or demands unique lived experience the model cannot have had.

Simulation handles these well:

Simulation handles these poorly:

Each calls for reasoning about preference, reaction, or evaluation criteria, something a well-conditioned model can approximate. The insurance-switch question instead demands invented autobiographical detail, and a model will fabricate specifics rather than admit it has none.

Simulated testing versus real-human fielding

DimensionSimulated testingReal-human fielding
Historical study-schedule exampleMinutes to hours in one inherited setupA 3 to 6 calendar-week range in one inherited planning example
Historical field-cost exampleAmortized across a subscription or platform cost in one inherited setupThousands to tens of thousands per field in one inherited planning example
Iteration designRepeat the same protocol to check stabilityTreat each additional field as a new study decision
Hard-to-reach audiencesStraightforward to configureOften impractical
Statistical validationDirectional signalDefensible population estimates with probability sampling; model-dependent for opt-in panels
Novel-category predictionUnreliable outside training dataGenuine signal
Sensory or emotional responseLimited: the model can reason about a reaction, not feel oneFull

The timing and cost figures are inherited historical planning examples. They are not current Subconscious or vendor prices, delivery estimates, guarantees, or service levels.

The pattern that holds up: route concept screening, message iteration, segment exploration, and multi-market comparison to simulation first. Reserve real-human fielding for population-level validation, hero claims, and any number that has to survive regulatory or PR scrutiny. Treating that split as roughly 80 percent simulation-first work and 20 percent real-fielding work is a reasonable planning heuristic for a research queue, not a fixed ratio every program will hit.

What a simulated study group looks like

Most teams run simulated buyers in groups rather than one at a time:

Where simulation is the wrong tool

Three situations call for real fielding instead:

Statistically validated population claims. Anything a team needs to defend as "X percent of the target population thinks Y" requires a real study designed to produce that estimate.

Genuinely novel categories. Products or events with no analog the model has seen before produce plausible-sounding guesses with no signal behind them.

Sensory or emotional response. Reactions to a physical product, a package design, or a video ad require real human perception. A model can reason about the likely reaction; it cannot feel one.

Moving from a simulated result to a validated one

The useful capability is testing the same causal question, does this concept, price, or message change buyer behavior, through simulation first, then real-human validation when the decision calls for it, without changing what the study measures. Subconscious can test or validate a study with real human participants, so a team that started with a simulated screen can move to real-human validation on the same causal question rather than starting over.

Subconscious's method is a controlled causal experiment on a simulated market, benchmarked for replication accuracy against real human studies.

Subconscious tests this against roughly 300 replicated human studies across 9 domains. Our best configuration reaches 87% of the measured human ceiling on one study: 0.832 rank correlation against the published human result, where two independent samples of real humans reach 0.959. Across all 43 studies that pass design filters the mean is 0.73. It is a validation result, not a guarantee for a new market. See the causal fidelity paper.

Where the term comes from, and its limits

The academic backbone is Argyle and colleagues' 2023 paper on conditioning a frontier model on a real respondent's demographic background to produce opinion distributions that match benchmark surveys, an approach the literature calls silicon sampling (Cambridge University Press, Political Analysis). The commercial category built products around that idea afterward.

The boundary that matters for a buyer is not whether the underlying model is impressive. A study design that tracks that boundary, and that can hand off a directional finding to a validated one without redefining the question, is what separates a research program from a stack of plausible-sounding guesses.

Next step: check the leaderboard for how simulated results have tracked against real-human studies, or see how Subconscious runs a study end to end.