What Is a Simulated Buyer, and When Should You Trust One?
A simulated buyer is a modeled stand-in that attempts to approximate responses from a defined audience. Inputs can include demographic profiles, supplied context or prior-response data. The definition does not guarantee fidelity: relevant held-out human responses or behavior must establish what the configuration can predict.
The decision this article answers: which research questions can move through simulation, and which ones still need real people. Route the wrong question to the wrong method and a team either ships on a guess it can't defend, or spends a field budget answering a directional question that simulation could have settled.
What makes a simulation trustworthy rather than a guess
Three layers determine whether a simulated buyer's answer means anything:
The response model. Record which model and version generates answers, what supplied context it uses and the response settings.
Audience conditioning. Record the profile data, provenance, coverage and weighting. Supplying an age or occupation does not by itself establish representation or calibration.
A response protocol. Define the stimuli, scales, task order, repeated draws and treatment assignment. Document exclusions and scorer behavior before seeing the result.
Four properties separate an audited simulation from a thin prompt wrapper:
- Matched fidelity. Compare held-out human outcomes for the intended audience, stimulus and task.
- Disagreement and pushback. A real respondent says "I would not buy this," misreads the question, or changes position under a follow-up. A model that agrees with everything is answering as a chatbot, not as the audience.
- Calibrated uncertainty. Distinguish self-reported confidence, variation between generated draws and empirical human agreement. If reliability probabilities or intervals are reported, request held-out calibration or coverage checks for the outcome.
- Stability with adequate heterogeneity. Check seeds, prompts and versions, alongside human variation. Stable repetition alone does not establish accuracy.
Which questions reward simulation, and which don't?
The dividing line is evidence for the outcome and audience, together with the consequence of a wrong answer. A preference question can be high stakes and outside the validated range.
Possible hypothesis-generation tasks include:
- "Would you buy any of these product concepts?"
- "What's off-putting about this messaging?"
- "Describe your process for evaluating a vendor switch."
- "What would push you from your current vendor to a competitor?"
- "Is this ad creative confusing in any way?"
A request that requires a real personal history is:
- "Describe the specific moment last summer when you switched insurance providers."
The first questions can help generate hypotheses, but a plausible answer does not establish commercial validity. The insurance-switch request asks the model to invent a personal event it did not experience.
Simulated testing versus real-human fielding
| Dimension | Simulated testing | Real-human fielding |
|---|---|---|
| Evidence scope | Conditional modeled responses; requires relevant external checks | Observed responses from a defined sample; sampling and measurement still matter |
| Field effort | Depends on configuration, design and required validation | Depends on recruitment, instrument and study scope |
| Iteration | Repeat prompts and stimuli with versioned settings | Recontact or new waves where consent and design permit |
| Specialized audiences | Easy to name in a prompt; representation still needs proof | Recruitment and coverage need an explicit plan |
| Population inference | Requires justified modeling and validation assumptions | Probability-sample or justified model-based inference with stated assumptions |
| Novel scenarios | Fresh context can be supplied; validate against relevant outcomes | Measure responses or behavior to the scenario |
| Sensory response | A generated description does not observe perception | Observe participant experience with the stimulus |
Compare actual scope, effort and evidence rather than inherited timing or price examples.
Route the research queue using matched validation, model/population mismatch, cost of error and budget. Preserve uncertain candidates when a false rejection would be costly. There is no justified default numerical split between simulation and human fielding.
What a simulated study group looks like
Most teams run simulated buyers in groups rather than one at a time:
- Set the nominal number of generated respondents from precision and stability checks; it is not automatically an effective sample size.
- Stratified across the demographic and behavioral dimensions that matter to the decision
- Calibrated against real prior data for that audience when it's available
- Run against a defined instrument: a concept test, an ad pretest, or a structured comparison
- Output as structured comparison data alongside open-ended qualitative response
Where simulation is the wrong tool
Some questions require an additional source of evidence:
Population claims. Specify the target population and the sampling or model-based assumptions needed for the estimate; validate the response model if synthetic outputs contribute.
Novel scenarios. Check fresh context, similarity to validated tasks and relevant human or market outcomes before treating a prediction as evidence.
Sensory or emotional response. Reactions to a physical product, a package design, or a video ad require real human perception. A model can reason about the likely reaction; it cannot feel one.
How do you move from a simulated result to a validated one?
Keep the stimuli and intended estimand aligned between a simulated comparison and an external check. A human stated-choice task, a prototype task and a live purchasing experiment measure different outcomes, so do not describe them as interchangeable validation.
Subconscious's method is a controlled causal experiment on a simulated market, benchmarked for replication accuracy against real human studies.
The public causal-fidelity working paper evaluates agreement on estimated choice parameters in published studies. That is task-specific benchmark evidence, rather than a guarantee that a simulated buyer represents your market or predicts realized sales.
Where the term comes from, and its limits
Argyle and colleagues' 2023 Political Analysis paper studies conditioning language models on respondent backgrounds for political-survey tasks. It provides one foundation for silicon sampling, rather than proof for all commercial questions.
The boundary that matters for a buyer is not whether the underlying model is impressive. A study design that tracks that boundary, and that can hand off a directional finding to a validated one without redefining the question, is what separates a research program from a stack of plausible-sounding guesses.
Next step: check the leaderboard for how simulated results have tracked against real-human studies, or see how Subconscious runs a study end to end.