What Is Synthetic Market Research? A Buyer's Guide to the Method and Its Limits
Synthetic market research conditions AI personas on demographic and behavioral inputs, then uses them to model the way a defined consumer or B2B audience might react to a survey, a concept test, ad creative, or messaging variants. The participants are modeled rather than recruited: a team describes the audience, configures the personas, and runs the session against a language model.
The method is also called AI market research, simulated market research, or virtual market research. The real decision it raises for a research or marketing leader isn't whether the method works. It's where in a research pipeline a simulated answer is good enough, and where the cost of being wrong requires a controlled experiment on real people before money moves.
Where the method comes from
The intellectual lineage is academic. In a 2023 study, Argyle et al. showed that a language model conditioned on a real survey respondent's demographic backstory could generate opinions whose spread lined up closely with how real Americans had actually answered benchmark surveys such as the ANES (Political Analysis, Cambridge University Press). That paper established silicon sampling as a viable technique, and commercial platforms packaged it into personas, panels, and workflows for marketing, product, and insight teams.
What a synthetic-first workflow looks like
A typical simulated study breaks into five steps:
- Define the audience. Age range, geography, household income, job role, industry, attitudes, prior brand exposure. The more specific the definition, the more useful the simulation.
- Configure personas. Assemble individual simulated participants into a research panel. Many teams run 50 to 500 personas per study, calibrated against any real-respondent data already held for that audience.
- Design the research instrument. The same survey, concept-test brief, or ad-pretest stimulus a team would field traditionally.
- Run the session. Each persona answers in natural language. Closed questions produce structured output; open prompts produce responses a team can read and theme.
- Synthesize and decide. Compare segments, identify a leading concept or message, and decide whether the finalists need a real-respondent check before the team commits budget.
The loop produces a directional simulated result. It does not produce evidence of actual behavior.
When directional screening is useful
Five situations, drawn from how research and marketing teams use the method:
- Directional narrowing. Cutting a dozen concepts down to a shortlist before commissioning real-respondent work.
- Iterative exploration. Applying the same instrument to successive concepts while treating the output as directional.
- Specialized audiences. Modeling hypotheses about senior B2B buyers, regulated professionals, or niche geographies before validating the decision on real people.
- Cross-market comparison. Applying the same instrument across several country definitions and comparing directional differences under a consistent setup.
- Sensitive topics. Exploring hypotheses about health, finance, or employment before designing research involving real participants.
Where correlation stops answering the question
A simulated response can track what a real audience tends to say. It does not, by itself, tell a team what happens when a specific input changes: a price, a headline, a feature. That is a causal question, and a correlation between simulated and real answers on past surveys is not evidence about a causal effect the team hasn't tested yet.
Three limits follow from that gap:
- Statistical validation. A simulated study does not produce a population estimate with a defensible confidence interval on its own.
- Genuinely novel behavior. A persona only echoes patterns already present in the model's training. It can't be trusted on a product, category, or event that has nothing comparable in that history.
- Final go/no-go calls. Capital allocation, regulatory filings, and public claims should not rest on simulated data alone.
The cost of skipping this step shows up later: a team that commits budget on the strength of a simulated correlation, without checking it against real holdout behavior, finds out the difference only after the launch.
Moving from directional screening to a causal test
The mature pattern is hybrid: use simulation to narrow the field, then run a randomized, controlled experiment on the audience in question for the option that will get funded. That second step is where Subconscious fits. Subconscious can test or validate studies with real human participants, and it can run controlled studies against a person-level audience graph covering 800 million real people, kept distinct from any panel a team recruits directly. The result of the controlled step is a causal effect with a confidence interval, not another correlated guess.
| Stage of the decision | Question being asked | Right tool |
|---|---|---|
| Directional screening | Which of many concepts looks promising? | Simulated iteration |
| Concept narrowing | Which one to three options deserve real budget? | Simulated iteration, then a shortlist |
| Real-world validation | Does the shortlisted option cause the outcome we want? | A randomized, controlled experiment on real people |
| Go/no-go decision | Do we commit budget to this option? | The controlled experiment's effect and confidence interval |
Subconscious does not replace direct observation of real behavior, and it does not claim any specific accuracy percentage against real respondents.
Privacy and compliance
Simulated participants are generated, not recruited, so a session usually has no real personal data to process in the first place, and that removes much of the consent and data-retention complexity traditional fielding carries. For organizations with strict compliance requirements, such as healthcare, finance, or the public sector, that can make simulated screening easier to deploy early.
Adjacent terms, clarified
- Synthetic data. Artificial datasets built to train models, or to pad out a real sample that's too small. A related but different problem.
- AI personas. The individual unit inside a simulated research panel. The persona is the agent; synthetic market research is the method built around it.
- AI focus groups. The qualitative format of simulated research, where personas respond as a group rather than individually.
- Multi-turn persona research. A newer extension where AI personas act and react to follow-up stimuli rather than only answering a fixed instrument.
Frequently asked questions
What is synthetic market research?
A method where AI-generated personas stand in for a defined consumer or B2B audience, so a team can gauge how that audience might react to surveys, concepts, ads, or messaging. Each persona is conditioned on demographic, psychographic, and behavioral inputs and queried in natural language, without the recruitment and fielding a traditional study requires.
When should a team use simulation versus a controlled experiment on real people?
Use simulation for directional input: early concept testing, ad pretesting, messaging iteration, specialized B2B audiences, and multi-market comparisons. Use a controlled experiment on real people for the claims that carry the launch decision, where a causal effect and a confidence interval, not a correlated guess, need to back the call.
Does simulation replace research on real people?
No. It supports directional iteration, not the step that proves an effect. The pattern that holds up is hybrid: narrow the field with simulation, then validate the finalists with a controlled study on the real audience before allocating spend.
A team that treats every stage in that path the same way can either overstate what simulation proves or apply controlled research before the question has been narrowed. Research shows how the controlled step works in practice, case studies show it applied to real launches, and a demo or a look at how Subconscious works shows where the causal step would sit in an existing pipeline.