Synthetic Research: When to Trust It, and When to Validate With Real People
A VP of Insights deciding whether to add synthetic research to the toolkit is really deciding which of this quarter's research questions simulation can answer, and which one still needs a recruited human study before capital moves. Get that split wrong: a recruited-research budget gets spent validating a question simulation could have answered in an afternoon, or a pricing, go-to-market, or public claim ships on a directional read that cannot produce a defensible confidence interval, and unravels under scrutiny.
What synthetic research actually is
Synthetic research uses AI-generated personas, conditioned on demographic, psychographic, and behavioral data, to simulate how a target population responds to a stimulus: a concept, a message, a poll question. Researchers assemble personas into panels, run a stimulus against the panel, and get back a distribution of responses with natural-language reasoning for each one.
Three terms often get conflated:
- Synthetic respondent: the individual AI agent that answers a single study, conditioned to hold a specific set of beliefs and background.
- Synthetic persona: the reusable profile behind a respondent, covering demographics, psychographic traits, and decision-making frameworks. Saved once, queried across projects.
- Synthetic panel: a structured group of personas, typically ranging from 8 to 100 or more, assembled to represent a market segment. A panel study might show that 60 percent of personas accepted a feature concept, 30 percent raised a specific objection, and 10 percent asked about pricing: a distribution, not a single verdict.
The premise traces to a 2023 finding: conditioning a language model on the detailed background of a real poll respondent produced opinion distributions that tracked actual human responses in benchmark national polls (Argyle et al., a study on prompting language models with individual respondent profiles to reproduce human sample distributions, published in Political Analysis by Cambridge University Press). That result moved the method from academic benchmarking into product, marketing, and insight teams.
Where the method stops being enough
Simulation is directionally useful and limited in predictable ways:
- It doesn't produce statistical validation. Simulation cannot generate a population estimate with a defended confidence interval. If a regulator, auditor, or public claim needs a stated figure, such as 34 percent of a population holding a given view, that number has to come from recruited research.
- It lags on novel behavior. Personas are built on historical patterns. A category with no real-world analog, or a sudden macroeconomic shift, can outpace it.
- It inherits training-data skew. Models trained heavily on English-language, Western text default to generalized assumptions for audiences underrepresented in that data. A community outside the training distribution needs real members validating the finding.
- It doesn't touch the physical world. A simulated persona doesn't pull out a credit card, sit through a shipping delay, or churn after a bad support call. Longitudinal behavioral tracking still needs real behavioral data.
The decision: synthetic alone, recruited alone, or both in sequence
| Situation | Right method |
|---|---|
| Early-stage concept, message, or ad-variant testing; competitive scoping; a target audience that's slow or expensive to recruit (senior B2B buyers, niche specialists); privacy-sensitive contexts | Synthetic alone |
| A single, capital-committing go-to-market or pricing decision; a quantitative claim for external publication or PR; a regulatory submission or legal evidence | Recruited alone |
| Narrowing many options down to the 1 to 3 that matter before a final call | Synthetic first, recruited second |
The third row is the pattern most teams underuse. Run synthetic research first to explore the landscape, test variations, and refine the research instrument. Then field a smaller, targeted study with recruited participants against only the finalists. That sequencing lowers recruitment cost because the human study only tests survivors, and raises confidence in the final number because the questions were already stripped of obvious flaws before a real person answered them.
Where Subconscious sits in that sequence
Subconscious runs controlled discrete choice experiments, using causal DCE, Mixed Logit, and ICLV methodology, that return causal effects with confidence intervals rather than a directional read. That's the rigor layer synthetic-alone methods can't supply: a defensible population estimate, with error bars, for the moment a finding needs to survive audit or public scrutiny.
The practical advantage: a team can move from a simulated experiment to a real-human study without changing the underlying causal question, so the narrow-with-simulation, confirm-with-people sequence runs on the same experimental design, rather than handing the finalists to a separate agency running a different methodology.
The fit is bounded, too: Subconscious's positioning is the validation and causal-inference layer for decisions that carry capital risk, not a claim to replace every recruited-human study with simulation. A regulatory submission, a legal filing, or a public percentage claim still needs the sourced, audited number a controlled study produces.
Running the sequence without wasting the human study
- Define the decision, not just the topic. Name the population, the alternatives being compared, and the outcome that decides the action, not just "test the messaging."
- Run the simulated pass first. Test the full set of variants, including the ones expected to lose, so the elimination is evidence-based rather than assumed.
- Narrow to the finalists. Carry forward only the options that survived the simulated round, typically one to three.
- Design the human study around the narrowed set. A smaller recruited sample against fewer options costs less than fielding every original variant.
- Run the causal experiment. A controlled discrete choice study against the finalists returns the effect size and confidence interval the earlier simulated round couldn't produce.
- Match the claim to the evidence. Directional findings stay internal; only the confidence-interval-backed result goes into a pricing decision, a public number, or a regulatory filing.
Skipping step 5 for a capital-committing decision is the mistake this framework exists to prevent: treating an unvalidated directional read as a statistically defended one.
Review Subconscious's published replication results and leaderboard, read how the causal experiments are built, or book time to scope a study before committing a budget line to either method.