Skip to content

Synthetic Research: When to Trust It, and When to Validate With Real People

Four-row list matching situations to methods: early tests to synthetic alone, capital-committing claims to recruited alone, narrowing to finalists to both in sequence, hard-to-recruit audiences to synthetic alone.
Which method to use depends on what the question is for, not a fixed preference for simulation or recruited research.

A VP of Insights deciding whether to add synthetic research to the toolkit is really deciding which of this quarter's research questions simulation can answer, and which one still needs a recruited human study before capital moves. Get that split wrong: a recruited-research budget gets spent validating a question simulation could have answered in an afternoon, or a pricing, go-to-market, or public claim ships on a directional read that cannot produce a defensible confidence interval, and unravels under scrutiny.

What synthetic research actually is

Synthetic research uses AI-generated personas, conditioned on demographic, psychographic, and behavioral data, to simulate how a target population responds to a stimulus: a concept, a message, a poll question. Researchers assemble personas into panels, run a stimulus against the panel, and get back a distribution of responses with natural-language reasoning for each one.

Three terms often get conflated:

The premise traces to a 2023 finding: conditioning a language model on the detailed background of a real poll respondent produced opinion distributions that tracked actual human responses in benchmark national polls (Argyle et al., a study on prompting language models with individual respondent profiles to reproduce human sample distributions, published in Political Analysis by Cambridge University Press). That result moved the method from academic benchmarking into product, marketing, and insight teams.

Where the method stops being enough

Simulation is directionally useful and limited in predictable ways:

The decision: synthetic alone, recruited alone, or both in sequence

SituationRight method
Early-stage concept, message, or ad-variant testing; competitive scoping; a target audience that's slow or expensive to recruit (senior B2B buyers, niche specialists); privacy-sensitive contextsSynthetic alone
A single, capital-committing go-to-market or pricing decision; a quantitative claim for external publication or PR; a regulatory submission or legal evidenceRecruited alone
Narrowing many options down to the 1 to 3 that matter before a final callSynthetic first, recruited second

The third row is the pattern most teams underuse. Run synthetic research first to explore the landscape, test variations, and refine the research instrument. Then field a smaller, targeted study with recruited participants against only the finalists. That sequencing lowers recruitment cost because the human study only tests survivors, and raises confidence in the final number because the questions were already stripped of obvious flaws before a real person answered them.

Where Subconscious sits in that sequence

Subconscious runs controlled discrete choice experiments, using causal DCE, Mixed Logit, and ICLV methodology, that return causal effects with confidence intervals rather than a directional read. That's the rigor layer synthetic-alone methods can't supply: a defensible population estimate, with error bars, for the moment a finding needs to survive audit or public scrutiny.

The practical advantage: a team can move from a simulated experiment to a real-human study without changing the underlying causal question, so the narrow-with-simulation, confirm-with-people sequence runs on the same experimental design, rather than handing the finalists to a separate agency running a different methodology.

The fit is bounded, too: Subconscious's positioning is the validation and causal-inference layer for decisions that carry capital risk, not a claim to replace every recruited-human study with simulation. A regulatory submission, a legal filing, or a public percentage claim still needs the sourced, audited number a controlled study produces.

Running the sequence without wasting the human study

  1. Define the decision, not just the topic. Name the population, the alternatives being compared, and the outcome that decides the action, not just "test the messaging."
  2. Run the simulated pass first. Test the full set of variants, including the ones expected to lose, so the elimination is evidence-based rather than assumed.
  3. Narrow to the finalists. Carry forward only the options that survived the simulated round, typically one to three.
  4. Design the human study around the narrowed set. A smaller recruited sample against fewer options costs less than fielding every original variant.
  5. Run the causal experiment. A controlled discrete choice study against the finalists returns the effect size and confidence interval the earlier simulated round couldn't produce.
  6. Match the claim to the evidence. Directional findings stay internal; only the confidence-interval-backed result goes into a pricing decision, a public number, or a regulatory filing.

Skipping step 5 for a capital-committing decision is the mistake this framework exists to prevent: treating an unvalidated directional read as a statistically defended one.

Review Subconscious's published replication results and leaderboard, read how the causal experiments are built, or book time to scope a study before committing a budget line to either method.