Skip to content

AI Research vs Real Users: A PM Decision Framework

A product manager chooses a research method for a specific decision: a feature bet, a price change, a message, or a positioning call. The method should match the evidence that decision requires.

Use a simulated behavioral experiment when the question is about causal choices across defined options and the result will be treated as directional evidence. Use real participants when the decision depends on actual product use, spending, lived context, or proof that a person understood something. Sequence both when a simulated finding will drive a consequential action and needs a study-specific check against human responses.

Choosing poorly creates two risks: shipping on a simulated signal that was never checked against real behavior, or committing to real-user fieldwork for a reversible, low-stakes call that did not require it. Two to four weeks is a historical planning assumption some teams have used for recruiting and scheduling, not a current service commitment.

The core tradeoff

A controlled behavioral experiment on a simulated population suits breadth: testing many concepts, framings, or messages before committing budget to one. A simulated respondent does not carry a real budget, a real procurement process, or years of lived context behind a workflow.

Real-user research, including interviews, usability sessions, and surveys, is grounded in responses and behavior nobody had to model. It can capture physical interaction with a product and the weight of a real financial or emotional decision.

When simulated testing is the right call

Early-stage concept exploration. A team narrowing five feature ideas to two can use a controlled behavioral experiment to compare all five under one defined causal question. The result can reduce the option set, but it does not prove how people will use the final feature.

Reversible decisions with bounded consequences. A feature flag, message, or concept suits simulated testing when the team can reverse the action and will treat the result as directional, not conclusive.

Sprint-speed calls. A prioritization decision due before a fixed deadline, with no time to recruit and schedule participants, gets a directional read from a simulated session it would not otherwise have in time.

Copy and messaging testing. Which value proposition changes preference, or which feature name is clearer, can be framed as explicit choices for a simulated population. The output is comparative evidence, not proof of market performance.

Pre-validation before a human study. A simulated test can identify which concepts deserve a real-human check. The human study still carries the burden of validating the selected causal question.

Competitive positioning. Modeling how buyers choose between two product descriptions across several framings can reveal which causal contrasts deserve further validation.

When you need real users

High-stakes pricing decisions. Willingness to pay depends on a real budget and a real pain of paying. A simulated population can model price sensitivity directionally, but a revenue-affecting price call needs a check against people who are actually spending money, and a defensible quantitative number needs a properly sized sample, not a handful of respondents (NN/g: Quantitative User Research: Study Guide).

Usability testing with complex interactions. Observing someone click through a multi-step workflow and hit real edge cases requires a real person in front of a real prototype. Simulated testing can evaluate a described flow, not the physical and cognitive experience of using the software. The reverse holds for open-ended usability problems: a small qualitative sample, as few as five participants, can already surface most of a workflow's usability issues, even though that same small sample would not support a quantitative claim (NN/g: Why 5 Participants Are Okay in a Qualitative Study, but Not in a Quantitative One).

Emotional and behavioral nuance. Whether someone trusts a feature with sensitive data, or how they feel about a change to a workflow they have used for years, involves personal context a model approximates rather than replicates.

Regulatory or compliance validation. Proof that users understood a consent flow or disclosure requires documented human evidence. Simulated research does not satisfy that bar.

Checking a simulated finding before it drives a decision. Periodically running the same question through both methods and comparing results is how a team learns which questions its simulated signal answers reliably.

Decision dimensionSimulated behavioral experimentReal-user research
Who or what respondsA simulated population under controlled conditionsRecruited people responding from lived context
Best useComparing concepts, messages, choices, or framingsObserving use, probing context, or confirming a consequential decision
Financial stakesDirectional because no real spend or procurement occursCan examine decisions made with an actual budget
Interaction complexityEvaluates defined choices and described experiencesCan observe live, multi-step product interaction
Compliance evidenceDoes not prove that a person understood a disclosureCan produce documented evidence from real participants

Sequencing both without redesigning the study

Subconscious can run a controlled behavioral experiment on a simulated population, then check the same causal question with real human participants without redesigning the study. The human step is a study-specific replication check, not a general endorsement of every simulated result.

A funnel that uses both stages this way: run a simulated experiment across 10 concepts to narrow to 3, then run 5 real-user interviews on those 3 to narrow to 1, then commit a full usability study to the winner. Each stage filters, so recruited-participant time goes only to what already cleared simulated screening.

The public replication metric behind that validation step is 93% replication accuracy: how often a simulated study reproduces the direction and outcome of the original human study, measured across 350+ published human studies in 20+ domains (go.subconscious.ai/paper). That figure describes the "does simulated signal match a real-human study" step specifically. It is not a claim about every possible research question.

Four questions before choosing a method

  1. How reversible is this decision? Easily reversible, such as copy or a feature flag, favors simulated testing. Hard to reverse, such as pricing, core architecture, or brand positioning, needs a real-user check.
  2. What does the deadline allow without lowering the evidence standard? Time available before deciding is a planning constraint, not evidence that one method is valid. If the decision requires observed human behavior and that evidence is unavailable, narrow, postpone, or mark the decision unresolved.
  3. Does the decision touch money or emotion? If people are paying for something, or the change reaches a workflow they rely on personally, lean toward real users.
  4. Are you exploring or confirming? Exploring a wide option set favors simulated testing. Confirming a final call before it ships favors real users.
A decision path: name the decision, then branch to a simulated experiment for reversible directional calls, real-user research for real spend or lived context, or a check against both for consequential findings.
Route the decision by what it depends on: directional and reversible goes to simulated testing, real spend or lived context goes to real users, consequential findings get checked against both.

What the validation step cannot settle

Real-human validation, run this way, is a study-specific replication check on one causal question. It is not open-ended usability testing of a live prototype, and it does not walk a person through a multi-step UI. A simulated population built from broad audience data is a targeting and simulation input, not a recruitable panel for interviews. Confidence intervals, segment breakdowns, and decision-memo outputs are specific to each study and need confirming per engagement, not treated as a standing guarantee.

Make confidence question-specific

One historical planning example is to compare simulated findings with existing human research for two to three sprints, recording where the methods agree and diverge. That cadence is an inherited planning benchmark, not a standing delivery commitment. Trust in the simulated signal should be earned for the specific question, not assumed globally.

Put the evidence burden on the decision

Do not use simulated testing merely because a deadline is close. Use it when the causal question, action, and evidence standard fit the method. When actual spending, physical interaction, compliance, or lived emotional context determines the answer, collect human evidence or treat the decision as unresolved.

See how one causal question can move from simulation to real-human validation, or bring a live decision to a working session to identify which evidence it requires.