AI Research vs Real Users: A PM Decision Framework
A product manager chooses a research method for a specific decision: a feature bet, a price change, a message, or a positioning call. The method should match the evidence that decision requires.
Use a simulated behavioral experiment when the question is about causal choices across defined options and the result will be treated as directional evidence. Use real participants when the decision depends on actual product use, spending, lived context, or proof that a person understood something. Sequence both when a simulated finding will drive a consequential action and needs a study-specific check against human responses.
Choosing poorly creates two risks: shipping on a simulated signal that was never checked against real behavior, or committing to real-user fieldwork for a reversible, low-stakes call that did not require it. Two to four weeks is a historical planning assumption some teams have used for recruiting and scheduling, not a current service commitment.
The core tradeoff
A controlled behavioral experiment on a simulated population suits breadth: testing many concepts, framings, or messages before committing budget to one. A simulated respondent does not carry a real budget, a real procurement process, or years of lived context behind a workflow.
Real-user research, including interviews, usability sessions, and surveys, is grounded in responses and behavior nobody had to model. It can capture physical interaction with a product and the weight of a real financial or emotional decision.
When simulated testing is the right call
Early-stage concept exploration. A team narrowing five feature ideas to two can use a controlled behavioral experiment to compare all five under one defined causal question. The result can reduce the option set, but it does not prove how people will use the final feature.
Reversible decisions with bounded consequences. A feature flag, message, or concept suits simulated testing when the team can reverse the action and will treat the result as directional, not conclusive.
Sprint-speed calls. A prioritization decision due before a fixed deadline, with no time to recruit and schedule participants, gets a directional read from a simulated session it would not otherwise have in time.
Copy and messaging testing. Which value proposition changes preference, or which feature name is clearer, can be framed as explicit choices for a simulated population. The output is comparative evidence, not proof of market performance.
Pre-validation before a human study. A simulated test can identify which concepts deserve a real-human check. The human study still carries the burden of validating the selected causal question.
Competitive positioning. Modeling how buyers choose between two product descriptions across several framings can reveal which causal contrasts deserve further validation.
When you need real users
High-stakes pricing decisions. Willingness to pay depends on a real budget and a real pain of paying. A simulated population can model price sensitivity directionally, but a revenue-affecting price call needs a check against people who are actually spending money, and a defensible quantitative number needs a properly sized sample, not a handful of respondents (NN/g: Quantitative User Research: Study Guide).
Usability testing with complex interactions. Observing someone click through a multi-step workflow and hit real edge cases requires a real person in front of a real prototype. Simulated testing can evaluate a described flow, not the physical and cognitive experience of using the software. The reverse holds for open-ended usability problems: a small qualitative sample, as few as five participants, can already surface most of a workflow's usability issues, even though that same small sample would not support a quantitative claim (NN/g: Why 5 Participants Are Okay in a Qualitative Study, but Not in a Quantitative One).
Emotional and behavioral nuance. Whether someone trusts a feature with sensitive data, or how they feel about a change to a workflow they have used for years, involves personal context a model approximates rather than replicates.
Regulatory or compliance validation. Proof that users understood a consent flow or disclosure requires documented human evidence. Simulated research does not satisfy that bar.
Checking a simulated finding before it drives a decision. Periodically running the same question through both methods and comparing results is how a team learns which questions its simulated signal answers reliably.
| Decision dimension | Simulated behavioral experiment | Real-user research |
|---|---|---|
| Who or what responds | A simulated population under controlled conditions | Recruited people responding from lived context |
| Best use | Comparing concepts, messages, choices, or framings | Observing use, probing context, or confirming a consequential decision |
| Financial stakes | Directional because no real spend or procurement occurs | Can examine decisions made with an actual budget |
| Interaction complexity | Evaluates defined choices and described experiences | Can observe live, multi-step product interaction |
| Compliance evidence | Does not prove that a person understood a disclosure | Can produce documented evidence from real participants |
Sequencing both without redesigning the study
Subconscious can run a controlled behavioral experiment on a simulated population, then check the same causal question with real human participants without redesigning the study. The human step is a study-specific replication check, not a general endorsement of every simulated result.
A funnel that uses both stages this way: run a simulated experiment across 10 concepts to narrow to 3, then run 5 real-user interviews on those 3 to narrow to 1, then commit a full usability study to the winner. Each stage filters, so recruited-participant time goes only to what already cleared simulated screening.
The public replication metric behind that validation step is 93% replication accuracy: how often a simulated study reproduces the direction and outcome of the original human study, measured across 350+ published human studies in 20+ domains (go.subconscious.ai/paper). That figure describes the "does simulated signal match a real-human study" step specifically. It is not a claim about every possible research question.
Four questions before choosing a method
- How reversible is this decision? Easily reversible, such as copy or a feature flag, favors simulated testing. Hard to reverse, such as pricing, core architecture, or brand positioning, needs a real-user check.
- What does the deadline allow without lowering the evidence standard? Time available before deciding is a planning constraint, not evidence that one method is valid. If the decision requires observed human behavior and that evidence is unavailable, narrow, postpone, or mark the decision unresolved.
- Does the decision touch money or emotion? If people are paying for something, or the change reaches a workflow they rely on personally, lean toward real users.
- Are you exploring or confirming? Exploring a wide option set favors simulated testing. Confirming a final call before it ships favors real users.
What the validation step cannot settle
Real-human validation, run this way, is a study-specific replication check on one causal question. It is not open-ended usability testing of a live prototype, and it does not walk a person through a multi-step UI. A simulated population built from broad audience data is a targeting and simulation input, not a recruitable panel for interviews. Confidence intervals, segment breakdowns, and decision-memo outputs are specific to each study and need confirming per engagement, not treated as a standing guarantee.
Make confidence question-specific
One historical planning example is to compare simulated findings with existing human research for two to three sprints, recording where the methods agree and diverge. That cadence is an inherited planning benchmark, not a standing delivery commitment. Trust in the simulated signal should be earned for the specific question, not assumed globally.
Put the evidence burden on the decision
Do not use simulated testing merely because a deadline is close. Use it when the causal question, action, and evidence standard fit the method. When actual spending, physical interaction, compliance, or lived emotional context determines the answer, collect human evidence or treat the decision as unresolved.
See how one causal question can move from simulation to real-human validation, or bring a live decision to a working session to identify which evidence it requires.