Skip to content

When a Synthetic-Customer Read Is Enough, and When You Need a Causal Experiment

A product or research leader has to decide, before engineering capacity or launch spend is committed, whether a synthetic-customer read is sufficient evidence for a feature, positioning, or pricing call, or whether it needs a controlled experiment first. The answer depends on the kind of question being asked, not on how confident the panel output sounds.

Two different questions get asked of the same panel

Synthetic-customer panels answer stated-preference questions: what a persona says it thinks, prefers, would choose, or would pay when shown a description. They do not answer observed-behavior questions: what a real customer does when a purchase, a workaround switch, or a real price is on the table. Prior work on whether simulated respondents built from language models can stand in for survey panels has found the match to human answers holds up better for some question types than others (Argyle et al., "Out of One, Many," on generating synthetic respondent samples with language models).

Treating a directional read from a synthetic panel as a causal answer is the failure mode: a team ships a feature, a message, or a price that never changed real buyer behavior, and only finds out after the engineering time and launch spend are gone. The fix is not to distrust synthetic reads; it's to match the method to the decision's stakes.

Where a directional read is the right first pass

Four situations are well suited to an open-ended synthetic-customer screen.

Pre-launch feature screening

Before committing engineering capacity to a feature, running it past a synthetic-customer panel surfaces whether the persona understands what it is, why it would matter, and how it compares to whatever workaround it already uses. It gives a directional read on whether the feature makes sense and which scoping choices matter, not a go/no-go verdict.

Positioning screening

Different positioning variants shown to the same persona set reveal which framings read as confident versus defensive, plainspoken versus jargon-heavy, on-brand versus off. This is useful triage before a positioning decision gets locked into launch collateral.

Pricing-tier screening

Asking a persona which tier feels right, which feels too cheap, and which feels too expensive produces a categorical signal, useful for choosing among tier structures or feature distributions. It is not precise enough to set a price point.

Segment-level reaction mapping

Running the same launch communication past personas built for each priority segment surfaces which segments respond well and which need different messaging, feeding sales-enablement and customer-success planning.

Running a screen that produces a usable signal

A screening read is only as good as its inputs; four choices determine whether the output is signal or noise.

Build the persona set around actual segments. Generic personas produce generic output. A persona library built from the team's real ICP segments, typically three to seven personas covering the priority segments, with a demographic profile, role context, and relevant attitudes, makes later panels comparable to each other.

Frame the stimulus as a task, not a preference check. A prompt like "do you like this feature" produces low-information output. Prompts that ask a persona to explain the feature in its own words, then name one workflow where it would use it and one where it would not, produce reasoning the team can act on. So does asking which of two options a persona would pick and why, or asking a persona to list its biggest objections before trying a product.

Read the panel as a distribution, not a single vote. One synthetic respondent is one data point. A panel of several personas is a distribution: where reactions cluster, where they diverge, and which segment breaks from the pattern is the useful output.

Decide to ship, kill, or refine, not just record a reaction. Most screening rounds should end in refinement and a second pass, not a binary verdict on the first read.

What a directional read can't tell you

It doesn't extend past what a language model has seen. A genuinely novel product category with no analog in the training data produces extrapolation, not measurement.

It isn't evidence for a regulatory or compliance filing. A claim filed with a regulator needs real human respondents on record, not a simulated read.

It thins out for niche audiences with little public signal. Mainstream consumer and common B2B roles are well represented in a language model's training; narrow roles in small industries are not.

It doesn't capture behavior under real stakes. A persona answering a hypothetical question behaves differently than a real customer facing a real purchase decision, a real time constraint, or a real switching cost.

When the decision needs a controlled experiment

For the launches that move revenue, the pricing changes that touch material spend, and the positioning that defines the brand, a directional screen is not the same evidence as a controlled test. Subconscious runs controlled discrete-choice experiments that compare defined alternatives against a defined population and a defined outcome, and return causal effects with confidence intervals rather than one aggregated reaction. Current results are tracked on the leaderboard.

Where warranted, Subconscious can test or validate studies with real human participants, so a team can move from a simulated experiment to real-human validation without changing the underlying causal question. Audience reach in a simulated experiment and a recruited real-human validation study are distinct claims and should not be blended into one number.

Four-item list: build personas from real ICP segments, frame the stimulus as a task not a preference check, read the panel as a distribution of reactions, end with a decision to ship, kill, or refine.
A screening panel only produces a usable signal when the persona set, the prompt, and the read-out method are deliberate choices, not defaults.

Limitations that carry over either way

A causal experiment answers which alternative changes the outcome and by how much. It does not replace customer discovery interviews, usability observation, or watching how a feature performs once it's in market. Real-human validation, when it's used, confirms a model's structure against real behavior; it does not turn a causal-choice test into an observed usability session, a clinical trial, or automatic proof of market performance.

A branching path: the question splits into a stated-preference route to a synthetic-customer screen, and an observed-behavior route to a controlled experiment. Both converge on ship, kill, or refine.
Whether a synthetic-customer read is sufficient depends on whether the question asks what a persona says versus what a real customer does.

Method choice should follow the decision's stakes, not the other way around (a Wiley review of generating consumer-research samples with large language models, covering the challenges, opportunities, and guidelines involved).