Skip to content

Directional Read or Decision-Grade Evidence? A Buyer's Test Before You Spend on Synthetic Research

If a synthetic panel gives you a plausible-sounding read on a campaign, a price, or a message, the question that matters before you commit budget is not which platform produced it. It is whether that read is a directional guess or a causal effect you can defend. Commit spend on the guess and the failure shows up later, as a mispriced offer or a message that tested well but flopped with real buyers.

Why the category exists

Traditional human research is slow and expensive in exactly the places a marketing or insights leader needs to move fast. A focus group typically runs 3 to 4 weeks from brief to readout (historical planning example). Recruiting a hard-to-reach segment, such as senior buyers or regulated professionals, can run into the thousands of dollars per study (historical planning example). A team limited to a handful of human sessions gets 5 to 10 questions answered; a faster method can ask hundreds in an afternoon (historical planning example). That speed and reach gap is real. It is why synthetic panels became a normal part of the research stack.

Speed does not make a read decision-grade. It makes it fast to obtain and easy to over-trust.

The two things a "read" can mean

Before shortlisting any research method, separate two outputs that look similar on a slide:

A buyer who cannot tell which one is on their desk cannot tell how much budget risk it is safe to attach to it.

What a synthetic read is weak at

External validity research on discrete choice and vignette methods is the relevant evidence here, not vendor marketing. Controlled comparisons between stated-preference experiments and real-world behavior find that these methods track actual choices reasonably well on routine, low-stakes decisions, but the fit degrades on unfamiliar or high-stakes choices (Hainmueller et al., PNAS). A systematic review of discrete choice experiments against real health-related choices reports similar external-validity gaps depending on the choice context (Springer, The European Journal of Health Economics). The pattern holds for synthetic panels built on top of these methods: directional accuracy is highest on routine, low-precedent-free questions, and weakest on emotionally charged behavior, novel markets, and decisions driven by social contagion, where there is no calibration data for the model to lean on.

That is not a reason to avoid synthetic methods. It is a reason not to treat their output as a finished answer on anything where being wrong is expensive.

Four questions to ask before you trust the number

  1. Is this a score or an effect? A single plausible number is a signal. An effect with a confidence interval is evidence. Ask which one you are looking at.
  2. What decision is riding on it? Early concept exploration and message iteration can run on directional signal alone. A pricing change, a launch decision, or a claim going into a regulated channel should not.
  3. Can the method be checked against real people without re-asking the question? A study designed to run as a synthetic experiment and be checked against recruited human respondents, without changing what is being tested, is a materially different guarantee than a synthetic-only read with no path to verification.
  4. Does the source publish where the method fails, not just where it works? A published failure mode (emotionally charged behavior, novel markets, network effects) is a sign the benchmark is honest. A claim with no stated limitation usually has one it isn't disclosing.
Directional signalDecision-grade evidence
What it answers"What might this segment say?""How much does this action move the outcome, and how sure are we?"
Typical outputA score or a summarized responseAn effect estimate with a confidence interval
Right useEarly exploration, message iteration, hypothesis generationSpend decisions: pricing, launch, positioning, claims
Verification pathNone, or informalCheckable against real recruited respondents on the same question
Failure mode if misusedWasted iteration cyclesCommitted budget against a result that does not hold with real customers

Where a causal experiment changes the answer

Subconscious runs controlled discrete choice experiments, including McFadden discrete choice, mixed logit, and integrated choice and latent variable models, and returns a causal effect estimate with a confidence interval rather than a single predicted score. That answers question 1 directly. Subconscious can also test or validate studies with real human participants, which answers question 3: a team can move from a simulated experiment to a real-human check without changing the causal question being asked. Recruiting real respondents for that check is distinct from the scale of any audience model behind the simulation; keep those two claims separate when you evaluate a research method's pitch.

This matters most exactly where the PNAS and European Journal of Health Economics findings above say synthetic-only methods are weakest: unfamiliar, high-stakes, or emotionally loaded decisions. A confidence interval tells you how much to trust a single run. A human check tells you whether the causal question still holds outside the simulation.

Where this does not change the answer

Not every decision needs this level of rigor. If the question is whether a headline direction feels closer to what your last three campaigns tested, a directional synthetic read answers it fine, and a full causal experiment with human validation is overkill. Reserve the four-question test above for decisions where a wrong answer costs more than the research would have. See how the underlying method works in more detail in our research documentation and how a study moves from design to a validated result in how a Subconscious study runs.

What this does not settle

Run the test on your next decision

Before committing budget on a synthetic read, apply the four-question test above to the last panel result you were handed. If it fails on question 1 or 3, that is the signal to ask for a causal effect estimate with a confidence interval, and a human check, before the spend goes out the door. See recent case evidence for what that verification path looks like end to end, or book a walkthrough of a controlled experiment on your own decision.

Two columns. Directional Signal: single score, no interval, no verification, for early exploration. Decision-Grade Evidence: effect with confidence interval, checkable against real respondents, for pricing and launch.
A directional signal answers what a segment might say; decision-grade evidence answers how much an action moves the outcome and how sure you can be.