Skip to content
Subconscious

Why Synthetic Consumer Ratings Need Better Elicitation

Synthetic consumers are easy to generate and hard to trust. Ask a language model for a score on a 1-5 purchase-intent scale and it may overuse the middle, avoid extreme answers, and produce a distribution unlike a human panel. The problem isn't just which model answers the question; the method used to elicit and score the answer can change the result.

Maier and colleagues' October 2025 preprint, involving PyMC Labs researchers, tests Semantic Similarity Rating (SSR). It maps natural-language responses to rating distributions through similarity to reference anchors.

The distinction matters for any team using simulated buyers to test a product, price, or message: a familiar survey format does not guarantee familiar response behavior.

Four steps: model gives a natural-language purchase-intent answer, it's compared with reference statements per scale point, that becomes a probability distribution, then checked against a human survey baseline.
Separating what the model says from how it gets scored is what makes the synthetic rating checkable against a human baseline.

Why direct numerical ratings break down

A prompt that asks a language model for a single rating forces it into a narrow response format, without checking whether the resulting distribution behaves like human survey data.

Direct ratings can cluster around 3 on a 1-5 scale and suppress disagreement, making concepts look more similar than they are. A team may get a clean table of scores while losing the variation needed to rank ideas or spot a weak concept.

SSR separates two tasks:

What the evaluation found

The evaluation uses 57 personal-care product surveys and 9,300 human responses. Those data support an elicitation comparison for aggregate purchase intent in that setting, rather than realized purchasing or causal price effects.

The evaluation also compared GPT-4o with Gemini 2.0 Flash and tested whether age, income, and product category changed the synthetic responses in plausible ways. The method did not require task-specific training data or fine-tuning.

The results support SSR in this evaluated setting relative to direct numerical prompting. Check scorer anchors, model versions and a matched held-out human comparator before extending it to a new category or outcome. See the methods and validation hub.

What a buyer should ask before using the result

An aggregate match can still hide errors that matter to a decision: before using a synthetic panel, inspect the validation at the level where the business choice will be made.

Does the method preserve distributions?

A matching average can conceal variance collapse, so compare the full distribution, not just the mean. Look for missing extremes, excessive neutral responses, and concepts that the synthetic sample fails to distinguish.

Does it preserve rankings?

Early product screening often depends on ordering concepts rather than predicting an exact score. Ranking agreement should be reported separately from distributional similarity because the two measures answer different questions.

Does it work by segment?

Aggregate prediction is usually easier than individual or subgroup simulation. A method can match the total sample while flattening differences by age, income, culture, or category experience; segment claims need their own evaluation.

Does the validation match the intended use?

Purchase intent is not the same as satisfaction, trust, relevance, or actual purchase behavior. A method validated on one response type should not be assumed to transfer to another: new questions, categories, scales, and populations need fresh checks.

Four-item checklist: distributions preserved, rankings preserved, works by segment, matches the intended question type.
A matching average can still hide collapsed variance, flattened segments, or validation done on the wrong question type.

A practical research sequence

Synthetic evidence is most useful when it narrows a decision before a more costly commitment. A disciplined sequence is:

The preprint gives the mathematical details. A proposed study should specify what human or behavioral comparison is included and which outcome it validates. Discuss that scope.

The useful lesson is simple: better prompts alone do not make synthetic research reliable. The elicitation method, mapping procedure, and human baseline determine whether the output can support a decision.