Why Synthetic Consumer Ratings Need Better Elicitation
Synthetic consumers are easy to generate and hard to trust. Ask a language model for a score on a 1-5 purchase-intent scale and it may overuse the middle, avoid extreme answers, and produce a distribution unlike a human panel. The problem isn't just which model answers the question; the method used to elicit and score the answer can change the result.
Maier and colleagues' October 2025 preprint, involving PyMC Labs researchers, tests Semantic Similarity Rating (SSR). It maps natural-language responses to rating distributions through similarity to reference anchors.
The distinction matters for any team using simulated buyers to test a product, price, or message: a familiar survey format does not guarantee familiar response behavior.
Why direct numerical ratings break down
A prompt that asks a language model for a single rating forces it into a narrow response format, without checking whether the resulting distribution behaves like human survey data.
Direct ratings can cluster around 3 on a 1-5 scale and suppress disagreement, making concepts look more similar than they are. A team may get a clean table of scores while losing the variation needed to rank ideas or spot a weak concept.
SSR separates two tasks:
- The model explains its purchase intent in ordinary language.
- A semantic mapping step compares that answer with reference anchors and assigns probabilities across the rating scale.
What the evaluation found
The evaluation uses 57 personal-care product surveys and 9,300 human responses. Those data support an elicitation comparison for aggregate purchase intent in that setting, rather than realized purchasing or causal price effects.
- A reported reliability ratio reaching 90 percent of the estimated human test-retest level; this compares correlations in survey-level purchase-intent means, rather than the fraction of surveys with exactly matching rankings.
- KS similarity greater than 0.85 for rating distributions. The metric is one minus the Kolmogorov-Smirnov distance, the largest gap between cumulative distributions; it is not a percentage of respondents who agree.
- Wider, more discriminating response patterns than direct numerical prompting
The evaluation also compared GPT-4o with Gemini 2.0 Flash and tested whether age, income, and product category changed the synthetic responses in plausible ways. The method did not require task-specific training data or fine-tuning.
The results support SSR in this evaluated setting relative to direct numerical prompting. Check scorer anchors, model versions and a matched held-out human comparator before extending it to a new category or outcome. See the methods and validation hub.
What a buyer should ask before using the result
An aggregate match can still hide errors that matter to a decision: before using a synthetic panel, inspect the validation at the level where the business choice will be made.
Does the method preserve distributions?
A matching average can conceal variance collapse, so compare the full distribution, not just the mean. Look for missing extremes, excessive neutral responses, and concepts that the synthetic sample fails to distinguish.
Does it preserve rankings?
Early product screening often depends on ordering concepts rather than predicting an exact score. Ranking agreement should be reported separately from distributional similarity because the two measures answer different questions.
Does it work by segment?
Aggregate prediction is usually easier than individual or subgroup simulation. A method can match the total sample while flattening differences by age, income, culture, or category experience; segment claims need their own evaluation.
Does the validation match the intended use?
Purchase intent is not the same as satisfaction, trust, relevance, or actual purchase behavior. A method validated on one response type should not be assumed to transfer to another: new questions, categories, scales, and populations need fresh checks.
A practical research sequence
Synthetic evidence is most useful when it narrows a decision before a more costly commitment. A disciplined sequence is:
- Define the exact product, pricing, messaging, or launch choice.
- Generate natural-language responses from a specified target population.
- Map those responses into a measure with stated reference anchors.
- Compare distributions, rankings, and segment behavior against a human baseline.
- Use the synthetic result to screen scenarios, not to claim certainty.
- Reserve human research for validation where the decision risk warrants it.
The preprint gives the mathematical details. A proposed study should specify what human or behavioral comparison is included and which outcome it validates. Discuss that scope.
The useful lesson is simple: better prompts alone do not make synthetic research reliable. The elicitation method, mapping procedure, and human baseline determine whether the output can support a decision.