Do LLMs Understand Real-World Prices? A Pricing Benchmark for Synthetic Consumers
An LLM-based synthetic panel can produce fluent, plausible-sounding survey answers without ever grounding those answers in real prices. Before a pricing or packaging decision leans on that kind of panel, the panel's price reasoning needs its own check.
The decision this bears on
A brand or insights team weighing an LLM-persona panel against a calibrated, validated experimental design is really deciding how much price-sensitive risk to accept. An uncalibrated panel that sounds confident about cost can still produce systematically biased demand or willingness-to-pay estimates. Shipping a pricing decision on that basis risks wasted research spend and a go-to-market call built on a number the panel never actually understood.
How does this benchmark test price reasoning?
A recent benchmark tested this directly by adapting a classic televised pricing-showcase game into a structured evaluation for large language models. Two models see the same showcase of consumer packaged goods, mostly everyday items like toothpaste and snack bars with a combined retail value around $20. Each model bids the total retail value; the closest bid without going over wins, and if both models overbid, neither wins.
Each round follows the same five steps: generate a random showcase of three items, hand each model ten example prices from similar products as calibration reference, send the identical prompt to two models, parse the bid and its rationale from a strict JSON response, then compare both bids to the actual retail price to score the round.
Naming this failure mode up front is what lets a buyer test each capability on its own. The design targets three related capabilities: estimating a real-world price, using reference examples to calibrate that estimate, and deciding how aggressively to act on it, though a single bid does not cleanly separate estimation from shading.
Accuracy and restraint are not the same skill
Dozens of models entered the tournament; two preliminary rounds narrowed the field to a small group of finalists competing across fifty rounds. The best-performing model in the finals estimated prices with an error noticeably below the error rate of human contestants on the original televised game, though the comparison is confounded: models bid on roughly $20 CPG showcases while contestants bid on showcases worth tens of thousands under real stakes; the weakest model's price estimates were only marginally better than random guessing, with almost no relationship between its bids and the actual prices.
Accuracy alone did not decide the outcome. The game penalizes overbidding by even a cent with automatic disqualification, so rational play means bidding a lower quantile of the estimate's distribution, not a small shave off a point estimate. Human contestants on the original show overbid about 25% of the time, a descriptive rate under real stakes and prior-round selection, not a normative benchmark. Among the finalist models, the most conservative overbid only 2% of the time, while the most aggressive overbid 72% of the time. The model with the single best price accuracy in the field still overbid well above the human benchmark, which cost it standing on the overall leaderboard.
Models varied in how well they used the ten reference prices they were given: models that incorporated the calibration examples tended to produce better-calibrated bids than models that largely ignored them, an observational pattern that does not isolate prompting effects from underlying model capability, though it echoes recent controlled research on prompting strategy. A controlled study of prompting approaches for choice-modelling tasks found that how calibration guidance and reference information get structured in a prompt materially changes the quality of a model's output, not just which model runs the prompt (Can large language models assist choice modelling? Insights into prompting strategies and current models' capabilities, arXiv, 2026-07-28).
"Our findings reveal that proprietary LLMs can generate valid and behaviourally sound utility specifications, particularly when guided by structured prompts (CoT)."
Sfeir, Nova, Hess, and van Cranenburgh, Journal of Choice Modelling (source)
What this means for evaluating a synthetic-consumer method
The historical benchmark numbers above describe one study's models and prompts at one point in time. They are not a live leaderboard and not a claim about any specific vendor's current product. Read them as evidence that price reasoning splits into distinct, testable parts, not as a ranking to shop from.
A synthetic-consumer method built around raw LLM roleplay inherits whatever calibration and restraint its base model happens to have on a given day. A method built around calibrated experimental design paired with validation against real human data tests those parts on purpose: it asks whether an estimate is grounded, whether the model uses reference information to correct itself, and whether its resulting choice behavior holds up before it drives a pricing or packaging decision. Subconscious's causal action-testing and discrete-choice experiment design is built around that kind of calibration and validated design rather than raw model access.
What are the limits of this benchmark?
A number without its limits is marketing. This benchmark measures price estimation and bidding restraint in one bidding game. It does not measure willingness-to-pay in a real purchase context, brand-specific price perception, or how a synthetic panel performs on a causal question outside pricing. A team evaluating a synthetic-consumer method for its own pricing or packaging question still needs evidence specific to that question and, when the decision is large enough, a way to check a simulated result against real-human testing. Subconscious can test or validate studies with real human participants, and a team can move from a simulated experiment to that kind of validation without changing the underlying causal question.
Next step
Teams comparing synthetic-consumer methods for a pricing decision can see current published comparisons on the leaderboard or read about the validated experiment workflow on how we work.