How CPG Teams Test a Price Increase Before a Shelf Reset
A revenue growth management or brand lead facing a retailer's category reset in six to twelve weeks needs to know whether a price increase on an already-listed SKU will hold, and the honest answer is that scanner history alone cannot tell them, because scanner history only reports prices that were actually charged, not the higher price under consideration. Household panel and point-of-sale data describe a market that existed; the shelf reset question is about a price that has never existed on that shelf. A randomized choice experiment can answer it, because the price a shopper sees is assigned by the experiment rather than pulled from history that was mostly generated by promotions.
- A price increase on a listed SKU should be tested with a randomized choice experiment, not scanner-derived elasticity, because scanner data reflects prices that were charged - largely promotional depths - not the new shelf price being proposed.
- Scanner-based elasticity is confounded with the display, feature, and seasonality that traveled alongside historical promotions, and the same categories can flip between "elastic" and "inelastic" depending on which data source produced the estimate.
- Gabor-Granger and Van Westendorp survey backstops ask a shopper to name a number with no budget and no tradeoff; stated willingness to pay runs an average of 21% above real willingness to pay (Schmidt and Bijmolt, Journal of the Academy of Marketing Science 48(3), 2020: https://research.rug.nl/en/publications/accurately-measuring-willingness-to-pay-for-consumer-goods-a-meta/).
- A randomized experiment assigns the price against the real, upcoming competitive set on the shelf, producing an estimate of the effect of that specific increase, plus the cross-price effect on the rest of the category and category dollars rather than brand volume alone.
- Every effect should ship with a confidence interval and a clear statement of which price points the design actually tested.
What decision does a CPG team need to test before a shelf reset?
The decision is narrow and consequential: whether to take a list price increase on a SKU that already has distribution, how large an increase, and whether to pair it with a pack-size change, all six to twelve weeks before the retailer's line review. The risk is not symmetric. Hold price and margin erodes quarter over quarter. Push the increase past what the category buyer will tolerate, and the SKU loses facings or gets delisted at the review - and winning that shelf space back typically takes a full planning cycle, not a quarter. That asymmetry is why the test has to happen before the reset, not after the retailer has already reacted.
Why the scanner data everyone reaches for can't answer this question
Retail scanner history and household panel data are records of what happened, not experiments about what would happen. The price variation embedded in that history is overwhelmingly promotional - temporary price cuts, feature and display activity - not permanent shelf-price movement, and the promotional periods carry their own lift from display and feature that gets folded into the elasticity estimate along with the price effect itself. An RGM team estimating elasticity from that history is really estimating the combined effect of a price cut plus a promotional event, and applying it to a permanent list price increase the SKU has never carried. The estimate answers a different question than the one the line review is going to ask.
Why does historical elasticity break down right when you need it most?
It breaks down because the estimate is fragile to the data source, not just the time period. Luke, Tonsor, and Schroeder's meat demand elasticity research (Agricultural and Resource Economics Review 55(1), 2026) finds the same meat categories flip between elastic and inelastic, and between luxury and necessity classification, depending on whether the analyst used public aggregate data or retail scanner data . A number that reverses sign depending on which dataset produced it is not a stable input for a decision with a year-long downside if it's wrong. The category buyer at the line review does not care which data source generated the RGM team's estimate; they care whether the shelf still sells at the new price.
The survey backstop carries its own bias
Most RGM teams pair scanner elasticity with a Gabor-Granger or Van Westendorp survey as a sanity check, and that backstop has a known, measured problem. Schmidt and Bijmolt's meta-analysis, spanning 77 studies in 47 papers with 24,347 hypothetical and 20,656 real willingness-to-pay observations, finds that stated willingness to pay overstates real willingness to pay by an average of 21% (Journal of the Academy of Marketing Science 48(3), 2020: https://research.rug.nl/en/publications/accurately-measuring-willingness-to-pay-for-consumer-goods-a-meta/). Asking a shopper to name a price with no budget constraint and no competing option in front of them produces a number that is systematically too generous. A team that anchors a price increase on that number is building in the hypothetical bias from the start, not correcting for it.
What does a randomized shelf-reset price experiment look like?
It looks like a discrete choice experiment where the price a respondent sees is randomly assigned rather than observed from history, run against the real, upcoming competitive set - the actual competitor SKUs, prices, and pack sizes expected to be on that shelf after the reset, not the current lineup. Because the assignment is random, the resulting effect is an estimate of what the price increase causes, not a correlation carried over from whatever price variation happened to occur historically. Randomized experiments analyzed with discrete choice models - McFadden discrete choice, Mixed Logit, and ICLV - are how the effect gets estimated once the manipulation is in place; the identification comes from the randomization in the design, not from the estimator itself. A flat multinomial logit model carries the IIA assumption, that a shopper's relative odds between two SKUs don't shift when a third option enters or leaves the set; Mixed Logit relaxes that assumption by letting preferences vary across the simulated population, which matters directly at a shelf reset because the competitive set itself is changing.
Conventional approach and randomized experiment, side by side
| Scanner elasticity + survey backstop | Randomized choice experiment | |
|---|---|---|
| Price variation used | Historical promotional depth, mostly | Assigned prices, including the increase never charged |
| Cross-price effect on category | Not directly estimable; requires separate models per SKU | Estimated directly against the real competitive set |
| Confound risk | Display, feature, and seasonality travel with historical price cuts | None from history; confounds are controlled by randomization |
| Willingness-to-pay bias | Survey backstop overstates by an average of 21% (Schmidt and Bijmolt, 2020) | Choice-based, no open-ended price question |
| Output granularity | Brand volume | Category dollars and brand volume, with confidence intervals |
| Best for: | Teams with only historical data and no time before the reset | Teams with six to twelve weeks to test a specific increase before the line review |
What does the retailer actually want to see at the line review?
The category buyer wants category dollars and the cross-price effect on the rest of the set, not just the effect on the RGM team's own brand. A randomized experiment against the real competitive set delivers both directly, because every simulated shopper chooses among all the SKUs the design includes, not just the client's own line. That is the number a category buyer can act on: not "our brand's volume at the new price," but "what happens to the category's dollars, and which competitor picks up the share that moves." Report each effect with its confidence interval, understanding that the interval covers the effect within the simulated population and design tested, not the entire real market unconditionally, and state plainly which price points the design actually covered so the buyer knows the range the result speaks to.
How much can you trust a randomized price experiment?
Trust it to the extent the validation record supports, and no further, which is why the record should always ship with its denominator rather than a bare percentage. On the Hainmueller immigration conjoint study, our best configuration reaches 0.832 rank correlation against the published human result, where two independent samples of real humans reach 0.959 - 87% of that measured human ceiling, on one study, at the best configuration. Across all 43 published randomized studies that pass design filters, the mean is 0.73 (https://fidelity.subconscious.ai/papers/causal-fidelity/causal_fidelity_paper.pdf). Published studies can sit in a model's training data, which is a real limitation the replication protocol behind that paper is built to address, not a problem that disappears by asserting validation. The full leaderboard, updated as new replications run, is public at leaderboard; the methodology behind the fidelity numbers and how they're built is covered on the methods and validation hub.
Before the next line review, pull the actual competitive set expected on the reset shelf - competitor SKUs, prices, and pack sizes - and write down the specific price points under consideration for the increase; that set is the input a randomized experiment needs, and having it in hand is the first concrete step regardless of which testing method gets used. For a walkthrough of how a shelf-reset design gets built against a specific category, book time on /meet.