Skip to content

How CPG Teams Test a Price Increase Before a Shelf Reset

A revenue growth management or brand lead facing a retailer's category reset in six to twelve weeks needs to know whether a price increase on an already-listed SKU will hold, and the honest answer is that scanner history alone cannot tell them, because scanner history only reports prices that were actually charged, not the higher price under consideration. Household panel and point-of-sale data describe a market that existed; the shelf reset question is about a price that has never existed on that shelf. A randomized choice experiment can answer it, because the price a shopper sees is assigned by the experiment rather than pulled from history that was mostly generated by promotions.

What decision does a CPG team need to test before a shelf reset?

The decision is narrow and consequential: whether to take a list price increase on a SKU that already has distribution, how large an increase, and whether to pair it with a pack-size change, all six to twelve weeks before the retailer's line review. The risk is not symmetric. Hold price and margin erodes quarter over quarter. Push the increase past what the category buyer will tolerate, and the SKU loses facings or gets delisted at the review - and winning that shelf space back typically takes a full planning cycle, not a quarter. That asymmetry is why the test has to happen before the reset, not after the retailer has already reacted.

Why the scanner data everyone reaches for can't answer this question

Retail scanner history and household panel data are records of what happened, not experiments about what would happen. The price variation embedded in that history is overwhelmingly promotional - temporary price cuts, feature and display activity - not permanent shelf-price movement, and the promotional periods carry their own lift from display and feature that gets folded into the elasticity estimate along with the price effect itself. An RGM team estimating elasticity from that history is really estimating the combined effect of a price cut plus a promotional event, and applying it to a permanent list price increase the SKU has never carried. The estimate answers a different question than the one the line review is going to ask.

Why does historical elasticity break down right when you need it most?

It breaks down because the estimate is fragile to the data source, not just the time period. Luke, Tonsor, and Schroeder's meat demand elasticity research (Agricultural and Resource Economics Review 55(1), 2026) finds the same meat categories flip between elastic and inelastic, and between luxury and necessity classification, depending on whether the analyst used public aggregate data or retail scanner data . A number that reverses sign depending on which dataset produced it is not a stable input for a decision with a year-long downside if it's wrong. The category buyer at the line review does not care which data source generated the RGM team's estimate; they care whether the shelf still sells at the new price.

The survey backstop carries its own bias

Most RGM teams pair scanner elasticity with a Gabor-Granger or Van Westendorp survey as a sanity check, and that backstop has a known, measured problem. Schmidt and Bijmolt's meta-analysis, spanning 77 studies in 47 papers with 24,347 hypothetical and 20,656 real willingness-to-pay observations, finds that stated willingness to pay overstates real willingness to pay by an average of 21% (Journal of the Academy of Marketing Science 48(3), 2020: https://research.rug.nl/en/publications/accurately-measuring-willingness-to-pay-for-consumer-goods-a-meta/). Asking a shopper to name a price with no budget constraint and no competing option in front of them produces a number that is systematically too generous. A team that anchors a price increase on that number is building in the hypothetical bias from the start, not correcting for it.

What does a randomized shelf-reset price experiment look like?

It looks like a discrete choice experiment where the price a respondent sees is randomly assigned rather than observed from history, run against the real, upcoming competitive set - the actual competitor SKUs, prices, and pack sizes expected to be on that shelf after the reset, not the current lineup. Because the assignment is random, the resulting effect is an estimate of what the price increase causes, not a correlation carried over from whatever price variation happened to occur historically. Randomized experiments analyzed with discrete choice models - McFadden discrete choice, Mixed Logit, and ICLV - are how the effect gets estimated once the manipulation is in place; the identification comes from the randomization in the design, not from the estimator itself. A flat multinomial logit model carries the IIA assumption, that a shopper's relative odds between two SKUs don't shift when a third option enters or leaves the set; Mixed Logit relaxes that assumption by letting preferences vary across the simulated population, which matters directly at a shelf reset because the competitive set itself is changing.

Conventional approach and randomized experiment, side by side

Scanner elasticity + survey backstopRandomized choice experiment
Price variation usedHistorical promotional depth, mostlyAssigned prices, including the increase never charged
Cross-price effect on categoryNot directly estimable; requires separate models per SKUEstimated directly against the real competitive set
Confound riskDisplay, feature, and seasonality travel with historical price cutsNone from history; confounds are controlled by randomization
Willingness-to-pay biasSurvey backstop overstates by an average of 21% (Schmidt and Bijmolt, 2020)Choice-based, no open-ended price question
Output granularityBrand volumeCategory dollars and brand volume, with confidence intervals
Best for:Teams with only historical data and no time before the resetTeams with six to twelve weeks to test a specific increase before the line review

What does the retailer actually want to see at the line review?

The category buyer wants category dollars and the cross-price effect on the rest of the set, not just the effect on the RGM team's own brand. A randomized experiment against the real competitive set delivers both directly, because every simulated shopper chooses among all the SKUs the design includes, not just the client's own line. That is the number a category buyer can act on: not "our brand's volume at the new price," but "what happens to the category's dollars, and which competitor picks up the share that moves." Report each effect with its confidence interval, understanding that the interval covers the effect within the simulated population and design tested, not the entire real market unconditionally, and state plainly which price points the design actually covered so the buyer knows the range the result speaks to.

How much can you trust a randomized price experiment?

Trust it to the extent the validation record supports, and no further, which is why the record should always ship with its denominator rather than a bare percentage. On the Hainmueller immigration conjoint study, our best configuration reaches 0.832 rank correlation against the published human result, where two independent samples of real humans reach 0.959 - 87% of that measured human ceiling, on one study, at the best configuration. Across all 43 published randomized studies that pass design filters, the mean is 0.73 (https://fidelity.subconscious.ai/papers/causal-fidelity/causal_fidelity_paper.pdf). Published studies can sit in a model's training data, which is a real limitation the replication protocol behind that paper is built to address, not a problem that disappears by asserting validation. The full leaderboard, updated as new replications run, is public at leaderboard; the methodology behind the fidelity numbers and how they're built is covered on the methods and validation hub.

Bar chart comparing three rank correlation values: human-to-human ceiling at 0.959, best configuration at 0.832, and mean across 43 studies at 0.73, all against the Hainmueller immigration conjoint benchmark.
87% is a ratio: 0.832 against a human ceiling of 0.959 on one study; the mean across 43 studies passing design filters is 0.73.

Before the next line review, pull the actual competitive set expected on the reset shelf - competitor SKUs, prices, and pack sizes - and write down the specific price points under consideration for the increase; that set is the input a randomized experiment needs, and having it in hand is the first concrete step regardless of which testing method gets used. For a walkthrough of how a shelf-reset design gets built against a specific category, book time on /meet.