Skip to content

How to Run a Shelf Simulation Before a Retailer Meeting

A category or shopper insights lead with a line review in four to eight weeks needs a planogram recommendation backed by a number the retailer will act on: the effect of a specific shelf change on category dollars, stated with a confidence interval. Run a shelf simulation before that meeting when the shelf itself is randomized across simulated respondents inside one design. A realistic-looking virtual aisle answers whether shoppers behave plausibly in a simulated store, and the published validation work on virtual shelves already answers that question reasonably well (Waterlander, Jiang, Steenhuis and Ni Mhurchu, Journal of Medical Internet Research 17(4):e107, 2015, https://pmc.ncbi.nlm.nih.gov/articles/PMC4429224/).

What decision does a shelf simulation need to answer before a retailer meeting?

The decision is narrow: which items get facings and where the block sits. The simulation also has to state what that planogram change does to category value if the retailer accepts it. The retailer is asking an effect question. A simulation that outputs "shoppers preferred cell B over cell A and C" has not answered it, because it cannot say which element of cell B, held everywhere else constant, produced the gain. The retailer in the room is deciding on one specific shelf change and wants its size, with a stated confidence interval.

What virtual shelf research already validates

Virtual shelf research validates reasonably well on the question it can answer. Waterlander, Jiang, Steenhuis, and Ni Mhurchu ran 60 participants through three virtual shopping trips and matched the results against real till receipts (Journal of Medical Internet Research, 17(4):e107, 2015, https://pmc.ncbi.nlm.nih.gov/articles/PMC4429224/). The same four categories carried the highest spend in both settings: fresh produce at 14.3% of spend virtual versus 17.4% real, dairy at 19.1% versus 12.6%, meat and fish at 16.5% versus 16.8%, bread and bakery at 10.0% versus 8.2%. The internal order shifted - dairy led virtual spend, produce led real spend - and across those 60 participants and three trips, six of the study's categories showed statistically significant gaps between the two settings. A shelf simulation is a legitimate place to observe shopping behavior at the category level, with that gap in mind.

Where a realistic shelf still misleads the room

The problem shows up in what the fielding constraint forces you to test. A shopper panel and a hand-built virtual store cost real recruit time per cell, so most projects cap out at three or four configurations, and those three or four are usually the layouts a category team already favored going in. The simulation then reports which of those specific layouts shoppers liked best. It cannot report what happens if you add one facing to the number two item, because that configuration was never built. Waterlander's own comparison found significant gaps in six categories against real receipts (Waterlander et al. 2015), so a raw purchase-intent number pulled from that kind of study deserves the same caution: treat it as directional.

How does randomizing the shelf change what you can claim?

Randomizing facings, adjacency, block position, price, and pack across simulated respondents turns each variable into something estimable on its own, separate from the hand-picked cell it would otherwise sit inside. Because an added configuration does not require a new recruit, the design can cover the range of planograms worth knowing about. The output changes shape: the result names the changed attribute directly, for example adding one facing to an item, and reports the estimated shift in category dollars with a stated interval, attributable to that change because assignment put a shopper in front of it.

The analysis method matters here. Randomized experiments analyzed with discrete choice models such as McFadden's model, Mixed Logit, and ICLV recover the effect because the shelf attributes were randomized in the design. A standard multinomial logit carries the independence of irrelevant alternatives assumption, which can misstate substitution between two similar items on the shelf; Mixed Logit relaxes that by letting preferences vary across simulated respondents, and ICLV adds a layer for latent constructs like health or convenience motivation behind the choice. A confidence interval from this design covers the effect within the simulated population tested; it does not extend automatically to the retailer's shopper base. Published human studies used to validate these models can sit inside a model's training data, which is why replication protocols test against newer, held-out studies rather than treating one match as proof.

Conventional cell testing versus a randomized shelf design

DimensionConventional virtual-shelf testRandomized shelf design
Configurations testedThree or four hand-built cells, capped by recruit cost and timeAs many attribute combinations as the design specifies, without adding recruits
What you learnWhich of the tested cells shoppers preferred overallThe effect of a specific change (facings, adjacency, price, pack) on category dollars, with an interval
Realism groundingCategory-level spend patterns validated against till receipts ([Waterlander et al. 2015](https://pmc.ncbi.nlm.nih.gov/articles/PMC4429224/))Same simulated-shelf mechanics; realism claim applies equally
Known biasCategory-level agreement is strong but imperfect against real receipts, per the same studySame caveat; the object of inference is the relative effect
Best for:A quick read on whether a small set of layouts someone already likes looks reasonableA category team that needs a specific, defensible planogram recommendation with a stated interval before a line review

How much of the measured human ceiling does this reach?

On one published benchmark study, our best configuration reaches 87% of the measured human ceiling: 0.832 rank correlation against the published human result, where two independent samples of real humans reach 0.959 against each other (causal fidelity paper). That is a best-case number on a single study, not a guarantee for a new retailer's category. Across all 43 published randomized studies that pass the paper's design filters, the mean rank correlation is 0.73. Bringing a category number into a retailer meeting means bringing it with that context and its denominator attached.

Bar chart showing human-to-human rank correlation ceiling at 0.959, best single-study configuration at 0.832, and mean across 43 published studies at 0.73, all sourced to the causal fidelity paper.
The 0.832 best-case figure is 87% of the 0.959 human ceiling on one study; the mean across all passing studies is lower, at 0.73.

Full study-by-study results sit on the public leaderboard, which is the place to check where a given method and market stand before treating any single number as settled.

What the interval covers, and what it doesn't

A confidence interval from this kind of design covers the estimated effect within the simulated population tested, given the randomized shelf attributes included in the design. It does not bound what happens in the retailer's actual store network, and it does not cover attribute combinations the design left out. If the design randomized facings, adjacency, block position, price, and pack, say so, and say which of those five the design held fixed, since a fixed attribute cannot produce an estimated effect. This is also where the IIA assumption matters: if the report leans on a flat multinomial logit for a preference-share or substitution claim between two similar items, name that assumption in the same sentence as the claim, because it is what shapes the substitution pattern the retailer will ask about directly.

What to bring into the line review

Bring the effect of each planogram change on category dollars, the interval around it, and a short statement of what the design covered. Skip a bare preference ranking across a small set of cells, and flag any purchase-intent or spend number pulled from a fixed-cell study as directional rather than exact. Methods pages on randomized experiments and validation walk through how the discrete choice models behind these estimates work, and the public leaderboard is the fastest way to check a method's validation record against a specific study before repeating its number in a room.

Before the review: list every attribute the planogram recommendation depends on (facings, adjacency, block position, price, pack), check which ones the simulation varied by assignment, and cut any claim resting on an attribute the design held fixed. For a second look at the design before it runs, book time with the team.