Skip to content

Source of Volume Analysis: Category Growth vs Brand Switching

Two respected decompositions of the same promotion data disagree by more than forty points on how much of the volume came from switching. An innovation or category lead facing a retailer listing decision - and a CFO asking whether the new item will eat the existing line - needs a pre-launch answer to that split. The answer comes from randomizing the shelf set the new item competes on, including the brand's own existing line, and reading the substitution shares directly from choices made under that assignment.

Why do two respected decompositions disagree on the same promotion data?

They disagree because source of volume is a substitution quantity, and substitution changes with the measure you choose to compute it. van Heerde, Gupta, and Wittink re-examined the standard elasticity-based decomposition used across the promotion literature and found it attributes about 74% of a sales bump to brand switching. Decomposing the identical data in unit-sales terms instead, they found only about 33% of the volume gain traced to losses by other brands in the category (van Heerde, Gupta, and Wittink, JMR 40(4), 2003). The two figures come from the same promotion and the same market; only the measurement convention changes, and the gap between them runs more than forty points.

Bar chart comparing two decompositions of the same promotion dataset: elasticity-based switching estimate at 74%, unit-sales-based switching estimate at 33%.
The same promotion data attributes about 74% of the volume bump to brand switching in elasticity terms, and about 33% in unit-sales terms.

Why are the retailer and the CFO asking the same underlying question?

The retailer funds the listing if the item grows the category rather than reshuffling share within it. The CFO funds it if the incremental margin survives whatever the item takes from the brand's own line. Both questions resolve to the same substitution structure: where does each unit of volume come from, and how much of it was already the brand's before launch. A category lead who answers the retailer's question with intent-to-buy data and the CFO's question with a separate cannibalization study is running two analyses on one underlying causal quantity, and the two answers can contradict each other if they are not built from the same design.

Why does a flat choice model get the cannibalization number wrong?

A flat choice specification imposes proportional substitution: when a new item enters, it draws share from every existing option in proportion to that option's size, a consequence of the independence-of-irrelevant-alternatives (IIA) assumption baked into a simple multinomial logit. That assumption is convenient and wrong for a line extension. The sibling item on the same shelf is the closest substitute, and proportional draws understate exactly the loss the CFO is asking about. This is the proportional-substitution failure covered in the methods and validation library. The fix is a model specification, such as Mixed Logit or ICLV, that allows substitution patterns to differ by alternative instead of assuming they are uniform.

How do you decompose source of volume before the listing is funded?

You randomize the shelf set the new item competes on, run the choice task, and read the substitution shares directly from the assignment rather than inferring them from history. The design has to include the brand's own existing line as one of the alternatives on the shelf, not just competitor items, because that is the only way self-cannibalization shows up as a choice-share number instead of an assumption. Randomized experiments analyzed with discrete choice models produce the estimate. The method is McFadden discrete choice, Mixed Logit, or ICLV, chosen by how much substitution flexibility and latent preference structure the category needs. The choice shares by arm split into three parts: category expansion, the new buyers entering the category; competitive switching, the share taken from other brands; and self-cannibalization, the share taken from the brand's own line. Each part carries a confidence interval. That interval covers the effect within the simulated population under the tested design; it does not bound the real market outcome unconditionally, and a launch decision should treat it as a pre-launch estimate, not a guarantee.

MethodWhen you get the answerHow it handles the sibling itemBest for
Panel gain-and-loss decompositionAfter launch, from household panel dataReads actual switching after the fact, no design control over what's comparedAuditing a launch that already happened; see the post-launch incremental-versus-cannibalistic read
Purchase-intent score scaled by awareness/distributionBefore launch, from a single-item concept testDoes not model competition directly; no substitution structure at allA rough go/no-go screen when no competitive read is needed
Randomized shelf-set experiment with own line includedBefore launch, from choice shares by armOwn line is a tested alternative, so the model measures cannibalization directlyThe retailer/CFO decision this article addresses, where both category growth and cannibalization must be funded from one number

How much can you trust a pre-launch causal read?

Our best configuration reaches a 0.832 rank correlation against the published human result on the Hainmueller immigration conjoint, against a measured human-to-human ceiling of 0.959 from two independent samples of real humans - 87% of that ceiling, stated with its denominator (Causal Fidelity paper). Across all 43 published randomized studies that pass design filters, the mean rank correlation is 0.73. These are validation results against studies that were already public, so some may sit inside a model's training data; the replication protocol is built to surface that risk rather than hide it. The Hainmueller study is an immigration conjoint, a political-preference task far from a retail shelf set, so its score validates the causal-inference machinery rather than performance on a category-growth-versus-switching decision specifically; a new category is a new test. The current standing of every published comparison is on the leaderboard, which is the place to check before treating any single number as settled.

What should the category lead bring to the retailer meeting?

Bring choice shares by arm. The retailer wants to see the category-expansion share; the CFO wants the cannibalization share against the existing line; both numbers come from the same randomized design and the same confidence interval, so there is nothing to reconcile between the two conversations. If a prior source-of-volume estimate on the table came from a flat choice model or a post-launch panel read, ask which measurement convention produced it before deciding what it means, because the gap between 74% and 33% on identical promotion data shows that the measurement convention can be doing most of the work (van Heerde, Gupta, and Wittink, 2003).

The concrete next step: before the next listing pitch, write down which of the three source-of-volume buckets - category growth, competitive switching, self-cannibalization - the current forecast actually measures, and which it assumes. If any bucket is an assumption rather than a measured share, that is the gap a randomized shelf-set experiment closes. For a walkthrough of how that design gets set up for a specific launch, meet the team.