Criticisms and counter-criticisms of Kano model
A product leader deciding whether to greenlight a feature for the next roadmap cycle needs a straight answer on whether Kano's basic, performance, and delighter categories can be trusted for that decision, or whether the mounting criticism against the model means it should be retired. Short answer: the instability and wording problems raised since 2015 are real and partially fixable, but they aren't the main issue. Kano's dual-response survey was built to record how satisfied someone says they'd feel about a feature in isolation, not to isolate the causal, traded-off effect that feature has on what the person actually chooses when price and competing options are on the table.
- The sharpest published critique, Chapman and Callegaro's 2022 Sawtooth Software Conference paper, finds Kano's dual-response items behave like unreliable survey questions, with category assignment unstable below roughly 200 respondents.
- Terry Grapentine's 2015 Quirks piece reaches a parallel conclusion from a psychometrics angle: no validated scale, and answers shift with wording, and he recommends conjoint-type trade-off methods instead.
- Most practitioners haven't switched anyway. Slevitch's 2025 systematic review finds hospitality and tourism studies still default to the original 1984 scoring table despite years of proposed categorization fixes.
- Every fix on offer, larger samples, adaptive dual-response scoring, AI-assisted classification, repairs measurement noise in the self-report design. None of them puts the feature into an actual trade-off against price or a competitor.
- A randomized trade-off experiment addresses that gap directly by making the choice itself the measurement, not a follow-up satisfaction rating.
What is the core criticism of the Kano model?
The core criticism is that Kano's dual-response items, the "how would you feel if this feature were present" and "how would you feel if it were absent" pair, function as low-quality survey questions rather than a validated measurement instrument. Chapman and Callegaro (2022) argue the response scale is multidimensional rather than unidimensional, meaning the same answer can reflect different underlying attitudes, and that category assignment (must-have, performance, delighter, indifferent) becomes unreliable below roughly 200 respondents. The Quant UX Blog's assessment makes the instability concrete: studies run at N=20, N=30, even N=100 are highly likely to produce different category assignments if repeated, with stable aggregate answers only showing up around N=200 and above. For a study that's typically run once, on a convenience sample, to decide which feature gets built, that's a real problem before any deeper issue is even on the table.
The counter-criticism: why product teams keep using it anyway
The standard defense is that Kano is simple, cheap, and communicates in three buckets an executive can act on without a statistics background, and that the sample-size problem is a known, fixable limitation, not a fatal flaw. That's the conventional treatment: run a bigger sample, tighten the item wording, and the categorization holds up. It's also, per Slevitch's 2025 review, what most teams actually do: despite a growing menu of categorization fixes published over the past decade, most hospitality and tourism studies in his systematic review still default to the original 1984 scoring table. The model survives less because the criticism has been answered and more because switching costs more than tolerating the noise.
Where the 2025 fixes still fall short
Slevitch's review catalogs the current wave of patches, including an adaptive dual-response variant and AI-assisted scoring, alongside the older fixes of bigger samples and cleaner item wording. Each one targets the same failure point: getting a more stable category out of noisy self-report data. None of them changes what the respondent is actually being asked to do, which is rate a feature they've often never used, alone, with no price attached and no competing option in view. A more stable estimate of an unstable question is still an estimate of the wrong thing for a roadmap or pricing decision.
The flaw underneath the scoring table
This is the structural issue the sample-size debate skips past. Satisfaction categorization was never built to isolate the causal, traded-off effect a feature has on what someone actually chooses; it was built to sort self-reported reactions into buckets. A person can honestly report that a feature would delight them and still not pick the product that has it once price or a competing feature enters the decision. Kano's survey design has no mechanism for surfacing that, because nothing in the question is ever varied or traded off.
Why delighters are the hardest thing to survey
The delighter category is definitionally a forecast of reaction to a surprise, and forecasting reactions to things people haven't experienced is precisely what stated-preference surveys do worst. Grapentine's conclusion follows from this: Kano lacks psychometric validation, scale wording measurably shifts outcomes, and he recommends conjoint-type trade-off methods for product design questions instead of satisfaction ratings. The same weakness shows up around pricing. When a delighter is really a premium feature, asking people to state what they'd pay for it runs into the same problem willingness-to-pay questions always have: stated willingness-to-pay comes in high relative to what people actually spend, unless the design incentive-aligns the choice, meaning it has real consequences for the respondent rather than a hypothetical answer.
Is Kano compatible with randomized, causal experiment design?
Not as it's currently practiced, but the underlying question it's trying to answer, does this feature change what someone chooses, is testable with randomized experiments analyzed with discrete choice models rather than dual-response items. McFadden discrete choice, Mixed Logit, and ICLV are estimators, not causal methods on their own; the causal identification comes from randomly varying the feature, price, and alternatives within the experiment design, not from the estimator that analyzes the resulting choices. When a flat multinomial logit is used, it carries the independence-from-irrelevant-alternatives assumption, which is part of why Mixed Logit and ICLV extensions exist: they relax that assumption when preferences vary across people or when a new option pulls share unevenly from existing ones.
Whether this kind of design actually reproduces real buying behavior is a testable claim, not an assumption. In a validation study, simulated randomized trade-off experiments reproduced the direction and outcome of the original human study 93 percent of the time, a validation-set result reported at go.subconscious.ai/paper, not a guarantee for a new market the design hasn't been tested against. That number also doesn't resolve whether any of the original studies in that validation set could have been part of a model's training data before the test; the replication protocol is built to check for that rather than assume it away. And a confidence interval reported from a simulated experiment covers the effect within the simulated population that was fielded, not an unconditional bound on the real market. The public leaderboard shows how different designs perform against held-out human studies, which is the level of scrutiny a Kano scoring table has never been put through.
What replaces it: a decision framework
The choice isn't Kano versus nothing; it's which tool matches the decision in front of you.
| Method | What it measures | Best for |
|---|---|---|
| Kano dual-response survey | Self-reported satisfaction with one feature, rated alone, no trade-off | Best for: fast internal sorting of raw feature ideas before any real spend is committed |
| McFadden discrete choice | Which option people pick when features and price are randomly varied against each other | Best for: a funded roadmap or pricing decision where you need to know which feature moves the actual choice |
| Mixed Logit | The same choice data, allowing preferences to vary across respondents and relaxing the IIA assumption | Best for: markets with distinct segments, or where a new feature might pull share unevenly from existing options |
| ICLV | Choice data linked to latent attitudes, like trust or status, that drive the observed choice | Best for: decisions where the reason behind the choice matters as much as the choice itself, such as messaging or positioning |
More detail on how these designs are validated is on the methods and validation blog.
If a Kano survey is already scheduled for an upcoming roadmap review, the concrete next step is to add one randomized trade-off task to the same fielding: the feature in question against price and a real competing option, then check which one actually predicts a pick rather than a stated satisfaction rating. Compare the result against the public leaderboard before treating either number as final. If it's easier to walk through setting one up, meet the team.