Kano or MaxDiff: Which is better for feature selection?
Kano and MaxDiff answer different questions, but neither is the right axis for a roadmap decision. Both are stated-preference surveys from the 1980s: they measure what respondents say, not what they do. The real choice is whether to keep debating survey format or run a randomized discrete choice experiment that prices each feature's causal effect on adoption, with a confidence interval attached.
The short version:
- Kano sorts features into must-be, performance, and delighter buckets from paired functional/dysfunctional questions. MaxDiff force-ranks features by relative importance from repeated best-worst choices. Neither produces a causal effect size.
- If you must pick one stated-preference tool, the standard advice holds: Kano to screen for must-be features, MaxDiff to rank the survivors. That combination still measures attitudes, not behavior.
- Swapping Kano for MaxDiff doesn't close the say-do gap. Neither survey puts a real tradeoff, money, a competing product, a budget, in front of the respondent, and neither has been benchmarked against real purchase or usage data.
- Subconscious's best-configuration study, a randomized discrete choice experiment analyzed with Mixed Logit and ICLV, reaches 87% of a measured human ceiling: 0.832 rank correlation against a human-to-human baseline of 0.959, with a mean of 0.73 across the 43 studies passing design filters, per the causal fidelity paper. That's a validation result on published studies at a point in time, not a guarantee for a new market. Kano and MaxDiff don't report a comparable number.
- Next step: before locking a roadmap, audit whether your last Kano or MaxDiff finding was ever checked against actual usage or purchase data. If it wasn't, pilot a small randomized discrete choice experiment on your top 5 to 8 features instead of running a third survey.
What does Kano actually measure?
Kano measures how a feature affects stated satisfaction, not whether it drives adoption. Noriaki Kano's 1984 method asks a paired functional/dysfunctional question for each feature ("how would you feel if this were present" and "how would you feel if this were absent") and sorts the answers into must-be, performance, and delighter categories. Quantilope's comparison frames this correctly as a satisfaction-impact taxonomy, distinct from a ranking task. The dysfunctional question is genuinely useful for one thing: mmrresearch's 2025 analysis argues it's still the only reliable way to detect true must-be features, because it separates "I wouldn't notice this" from "I'd be furious without it." That's a real strength. It's still a respondent guessing at a hypothetical feeling, not a person facing a real tradeoff.
What does MaxDiff actually measure?
MaxDiff measures relative importance among stated preferences, with no satisfaction label attached. Best-worst scaling, developed by Jordan Louviere and commercialized by Sawtooth Software, presents respondents with subsets of features and asks them to pick the best and worst of each set, repeatedly. The output is a clean rank order across a long list, without the scale-use bias that plagues simple rating questions. Qualtrics, Displayr, QuestionPro, and quantilope all ship MaxDiff modules, with Sawtooth still treated as the reference implementation. What MaxDiff can't do is explain why a feature ranks where it does: mmrresearch notes it struggles to distinguish a must-be feature (furious without it, indifferent with it) from an excitement feature (thrilled with it, indifferent without it), because both can produce similar mid-pack importance scores.
Kano vs. MaxDiff at a glance
| Dimension | Kano | MaxDiff |
|---|---|---|
| Question format | Paired functional/dysfunctional questions per feature | Repeated best/worst choices across feature subsets |
| Output | Satisfaction category: must-be, performance, delighter | Rank order of relative importance |
| Strongest at | Detecting true must-be features via the dysfunctional question ([mmrresearch](https://www.mmrresearch.com/post/is-kano-modeling-interchangeable-with-maximum-differentiation)) | Fine-grained ranking across long lists without scale bias ([quantilope](https://www.quantilope.com/resources/kano-va-maxdiff)) |
| Weakest at | No rank order across a full feature set | Struggles to separate must-be from excitement-type features |
| Tested against real choice | No | No |
| Best for | A short candidate list needing a satisfaction taxonomy before deeper testing | A long backlog you need to force-rank, with no causal claim attached |
Kano or MaxDiff: which is better for feature selection?
Neither, because "better" implies one returns something the other doesn't, and both return a stated-preference label with no attached uncertainty about real-world adoption. The conventional decision tree, Kano for satisfaction categories, MaxDiff for rank order, both if budget allows, is reasonable advice for choosing between two survey formats. It answers the wrong question for a roadmap decision. A rank order or a satisfaction bucket doesn't tell you what happens to adoption, retention, or revenue if the feature ships. It tells you how a respondent answered a hypothetical question about a feature they weren't paying for, weren't trading off against a competitor, and weren't buying under a real budget constraint. That's the axis that matters, and Kano-versus-MaxDiff never crosses it.
Why doesn't switching from Kano to MaxDiff close the say-do gap?
It can't, because the say-do gap is a property of stated-preference methodology itself, not a defect specific to either survey format. A respondent resolving "how would you feel if this were missing" and a respondent resolving "pick the best and worst of these four" are both answering a question with no money, no competing product, and no real budget attached. Swapping one instrument for the other changes the shape of the output, a category versus a rank, but it doesn't put a real tradeoff in front of the respondent. Neither Kano nor MaxDiff has been benchmarked against real purchase or usage data in the comparisons reviewed here, so there's no method-specific number to report, only the shared absence of one.
Is "Tandem MaxDiff" a real fix?
No, because it combines two stated-preference outputs into one study instead of adding a tested behavioral signal. KSR's 2025 argument for "Tandem MaxDiff" proposes running MaxDiff instead of, or alongside, Kano to get both satisfaction-style signal and rank order from a single instrument. That's a legitimate efficiency gain if you were going to run two separate stated-preference studies anyway. It's still two flavors of the same measurement problem stacked together: a satisfaction inference layered onto a best-worst ranking, both derived from hypothetical questions, neither checked against a purchase, a signup, or a churn event. Combining the surveys makes the research cheaper to run. It doesn't make the output causal.
What would a causal test of feature value actually look like?
It looks like a randomized experiment where feature bundles and prices vary across respondents by design, analyzed with a discrete choice model that estimates the effect of each manipulated attribute. McFadden discrete choice, Mixed Logit, and ICLV are estimators; they extract a coefficient and a confidence interval from choice data. The causal claim comes from the randomization in the experimental design, not from the estimator, so the correct description is randomized experiments analyzed with discrete choice models, not "causal methods like DCE." Subconscious runs this design on a simulated population and validates it against published human studies: the best-configuration result reaches 0.832 rank correlation against a human benchmark of 0.959, an 87% ratio, with a mean of 0.73 across the 43 studies that passed design filters, per the causal fidelity paper. That's a validation result on studies published at a point in time, not a guarantee that holds for every new market, and because published studies can overlap with a model's training data, the replication protocol is built specifically to test for that overlap rather than assume it away. A confidence interval produced this way covers the effect within the simulated population tested; it doesn't bound the real market unconditionally.
How does this compare against a real human baseline, in practice?
Check it the way you'd check any measurement instrument: against a public, repeatable benchmark rather than a vendor's internal claim. The leaderboard tracks replication performance study by study, so a buyer can see where a given design falls relative to the human baseline instead of taking a single average on faith. For adjacent comparison decisions, the comparisons hub and methods and validation hub cover related tradeoffs in feature and pricing research.
If you're choosing between Kano and MaxDiff this week, run the one that fits your immediate need: Kano for a short list needing satisfaction categories, MaxDiff for a long backlog needing a rank order. Treat that result as a screening pass, not a roadmap-ending answer. Before you commit budget against it, price a randomized discrete choice pilot on your top 5 to 8 features and see whether the effect estimate and its confidence interval change your priority order. If you want a second opinion on the design, meet with the team.