Skip to content

Classifying Types of Conjoint Analysis

---

Title: Classifying Types of Conjoint Analysis: Why Format Isn't the Whole Decision

A senior research buyer picking a conjoint format is actually making two separate decisions, and only one of them usually gets written up. Four working types cover most cases: choice-based (CBC), adaptive choice-based (ACBC), MaxDiff, and ranking- or rating-based. Picking among them decides what you can ask a respondent. It says nothing about whether the resulting answer would hold up against real behavior.

- CBC is the default because it mirrors actual purchase behavior, asking respondents to choose one of three to five profiles, repeated across screens ([Qualtrics](https://www.qualtrics.com/experience-management/research/types-of-conjoint/)).
- ACBC exists for products with more attributes than a respondent can process in one CBC screen; it adapts the choice set to each respondent's prior answers ([Sawtooth Software](https://sawtoothsoftware.com/resources/knowledge-base/sales-questions/which-conjoint-method-grid)).
- MaxDiff prioritizes long attribute or message lists. It does not simulate a market choice the way CBC does.
- Ranking and rating formats ask for a rating on every profile instead of a choice among a few, which raises respondent burden as the list grows.
- None of these four formats tell you whether the design was checked against real behavior. That is a separate decision, and it is the one that determines whether the resulting numbers are trustworthy.

## How do the four conjoint types differ?

Format determines two things: how many attributes a respondent can realistically hold in one screen, and how closely the task resembles an actual purchase decision.

| Type | When it fits |
|---|---|
| CBC | Default for market simulation and price sensitivity. Mirrors real purchase behavior by asking respondents to pick one of three to five profiles, repeated across screens. |
| ACBC | Products with more attributes than a respondent can hold in a single CBC screen. Adapts the choice set to each respondent's prior answers. |
| MaxDiff | Prioritizing a long list of features, claims, or messages against each other, not simulating a market choice. |
| Ranking or rating | The oldest format. Asks for a rating on every profile individually rather than a choice among a small set, which raises per-item burden as the list grows. |

The taxonomy keeps expanding past these four. Independent catalogs now list more than a dozen named variants ([OpinionX](https://www.opinionx.co/blog/conjoint-analysis-types)), which is part of why platforms like Sawtooth publish decision grids just to give buyers a starting point.

## Why isn't picking the right format enough?

Format choice controls what you can ask. It does not test whether the answer describes real behavior.

Meta-analyses across the stated-preference literature find stated willingness-to-pay running roughly two to three times higher than revealed willingness-to-pay, with no reliable field-wide fix ([ScienceDirect](https://www.sciencedirect.com/science/article/abs/pii/S1755534521000555)). That gap does not shrink because the format suits the attribute count. It is a property of asking people to state a preference at all.

One documented exception is instructive. Hainmueller, Hangartner, and Yamamoto compared conjoint-estimated effects on support for immigrant naturalization against real Swiss referendum outcomes, and the survey-based estimates matched behavior closely ([PNAS](https://www.pnas.org/doi/10.1073/pnas.1416587112)). That result held because the design was checked against an external outcome, not because of which format was chosen.

A validated design carries a specific, sourced number, not a promise. Our best configuration reaches [87% of the measured human ceiling](https://fidelity.subconscious.ai/papers/causal-fidelity/causal_fidelity_paper.pdf) (0.832 rank correlation against the published human result, where two independent human samples reach 0.959; mean 0.73 across the 43 studies passing design filters).

![Bar chart of three rank correlation values: human ceiling at 0.959, best configuration at 0.832, and mean across 43 studies at 0.73.](/images/authority/classifying-types-conjoint-analysis.svg "The best configuration reaches 0.832 of a 0.959 human ceiling, 87% of that benchmark. The mean across 43 studies passing design filters is 0.73.")

Even a validated result needs one more caveat. Published human studies can sit in a model's own training data, so a match against a public benchmark does not by itself rule out memorization. The replication protocol behind the figure above tests for this directly. That protocol reduces the memorization risk; it does not eliminate the general one, and it does not turn a validation result into a guarantee for a new market. The figure above measures fit against studies that have already happened. A market you have not yet tested has no equivalent published number, and none of the figures above cover it in advance. Ask any vendor how they control for training-data leakage before you trust their number.

These four formats are typically estimated with discrete choice models: McFadden's conditional logit, for which McFadden won a share of the 2000 Nobel Prize in Economics, Mixed Logit, or ICLV, a later extension that adds latent variables to the choice model ([Econlib](https://www.econlib.org/library/Enc/bios/McFadden.html)). The estimator fits the functional form. The causal read comes from the randomized profiles in the experiment design, not from the model. When CBC results are used to simulate market share or substitution between profiles, a flat logit assumes independence of irrelevant alternatives (IIA), which can distort predicted share shifts between similar options.

## Which conjoint type should you pick?

Match format to attribute count and respondent burden first. Confirm validation second, as a separate check, not a formality.

Run the format decision through Sawtooth's or Qualtrics' published grids. Then ask your vendor for a validation number in the same ratio grammar used above: a rank correlation against a published human benchmark, stated alongside that benchmark's own ceiling. A confidence interval from a simulated experiment covers the effect within that simulated population; it is not an unconditional bound on the real market until checked against an outcome like the Swiss referendum comparison above. If a vendor can't produce a sourced number in that form, you are buying a format, not proof.

See the [leaderboard](/leaderboard) for how these validation numbers get published and compared, and the [methods and validation hub](/blog/methods-and-validation) for the rest of this argument. Or talk it through directly: [book time](/meet).