How to manually calculate partworth utilities
Buyers evaluating a conjoint vendor's output have one real decision to make: whether to trust the part-worth utilities enough to act on them, not whether the underlying regression ran correctly. You can manually calculate partworth utilities from a rating-based (CVA) study in three steps: compute each attribute's utility range from the regression coefficients, zero-center the part-worths within each attribute, then rescale across attributes to get relative importance. That check confirms the estimator was applied correctly to the data it was given. It says nothing about whether the preferences it recovered are the ones respondents would actually act on.
- The manual process is standard: compute attribute-level utility ranges, zero-center part-worths within each attribute, rescale across attributes for relative importance.
- OLS regression handles rating-based CVA designs; hierarchical Bayes (HB) is now standard for choice-based (CBC) designs because CBC tasks are too sparse per respondent for OLS to solve individually.
- A correct calculation proves the estimator was applied correctly. It does not prove the responses predict real-world behavior.
- Well-designed paired discrete-choice conjoint has matched real referendum outcomes in published research; ranking and rating tasks degraded badly by comparison.
- Subconscious's best configuration reaches 87% of a measured human ceiling on one study: 0.832 rank correlation against a published human result, where two independent human samples reach a rank correlation of 0.959 with each other. Mean across the 43 studies that passed design filters is 0.73 (Causal Fidelity paper). That is a validation result, not a guarantee for a new market.
How do you manually calculate partworth utilities?
You run the process every conjoint textbook and vendor help center describes the same way: fit a regression of ratings or choices on dummy-coded attribute levels, compute the utility range within each attribute (highest coefficient minus lowest), zero-center the part-worths so they sum to zero within each attribute, then divide each attribute's range by the sum of all attribute ranges to get relative importance weights (QuestionPro, Part Worths - Conjoint Analysis). For a simple CVA study with dummy-coded levels and individual-level data, this is a spreadsheet-tractable OLS regression. Academic guides from WU Vienna and MIT Sloan teach the identical sequence. None of these guides claim the output is anything more than a correctly transformed version of the input data.
Why most CBC studies use hierarchical Bayes instead of OLS
Because modern choice-based conjoint designs don't give OLS enough data per respondent to solve. Hierarchical Bayes was introduced to conjoint estimation around 1995 specifically to recover individual-level part-worths from CBC designs where each respondent sees too few profiles for an aggregate or per-person OLS fit to converge (Sawtooth Software, Hierarchical Bayes Estimation). HB borrows statistical strength across respondents to stabilize individual estimates, which is why it dominates CBC pipelines today while OLS remains fine for the older rating-based CVA format. This is a real methodological choice, and vendors are right to explain it. It is also a choice entirely internal to the estimator. Neither OLS nor HB has any mechanism for checking whether the choices respondents made in the survey resemble the choices they'd make with real money and real consequences.
Auditing the arithmetic vs validating the preferences
These are different checks, and a buyer who only runs the first one has not touched the second.
| Manual arithmetic check | Behavioral validation | |
|---|---|---|
| What it confirms | The estimator (OLS or HB) was applied correctly to the design matrix and response data | The resulting preferences correspond to choices made outside the survey instrument |
| What it can't confirm | Whether respondents' stated preferences reflect what they'd actually do | Anything beyond the population and design the benchmark was measured against |
| Evidence required | Design matrix and response data only | An independent behavioral benchmark: real referendum results, a holdout human study, purchase data |
| Best for | Buyers confirming a vendor didn't make a computational error | Buyers deciding whether to act on the result |
What does a correct part-worth calculation actually prove?
It proves the arithmetic is right, not that the preferences are real. Reverse-engineering a vendor's coefficients and finding they "check out" confirms internal consistency between the estimator and the input data. It says nothing about external validity, which is whether the input data (the choices respondents made in a survey) resembles the choices those same people would make outside it. A buyer who stops at the spreadsheet has validated the layer of the claim that was never in question. The layer that matters, whether the recovered preferences predict real behavior, sits one step further out and requires a different kind of check entirely.
Do conjoint results predict real-world behavior?
Sometimes. It depends on task design, not arithmetic. A 2015 PNAS study by Hainmueller, Hangartner, and Yamamoto compared conjoint-derived effects on support for immigrant naturalization against actual outcomes from Swiss municipal referendums and found close correspondence, but only for paired discrete-choice tasks (PNAS, 2015). Ranking and rating tasks, the same tasks that OLS was built to handle, degraded badly by comparison in the same study. That result was earned through design discipline and an out-of-sample comparison against real behavior, not through auditing a regression. It is also a single validated domain (referendum voting), which is a limitation worth stating plainly rather than generalizing past.
Where does the say-do gap hide inside a correct regression?
It hides in the survey responses themselves, before the estimator ever runs, which is why no amount of recomputation catches it. Hypothetical bias, the tendency for stated willingness to act to run higher than actual behavior, is a documented and unresolved problem in the stated-choice literature. A part-worth calculation has no step that detects this. It takes whatever the respondent said as ground truth and transforms it faithfully. The transformation can be flawless and the ground truth can still be wrong.
How does Subconscious validate part-worths against real behavior instead of just the math?
By running randomized experiments analyzed with discrete choice models on a simulated population and checking the result against a measured human ceiling before treating it as usable. Subconscious's methods are McFadden discrete choice, Mixed Logit, and ICLV. These are estimators, not causal methods. The causal identification comes from the randomized manipulation built into the experiment design, not from the choice of model. Flat multinomial logit assumes independence of irrelevant alternatives, which is one reason Mixed Logit and ICLV are used for preference-share and substitution questions rather than a base logit alone.
On the replication protocol, the best configuration reaches 87% of the measured human ceiling on one study: 0.832 rank correlation against a published human result, where two independent human samples reach a rank correlation of 0.959 with each other. Across the 43 studies that passed design filters, the mean is 0.73 (Causal Fidelity paper). Full results by study are on the leaderboard.
That is a validation result on studies already measured, not a guarantee for a new market. Published human studies can sit inside a model's training data, so the replication protocol is built to check for that rather than assume it away. The general validation approach is covered in the methods and validation hub.
If you're deciding whether to trust a set of part-worth utilities, don't stop at recomputing the coefficients. Ask the vendor what independent behavioral benchmark, if any, the output was checked against, and whether that check used paired discrete-choice tasks or ratings and rankings, since the PNAS results show that distinction matters more than the estimator does. If you want a second read on how a study's replication protocol was structured before you commission it, book a working session with the team.