Skip to content

How to manually calculate partworth utilities

Buyers evaluating a conjoint vendor's output have one real decision to make: whether to trust the part-worth utilities enough to act on them, not whether the underlying regression ran correctly. You can manually calculate partworth utilities from a rating-based (CVA) study in three steps: compute each attribute's utility range from the regression coefficients, zero-center the part-worths within each attribute, then rescale across attributes to get relative importance. That check confirms the estimator was applied correctly to the data it was given. It says nothing about whether the preferences it recovered are the ones respondents would actually act on.

How do you manually calculate partworth utilities?

You run the process every conjoint textbook and vendor help center describes the same way: fit a regression of ratings or choices on dummy-coded attribute levels, compute the utility range within each attribute (highest coefficient minus lowest), zero-center the part-worths so they sum to zero within each attribute, then divide each attribute's range by the sum of all attribute ranges to get relative importance weights (QuestionPro, Part Worths - Conjoint Analysis). For a simple CVA study with dummy-coded levels and individual-level data, this is a spreadsheet-tractable OLS regression. Academic guides from WU Vienna and MIT Sloan teach the identical sequence. None of these guides claim the output is anything more than a correctly transformed version of the input data.

Why most CBC studies use hierarchical Bayes instead of OLS

Because modern choice-based conjoint designs don't give OLS enough data per respondent to solve. Hierarchical Bayes was introduced to conjoint estimation around 1995 specifically to recover individual-level part-worths from CBC designs where each respondent sees too few profiles for an aggregate or per-person OLS fit to converge (Sawtooth Software, Hierarchical Bayes Estimation). HB borrows statistical strength across respondents to stabilize individual estimates, which is why it dominates CBC pipelines today while OLS remains fine for the older rating-based CVA format. This is a real methodological choice, and vendors are right to explain it. It is also a choice entirely internal to the estimator. Neither OLS nor HB has any mechanism for checking whether the choices respondents made in the survey resemble the choices they'd make with real money and real consequences.

Auditing the arithmetic vs validating the preferences

These are different checks, and a buyer who only runs the first one has not touched the second.

Manual arithmetic checkBehavioral validation
What it confirmsThe estimator (OLS or HB) was applied correctly to the design matrix and response dataThe resulting preferences correspond to choices made outside the survey instrument
What it can't confirmWhether respondents' stated preferences reflect what they'd actually doAnything beyond the population and design the benchmark was measured against
Evidence requiredDesign matrix and response data onlyAn independent behavioral benchmark: real referendum results, a holdout human study, purchase data
Best forBuyers confirming a vendor didn't make a computational errorBuyers deciding whether to act on the result

What does a correct part-worth calculation actually prove?

It proves the arithmetic is right, not that the preferences are real. Reverse-engineering a vendor's coefficients and finding they "check out" confirms internal consistency between the estimator and the input data. It says nothing about external validity, which is whether the input data (the choices respondents made in a survey) resembles the choices those same people would make outside it. A buyer who stops at the spreadsheet has validated the layer of the claim that was never in question. The layer that matters, whether the recovered preferences predict real behavior, sits one step further out and requires a different kind of check entirely.

A left-to-right chain showing survey response feeding an estimator, producing part-worths, where a manual arithmetic check stops, followed by an unverified gap before real-world choice.
Recomputing the regression only checks the first half of this chain; nothing about it touches the second half.

Do conjoint results predict real-world behavior?

Sometimes. It depends on task design, not arithmetic. A 2015 PNAS study by Hainmueller, Hangartner, and Yamamoto compared conjoint-derived effects on support for immigrant naturalization against actual outcomes from Swiss municipal referendums and found close correspondence, but only for paired discrete-choice tasks (PNAS, 2015). Ranking and rating tasks, the same tasks that OLS was built to handle, degraded badly by comparison in the same study. That result was earned through design discipline and an out-of-sample comparison against real behavior, not through auditing a regression. It is also a single validated domain (referendum voting), which is a limitation worth stating plainly rather than generalizing past.

Where does the say-do gap hide inside a correct regression?

It hides in the survey responses themselves, before the estimator ever runs, which is why no amount of recomputation catches it. Hypothetical bias, the tendency for stated willingness to act to run higher than actual behavior, is a documented and unresolved problem in the stated-choice literature. A part-worth calculation has no step that detects this. It takes whatever the respondent said as ground truth and transforms it faithfully. The transformation can be flawless and the ground truth can still be wrong.

How does Subconscious validate part-worths against real behavior instead of just the math?

By running randomized experiments analyzed with discrete choice models on a simulated population and checking the result against a measured human ceiling before treating it as usable. Subconscious's methods are McFadden discrete choice, Mixed Logit, and ICLV. These are estimators, not causal methods. The causal identification comes from the randomized manipulation built into the experiment design, not from the choice of model. Flat multinomial logit assumes independence of irrelevant alternatives, which is one reason Mixed Logit and ICLV are used for preference-share and substitution questions rather than a base logit alone.

On the replication protocol, the best configuration reaches 87% of the measured human ceiling on one study: 0.832 rank correlation against a published human result, where two independent human samples reach a rank correlation of 0.959 with each other. Across the 43 studies that passed design filters, the mean is 0.73 (Causal Fidelity paper). Full results by study are on the leaderboard.

That is a validation result on studies already measured, not a guarantee for a new market. Published human studies can sit inside a model's training data, so the replication protocol is built to check for that rather than assume it away. The general validation approach is covered in the methods and validation hub.

If you're deciding whether to trust a set of part-worth utilities, don't stop at recomputing the coefficients. Ask the vendor what independent behavioral benchmark, if any, the output was checked against, and whether that check used paired discrete-choice tasks or ratings and rankings, since the PNAS results show that distinction matters more than the estimator does. If you want a second read on how a study's replication protocol was structured before you commission it, book a working session with the team.