Methodology · 6 min read
MaxDiff Analysis (Best-Worst Scaling)
A relative priority ranking from best-worst choices.
Jordan J. Louviere, George Woodworth, Terry N. Flynn and A.A.J. Marley · View sources ↓At a glance
- Use this when
- You need to rank attributes without every item being rated important.
- What you will work towards
- A relative priority ranking from best-worst choices.
- Bring to the reading
- A specific decision from your work and the customer evidence you have so far.
What it is
A trade-off-based survey method, developed by Jordan Louviere (with Woodworth) in the early 1990s and now a well-established discrete-choice technique in market research, that ranks how much respondents value a list of items by repeatedly asking a simpler question than a full conjoint study does: from a small set of items shown together, which is most important or valuable, and which is least. Repeating this across several different subsets, with every item appearing in multiple sets, produces a forced ranking of the whole list, feature by feature, without ever asking a respondent to name a price or rate every item on its own scale. Where Conjoint Analysis isolates the price value of individual features by testing full bundles at different price points, a heavier and more expensive study, MaxDiff answers a narrower, cheaper question: of these items, which ones actually matter most, full stop, with no price attached. It is the natural first pass before deciding which features are even worth pricing rigorously with conjoint.
When to use it
- You have a long list of candidate features, benefits, or messaging claims and need to know which ones customers actually care about, before committing conjoint-analysis or copywriting budget to all of them.
- A conjoint study is being planned but the candidate feature list is too long to test directly. Conjoint degrades past roughly 6 attributes; MaxDiff is the standard pre-study to narrow a longer list down to the genuinely contested few before the full conjoint design.
- Internal teams disagree about which of several roadmap items, benefits, or messaging pillars should be prioritised, and the debate needs a forced-ranking answer rather than a rating scale, where everything tends to score "important."
- You are choosing which value pillars to lead with in a Message Architecture, and need evidence for which claims customers rank highest, rather than which ones internal stakeholders personally favour.
- You need a fast, respondent-friendly research method (lower burden than conjoint, no price points to reason about) that still produces a genuine forced ranking rather than a Likert-scale rating exercise, where most items cluster at "somewhat important" and fail to differentiate.
How to run it
- List the candidate items. Gather the full list of features, benefits, or messaging claims under consideration, typically 10 to 20 items; MaxDiff scales well to longer lists than conjoint precisely because each individual choice task is simpler.
- Design the choice sets. Use MaxDiff survey software (Sawtooth, Qualtrics, or a comparable tool) to generate a balanced set of subsets, usually 4 to 5 items per set, so that every item appears in multiple sets and is compared against a rotating mix of the others. Do not hand-build the sets; a non-randomised design biases which items get compared against which.
- Field the choice tasks. Present each respondent with a sequence of sets (typically 10 to 15), asking in each: "Which of these is most important to you?" and "Which is least important?" Respondents answer both questions for every set, which is what produces the forced ranking, not just a single "most important" pick per set.
- Recruit a representative sample. Aim for at least 100 respondents per persona or segment being analysed separately; below roughly 75 the resulting ranking is too noisy to act on with confidence.
- Score the results. Standard MaxDiff scoring counts how often each item was picked as most important minus how often it was picked as least important, across all sets and respondents, producing a relative importance score for every item on the same scale, directly comparable across the full list.
- Segment the analysis if the sample spans personas. Run the scoring separately by persona or segment where sample size allows; a top-five ranking can differ meaningfully between a self-serve buyer and an enterprise buyer, and a single blended ranking can mask both.
- Translate the ranking into a decision. Feed a low-priority item's ranking into a decision to deprioritise it, whether that means dropping a feature from a launch, cutting a messaging pillar, or shortening the feature list a follow-up Conjoint Analysis study will test at full rigour.
Cadence & ownership
PMM typically commissions and interprets a MaxDiff study, often in partnership with a research vendor or an in-house insights team for survey design and statistical scoring, though the scoring itself is simpler than conjoint's part-worth modelling and can often be run without a dedicated data scientist. Run it as a discrete project ahead of a specific decision, narrowing a feature list before a conjoint study, prioritising message pillars before a Message Architecture refresh, rather than as a continuous programme. Budget one to two weeks end to end, materially faster than conjoint analysis, and re-run whenever the candidate list changes substantially or a major product or competitive shift makes the last ranking stale.
Example
Fictional workplace-scheduling SaaS company Shiftwell was preparing to test twelve candidate premium features in a full Conjoint Analysis study but recognised the list was too long to test rigorously in one pass. PMM ran a MaxDiff study first, with 210 respondents split across its two personas (small-team owners and mid-market operations managers), presenting each with 12 sets of four features drawn from the full list of twelve. The scoring showed a clear top four for mid-market operations managers, automated shift-conflict detection, payroll integration, compliance reporting, and multi-location scheduling, each scoring at least three times higher than the bottom four items on the list, which included several features product had assumed were strong candidates, such as a built-in team chat feature that scored second-to-last. Small-team owners' ranking looked different: mobile clock-in and simple shift-swap approval scored highest, with payroll integration ranking far lower than it had for the mid-market persona. PMM used the MaxDiff ranking to cut the twelve-item list down to the six items that scored meaningfully above the noise floor for at least one persona, and dropped the team-chat feature entirely from the roadmap discussion rather than carrying it into the more expensive conjoint study. The subsequent Conjoint Analysis study, run on the narrowed six-item list, produced cleaner, less noisy part-worth estimates than a twelve-item design would have, and the whole two-stage process cost less in combined research budget than a single overloaded twelve-item conjoint study would have.
Pitfalls
- Treating MaxDiff's importance ranking as a price signal. MaxDiff tells you which items customers rank highest in relative importance; it says nothing about how much they would pay for any of them, since no price is ever shown. Recovery: use MaxDiff to decide which items are worth pricing rigorously, then hand the narrowed list to Conjoint Analysis, Van Westendorp, or Gabor-Granger to answer the pricing question those methods are built for.
- Running MaxDiff with too short a candidate list to justify the method. MaxDiff's advantage over a simple rating scale is handling longer lists (10-plus items) that a rating scale would compress into an undifferentiated cluster of "important." For a very short list (four or five items), a simpler ranking exercise gets a comparable answer without the survey-design overhead. Recovery: reserve MaxDiff for lists of roughly 10 items or more; use a direct ranking question for anything shorter.
- Reading the blended, unsegmented ranking when personas diverge. As Shiftwell's example shows, two personas can rank the same list very differently, and a single blended score can average away a genuine, actionable difference between them. Recovery: always score by persona or segment separately when the sample allows, and treat a blended-only result as provisional until the segmented cuts have been checked.
How the ideas connect
Choose where to go next
Make it useful
Bring it back to your work.
Name one decision this guide could help you make. Write down the evidence you need, the output you would produce, and how you would know it was useful.
Check your understanding
Practise applying MaxDiff Analysis (Best-Worst Scaling) in five short scenarios.
5 practical scenarios. Choose an answer, explore the reasoning, and revisit the guide whenever you need.
Sources
- Jordan J. Louviere and George Woodworth developed the method in an unpublished 1983 (some sources date it 1990-91) working paper at the University of Alberta that was never formally published; it has since been documented extensively in market-research literature, most authoritatively in Jordan J. Louviere, Terry N. Flynn, and A.A.J. Marley, Best-Worst Scaling: Theory, Methods and Applications, Cambridge University Press (2015)
← All entries in Pricing & Packaging · Try the category quiz