Methods · Subjective weighting
MACBETH (Measuring Attractiveness by a Categorical Based Evaluation Technique)
MACBETH never asks an expert for a number; it only asks which of the categories "no difference, very weak, weak, moderate, strong, very strong, extreme" the attractiveness gap between two alternatives falls into, and turns these verbal judgements into a consistent numerical scale.
Base method's data type: Classical
What Is the Method?
MACBETH is a method used to determine criterion weights or alternative values, built entirely on verbal, non-numerical judgements. This card describes the method's use for criterion weighting: its output is a weight vector summing to one. AHP also makes pairwise comparisons, but asks the expert for a number such as "how many times more important"; MACBETH instead asks only "which category does the gap between these two reference profiles fall into," and generates the number itself through a linear-programming solution. Proposed by Bana e Costa and Vansnick in 1994, the method is used in fields such as health policy, environmental assessment and public investment, where decision-makers struggle to give numerical judgements but can make verbal comparisons.
The Philosophy Behind It
MACBETH's underlying idea is that people can be inconsistent when giving ratio judgements such as "A is three times more important than B," but are more reliable in verbal comparisons such as "the gap between A and B is strong, the gap between B and C is weak." For criterion weighting, the method asks the question "how attractive is moving from the worst to the best on this criterion" through a set of reference profiles. This set consists of a neutral profile, at the lowest level on every criterion, together with profiles that each rise to the highest level on only one criterion. The expert assesses these gaps using verbal categories; MACBETH tests these categorical judgements through linear programming, searching for a numerical scale consistent with all of them.
This has one consequence: MACBETH tests consistency before generating a number. The expert's judgements may contradict one another; for instance, one judgement might state that the gap between A and B is stronger than that between C and D, while another implies the opposite. In that case the linear programme has no solution, and the method exposes this contradiction before producing any number. This makes MACBETH not merely a calculation tool but also a consistency checker.
How It Works
The method proceeds through four steps.
First, building the verbal difference matrix. A neutral reference profile (at the lowest level on every criterion) is defined, along with one profile per criterion that rises to the highest level on only that criterion. The expert assesses every pairwise gap between these profiles using one of seven categories: no difference, very weak, weak, moderate, strong, very strong, extreme. A pair on which no judgement has been given is marked "no judgement."
Second, the consistency test. A linear programme tests whether at least one numerical scale exists that agrees with these verbal judgements. Pairs marked "no difference" must take equal values; pairs marked with a difference must carry a gap consistent with the category order. If no solution is found, the judgements are contradictory and must be revised.
Third, building the basic scale. Using consistent judgements, a numerical scale is chosen that maximises the interval between the most attractive reference profile (the high anchor) and the least attractive profile (the low anchor, neutral). This is MACBETH's "basic scale," usually reported after being rescaled to run between 0 and 100.
Fourth, converting to weights. For every criterion, the gap on this scale between the profile that rises on that criterion alone and the neutral profile (the swing) is computed. These swing values are divided by their own sum to convert them into criterion weights summing to one.
The formulas behind these steps are given on the DecisionMind MACBETH method page; this card carries no formulas.
How to Read the Output
The weight shows how large a swing moving from the worst to the best on a criterion creates on a numerical scale consistent with the expert's verbal judgements; it measures the attractiveness gap the expert attributes to that transition, not the criterion's "natural importance." If two criteria's weights are close, this means the expert attributed a similar magnitude of attractiveness to the worst-to-best gap on both. MACBETH's scale is not unique: more than one numerical scale can be consistent with the judgements, and the method selects only one of them (the basic scale); for this reason, the ratios and ordering between the numbers should be relied on, not the absolute figures.
For this reason, instead of writing:
"MACBETH calculated this criterion's weight exactly"
the report should read:
"On a scale consistent with the expert's verbal judgements, this criterion's swing came out largest; if the judgements were changed, or another consistent scale chosen, the ratios would stay similar but the absolute figures could change"
Data Type and Inputs
MACBETH works with crisp (categorical) data: every cell is either a category number from 0 to 6, or a mark showing that no judgement was given. DecisionMind carries no separately registered extension member within MACBETH's own family. You need at least three reference profiles (including the neutral profile), an above-diagonal verbal-judgement matrix between these profiles, information on which profile corresponds to which criterion, and which profiles are the neutral and the most attractive. The method produces weights, it does not take weights from outside. Three to twelve criteria is typical; as the number of judgements grows, so does the risk of the consistency test finding no solution.
When to Use It, When Not To
MACBETH is a suitable choice if the expert struggles to give numerical ratios but can make verbal comparisons such as "this gap is small, that gap is large," and wants the consistency of these judgements tested from the outset. If the expert can already give numerical ratio judgements comfortably (as with AHP's nine-point scale), MACBETH's extra consistency test becomes an unnecessary step. If the number of criteria is very large, defining a separate reference profile for every criterion and collecting judgements between them becomes burdensome; methods requiring fewer judgements, such as LBWA or BWM, may be preferred in that case.
The expert can make verbal comparisons, consistency must be tested from the outset → MACBETH
The expert can give numerical ratios → AHP, BWM
Too many criteria, few judgements wanted → LBWA, BWM
Not weights but the data's own variability matters → Entropy, CRITIC
Strengths
MACBETH's principal strength is that it can work with purely verbal categories, asking the expert for no number at all; this is a more natural input form for decision-makers who are hesitant to give numerical judgements. Before producing any number, the method tests whether the judgements are mutually consistent through linear programming, and clearly exposes contradictory judgements (cycles); this allows an error to be caught from the outset rather than noticed later (Bana e Costa, De Corte and Vansnick, 2012). Over the years the method has also been supported by software such as M-MACBETH, and has built up a broad body of applied literature in fields such as health, environment and public policy (Ferreira and Santos, 2018).
Weaknesses
The method's greatest limitation is that the numerical counterpart of the seven verbal categories is chosen by the linear programme; because more than one scale can be consistent with the judgements, the selected "basic scale" is one possible solution, not the only correct one. Second, as the number of reference profiles grows, so does the number of pairwise judgements that must be collected, which can create a fatigue similar to AHP's pairwise-comparison burden. Third, when judgements turn out contradictory (when the linear programme has no solution), finding which judgement must be changed is left to the expert; the method exposes the contradiction but does not resolve it automatically. Fourth, because the numerical distance between categories (whether the gap between "moderate" and "strong" equals the gap between "strong" and "very strong") carries ordinal information rather than a cardinal distance, the result is sensitive to how the category boundaries are interpreted.
Common Mistakes
The most common mistake is mistakenly confusing one of the reference profiles with something other than the neutral (lowest) profile; the neutral profile must represent the lowest level on every criterion, otherwise the swing calculation loses its meaning. A second mistake is ignoring inconsistent judgements (the case where the linear programme has no solution) and declaring "I found a solution" after randomly changing one of the judgements; which judgement to revisit must be chosen carefully. A third mistake is presenting the basic scale MACBETH produces as an absolute reality; this scale is one of the possible scales consistent with the judgements. A fourth is confusing the criterion-weighting mode with the alternative-evaluation mode; a criterion's swing weight and an alternative's local value score are different concepts and cannot be substituted for one another.
The governing principle is this:
MACBETH's weight is the swing magnitude of a numerical scale found consistent with the expert's verbal judgements; if the judgements are consistent, the calculation is reliable, but the numbers produced are one of several possible scales.
Cases
Each case opens with a decision table, describes in words what the method does to it, and shows how to read the result. The first case is drawn from the method's founding literature; the figures are the paper's own. The remaining cases are illustrative constructions.
1. Method validation: Criterion weighting for three features of a product (Bana e Costa, De Corte and Vansnick)
In an example frequently used in the MACBETH literature, criterion weights are sought for three features of a product (colour, speed, and a black colour option). Four reference profiles are defined: a profile high only on "black," a profile high only on "colour," a profile high only on "speed," and a neutral profile high on none of them. The expert assesses the attractiveness gap between these profiles using verbal categories: for example, the "black" profile is "extremely" more attractive than the neutral profile (category 6), and "strongly" more attractive than the "colour" profile (category 4).
| Black | Colour | Speed | Neutral | |
|---|---|---|---|---|
| Black | no difference | strong | very strong | extreme |
| Colour | blank | no difference | very weak | moderate |
| Speed | blank | blank | no difference | weak |
| Neutral | blank | blank | blank | no difference |
The method first tests whether a numerical scale consistent with these judgements exists; it proves consistent. It then builds a basic scale that maximises the interval between the neutral profile and the black profile: black 7, colour 3, speed 2, neutral 0 (one possible consistent solution). Each criterion's swing (its gap from neutral) is divided by their sum to convert it into a weight.
| Criterion | Swing | Weight |
|---|---|---|
| Black | 7 | 0.583 |
| Colour | 3 | 0.250 |
| Speed | 2 | 0.167 |
The result reads as follows. The black colour option criterion takes the highest weight because it carries an "extreme" gap relative to the neutral profile. The speed criterion takes the lowest weight because it carries only a "weak" gap relative to the neutral profile. This ordering is a direct reflection of the expert's verbal judgements.
The team hesitates here: the numerical interval between the black profile's "strong" gap relative to colour and its "very strong" gap relative to speed is a single solution chosen by the linear programme. Another scale consistent with the judgements, for instance black 8, colour 3, speed 2, could equally agree with the same verbal judgements and shift the weight ratios slightly.
In the report: "The expert's verbal gaps are consistent with a scale giving the black colour option the highest weight (0.583); this scale is one of the possible scales consistent with the same verbal judgements."
Source: this is based on Bana e Costa, De Corte and Vansnick's criterion-weighting example from the MACBETH literature (1994). The same example is also carried in a 2016 Springer chapter. The weights were generated and verified by running DecisionMind's engine.
2. Healthcare: Weighting service-quality criteria for a family medicine unit
A family health centre manager must decide which area to prioritise in order to improve service quality. Three criteria apply: shortening appointment waiting time, improving patient communication, and the comfort of the physical waiting area. Four reference profiles are defined: a profile improved only on waiting time, a profile improved only on communication, a profile improved only on comfort, and a neutral profile improved on none.
The manager states that the improvement in waiting time creates a "very strong" gap relative to the neutral profile, the improvement in communication a "moderate" gap, and the improvement in comfort a "weak" gap. The method confirms these judgements are consistent and builds a scale giving waiting time the highest weight.
The manager hesitates here: rating the improvement in comfort as "weak" may present waiting-area satisfaction as low priority for patients in general, whereas this criterion may be more decisive for the elderly patient group. The manager considers weighting separately by patient group.
In the report: "The improvement in appointment waiting time carries the strongest gap relative to the neutral profile and so takes the highest weight; the weight for waiting-area comfort came out low in the overall assessment, and separate assessment for the elderly patient group is recommended."
3. Environment: Weighting criteria for a municipality's wastewater treatment investment
A municipality must decide which performance dimension to prioritise in a wastewater treatment plant investment. Three criteria apply: increasing treatment efficiency, reducing operating cost, and reducing odour and visual pollution. Four reference profiles (including neutral) are defined, and council members assess the gaps between these profiles using verbal categories.
Council members disagree over the category of the gap between treatment efficiency and odour reduction; some say "strong," others say "very strong." This disagreement does not create a contradiction, but the method produces a different basic scale under each choice.
The council hesitates here: the two different category choices noticeably change the weight of the treatment-efficiency criterion. The council decides to hold an additional round of discussion to clarify the judgement on this criterion.
In the report: "The increase in treatment efficiency takes the highest weight; because the verbal category chosen between this criterion and odour reduction was not settled among council members, the weight is treated as provisional pending an additional round of discussion."
4. What Not to Do
In the first case's example, skipping the consistency test and simply assuming a scale is wrong; the judgements may be contradictory, and this can only be established by solving the linear programme. A second error is presenting the resulting basic scale (black 7, colour 3, speed 2) as the only correct numbers; these figures are one possible scale consistent with the judgements, not the only possible one. A third error is mistakenly defining the neutral profile as representing the highest level on every criterion; in that case the swing calculation is reversed and the weights become meaningless.
Sources
For the formulas behind each step, the intermediate tables and citation formats, see the DecisionMind method page: decisionmind.app/library/macbeth
Bana e Costa, C. A., & Vansnick, J.-C. (1994). MACBETH: An interactive path towards the construction of cardinal value functions. International Transactions in Operational Research, 1(4), 489-500. DOI: 10.1016/0969-6016(94)90010-8
Bana e Costa, C. A., De Corte, J.-M., & Vansnick, J.-C. (2016). On the mathematical foundations of MACBETH. In S. Greco, M. Ehrgott, & J. R. Figueira (Eds.), Multiple Criteria Decision Analysis: State of the Art Surveys (2nd ed., Chapter 11, pp. 421-463). Springer. DOI: 10.1007/978-1-4939-3094-4_11
Bana e Costa, C. A., De Corte, J.-M., & Vansnick, J.-C. (2012). MACBETH. International Journal of Information Technology & Decision Making, 11(2), 359-387. DOI: 10.1142/s0219622012400068
Ferreira, F. A. F., & Santos, S. P. (2018). Two decades on the MACBETH approach: A bibliometric analysis. Annals of Operations Research, 296(1-2), 901-925. DOI: 10.1007/s10479-018-3083-9