Methods · Sensitivity
Kendall's W (Kendall's Coefficient of Concordance)
A method that reduces how closely the rankings given by three or more rankers (experts, methods, criteria) resemble one another to a single concordance coefficient between zero and one.
Base method's data type: Classical
What Is the Method?
Kendall's W is not a ranking method; it does not rank alternatives. If you hold rankings produced by more than one ranker for the same alternatives, orders given by several experts, orders produced by several weight scenarios, or the output of several methods, W summarises how closely all of these resemble one another in a single figure. Its output is a coefficient between zero and one: one means every ranker gave exactly the same order, and zero means the rankers show a disagreement indistinguishable from random scatter.
The method was defined by Kendall and Babington Smith in 1939, under the title "the problem of m rankings," and is the extension of Kendall's tau to more than two rankers. It is widely used to measure the concordance among the rankings produced by expert panels, juries and multi-criteria decision methods.
The Philosophy Behind It
The idea behind Kendall's W is to measure how far more than one independent opinion points in the same direction. Every ranker gives the alternatives their own order. If an alternative comes out at a similar rank across all rankers, always near the top or always near the bottom, there is strong concordance on that alternative's position. If the rankers sometimes place an alternative first and sometimes last, there is no concordance on its position.
The consequence of this philosophy is that W does not answer the question of which order is correct; it answers the question of how far the rankers agree. A high W shows that the rankers converge on something, but it does not prove which alternative is genuinely good; they could all be mistaken together. A low W shows that a genuine difference of opinion exists among the rankers, and the reason for that difference should be investigated.
How It Works
The method proceeds through four steps.
First, building the rank matrix. The rank each ranker gives to every alternative is written into a table; the rows represent the rankers and the columns represent the alternatives.
Second, column totals. The sum of the ranks every alternative receives across all rankers is calculated. An alternative that always comes out near the top gives a small total; one that always comes out near the bottom gives a large one.
Third, dispersion. How far the column totals deviate from their own average is calculated, and the sum of the squares of these deviations is taken. If the rankers agree, the totals differ greatly from one another, because everyone places the same alternatives at the top and the same ones at the bottom; if the rankers do not agree, the totals move closer to one another.
Fourth, scaling. This dispersion is divided by the maximum possible dispersion for the given number of rankers and alternatives, scaling it to between zero and one. The result is the Kendall W coefficient.
The formulas behind each step and the intermediate tables are given on the DecisionMind method page; this card carries no formulas.
How to Read the Output
The W coefficient shows how far the rankers agree as a whole. A W close to one shows that the rankers arranged the alternatives in almost the same order; a W close to zero shows that the orders are scattered independently and at random. W is not a percentage such as "what share of the rankers agree," and it does not mean that every pair of rankers gave exactly the same order; it is an overall degree of concordance.
A high W does not, on its own, validate any alternative. If all the rankers rest on the same wrong information, the same bias, or the same missing data, the result can be wrong despite a high concordance. W measures only the consistency among the rankers, not the accuracy of the ranking.
Thus instead of writing:
"Kendall's W came out at 0.78, so 78 per cent of the experts agree"
the report should read:
"Kendall's W is 0.78; this means the experts' rankings show a generally strong concordance, not that every pair of experts gave exactly the same order"
Data Type and Inputs
Kendall's W works with crisp data; its input is not a numerical decision table but a list of rankings produced by more than one ranker for the same alternatives. DecisionMind carries no separate extension of this method; if the agreement of only two rankings is to be measured, Kendall's tau is used instead.
You need rankings from at least three separate rankers for the same set of alternatives (with fewer than two rankers, Kendall's tau is sufficient). If tied alternatives exist, this must be handled separately. A minimum of three alternatives is required; as the number of alternatives and the number of rankers grow, W gives a more reliable measure of concordance.
When to Use It, When Not To
If you want to summarise, in a single figure, the overall degree of concordance among the rankings produced by more than one expert, more than one method, or more than one weight scenario, Kendall's W is exactly the right tool. In expert-panel or jury evaluations in particular, it gives a concrete answer to the question of how much the panel agrees.
There are also cases where it should not be used. If only two rankings are to be compared, W is unnecessary; Kendall's tau is sufficient and its interpretation is more direct. Reading a high W as meaning "the result is correct" is wrong; W measures only concordance, not accuracy. W is undefined with two rankers; at least three are required.
Want to summarise the concordance of three or more rankers in a single figure → Kendall's W
Only two rankings are to be compared → Kendall's tau
Not concordance but which pair of alternatives is the source of disagreement is of interest → Kendall's tau computed separately between pairs of rankers
Robustness to randomness in the alternative set is of interest → Bootstrap Resampling
Fewer than two rankers → Kendall's W is undefined
Strengths
Kendall's W's most important strength is that it reduces the concordance of more than two rankers to a single interpretable figure; this is far more practical than reading through many pairwise Kendall tau comparisons one by one. It makes no distributional assumption, resting purely on rank information. It can be applied directly in many settings, such as expert-panel evaluation, jury assessment or the comparison of multiple methods (Legendre, 2005).
Weaknesses
Its limitations arise from this same design. First, W is a measure of concordance, not of accuracy; if all the rankers rest on the same wrong assumption, W comes out high even though the result is wrong. Second, W does not show which pair of alternatives carries the disagreement; it gives only an overall degree of concordance, and for detail Kendall's tau must be computed separately between pairs of rankers. Third, tied alternatives require a correction in the classic formula. Fourth, with few alternatives and few rankers, W can take only a handful of discrete values, and a single ranker behaving differently can shift W considerably (Legendre, 2005, examines this sensitivity in detail using ecological data).
Common Mistakes
The most common mistake is reading a high W as meaning "the result is correct." The correct statement is that a high W shows only that the rankers gave each other close rankings; they could all be mistaken together.
A second mistake is presenting W as a percentage; the sentence "W is 0.78, so 78 per cent of the experts agree" is wrong, and the correct statement is that W is an overall degree of concordance. A third mistake is reporting only "concordance is weak" when a low W comes out, without investigating which alternative is the source of disagreement; behind a low W there is usually a disagreement concentrated on one or two alternatives. A fourth mistake is trying to calculate W for a comparison with only two rankers; the correct tool in that case is Kendall's tau.
The governing principle is this:
Kendall's W measures the degree of concordance among rankers, not the accuracy of the ranking; a high W tells you not which alternative is genuinely good, but how much the rankers agree on the matter.
Cases
Each case opens with a rank matrix, describes in words what W does to it, and shows how to read the result. The first case is DecisionMind's validation example; the figures were recalculated and verified in Python. The remaining cases are illustrative constructions.
1. Illustrative example: Three rankers, three alternatives (DecisionMind validation example)
This example is not a case from the literature; it is a small example constructed to make W's logic traceable by hand, used in DecisionMind's own validation test. Three rankers (R1, R2, R3) have ranked three alternatives (A1, A2, A3).
| Ranker | A1 | A2 | A3 |
|---|---|---|---|
| R1 | 1 | 2 | 3 |
| R2 | 1 | 3 | 2 |
| R3 | 1 | 2 | 3 |
The method first takes the column total for each alternative: for A1, 1+1+1=3; for A2, 2+3+2=7; for A3, 3+2+3=8. The average of these three totals is 6. The sum of the squared deviations of the totals from this average is (3-6)²+(7-6)²+(8-6)²=9+1+4=14. This figure is expressed as a ratio of the maximum possible dispersion for three rankers and three alternatives (216), giving W = 12×14/216 = 0.778.
The result reads as follows. All three rankers placed A1 first; there is complete concordance on this point. The disagreement lies only in the positions of A2 and A3: R1 and R3 place A2 second and A3 third, while R2 does the exact opposite. This single-source disagreement pulls W down from one, complete concordance, to 0.778, strong but not complete concordance.
The analyst's hesitation is this: a W of 0.778 does not mean "seventy-eight per cent of the rankers agree"; all three rankers are in complete agreement on one point, A1's first place, and split two against one on only one point, the order of A2 and A3. The report should also show this distinction behind W's single figure.
In the report: "Among the three rankers, Kendall's W is 0.778; there is complete concordance on A1's first place, and the disagreement lies only in the order of A2 and A3, in a two-against-one split."
Source: the DecisionMind KENDALL-W manifest, validation example. The manifest's J.expected_primary field was produced by the repository's own audit validator; it is not an example taken directly from a paper. The method's definition is Kendall and Babington Smith (1939).
2. Examination Centre: Three assessors' concordance on essay rankings
An examination centre had three independent assessors each rank six essays from a written examination. The centre wanted to check the concordance among the assessors with Kendall's W before releasing the marks.
The method extracted the column totals from the three assessors' rankings and calculated W. Suppose W came out high, at 0.85; this shows that the three assessors ranked the essays in a largely similar order.
The centre's hesitation: a high W does not prove the assessment is fair; all three assessors could be scoring by the same writing pattern or the same bias. While treating the high concordance as a positive sign, the centre noted that the assessment criteria themselves should also be reviewed separately.
In the report: "Among the three assessors' essay rankings, Kendall's W is 0.85; there is strong concordance among the assessors, and this concordance shows the consistency, not the accuracy, of the assessment criteria."
3. Maritime: Four experts' concordance on port investment prioritisation
A port operator had four independent experts each rank five investment projects: quay extension, crane renewal, a digital tracking system, a fuel facility, and an environmental treatment plant. The operator wanted to measure the degree of concordance among the experts before proceeding to a decision.
The method extracted the column totals from the four experts' rankings and calculated W. Suppose W came out at a moderate-to-low value, 0.42; this shows a marked difference of opinion among the experts.
The operator's hesitation: a low W does not, on its own, say which project is disputed. Comparing the four experts' rankings pairwise with Kendall's tau, the operator found that the disagreement is concentrated mainly on the order between the digital tracking system and the fuel facility; all four experts agree on the priority of the quay extension.
In the report: "Among the four experts, Kendall's W is 0.42; overall concordance is weak. Detailed comparison showed that the disagreement is concentrated mainly on the order between the digital tracking system and the fuel facility, while there is complete concordance on the priority of the quay extension."
4. What Not to Do
Reporting the illustrative example's W value of 0.778 as "seventy-eight per cent of the three assessors agree" is wrong; W is not a percentage but an overall degree of concordance. A second error is seeing a high W and declaring "the result is correct" without questioning the assessment criteria at all; W measures only consistency, not accuracy. A third error is closing the matter with "the experts could not agree" when a low W comes out, without investigating which alternative is the source of disagreement; for detail, Kendall's tau should be examined separately between pairs of rankers.
Sources
For the formulas behind each step, the intermediate tables and citation formats, see the DecisionMind method page: decisionmind.app/library/kendall-w
Kendall, M. G., & Babington Smith, B. (1939). The problem of m rankings. The Annals of Mathematical Statistics, 10(3), 275-287. DOI: 10.1214/aoms/1177732186
Legendre, P. (2005). Species associations: The Kendall coefficient of concordance revisited. Journal of Agricultural, Biological, and Environmental Statistics, 10(2), 226-245. DOI: 10.1198/108571105x46642
Kendall, M. G. (1938). A new measure of rank correlation. Biometrika, 30(1-2), 81-93. DOI: 10.2307/2332226
García-Cascales, M. S., & Lamata, M. T. (2012). On rank reversal and TOPSIS method. Mathematical and Computer Modelling, 56(5-6), 123-132. DOI: 10.1016/j.mcm.2011.12.022