Methods · Ranking
TAXONOMY (Wrocław Taxonomic Measure of Development)
TAXONOMY standardises alternatives on each criterion against their own mean and standard deviation, then converts their distance to a hypothetical "best" reference point into a single measure of development.
Base method's data type: Classical
What Is the Method?
TAXONOMY answers the question "how do I combine a group of units (regions, institutions, countries) on several indicators into a single development ranking?" Its output is a development share for every alternative and the rank that share produces. The method was proposed in 1951 by Florek, Łukaszewicz, Perkal, Steinhaus and Zubrzycki of the Wrocław school, and developed further by Hellwig in 1968 in a study of country typology. It is used in fields such as regional development comparison, institutional performance indices and supplier evaluation; official statistical bodies still apply it in regional development indices today (Roszkowska, 2024).
The Philosophy Behind It
TAXONOMY, much like TOPSIS, builds an ideal reference point and measures every alternative's distance to it. The difference lies in how it scales the data. Where TOPSIS scales the raw data by its magnitude, TAXONOMY subtracts each criterion's own mean from it and divides by its own standard deviation; this is called a standard score, or z-score. It answers the question "how many standard deviations is the alternative above or below the mean," looking at relative position rather than raw magnitude.
An analogy makes this clear. A student's exam performance can be measured not just by the raw score but by how many standard deviations it sits from the class average. A score of 80 might count as middling on an easy exam, yet the same score might count as outstanding on a difficult one. TAXONOMY converts every criterion to this relative scale, builds a single hypothetical "most developed" reference alternative, and turns every real alternative's distance to that reference, scaled against the typical spread of distances in the data, into a development share. Its philosophical consequence is that the method is compensatory: a weakness on one criterion can be balanced out by strength on another, and criteria are treated as independent of one another.
How It Works
The method proceeds through four steps.
First step, standardisation. Every criterion has its own mean subtracted and is then divided by its own standard deviation; this carries all criteria onto a unit-free, comparable scale (the z-score). For cost criteria this standardised value's sign is reversed, so that a larger value always means "good" on every criterion.
Second step, building the reference point. The highest value of every criterion in the standardised table is taken to build a hypothetical "most developed" reference alternative.
Third step, measuring the distance. Every real alternative's straight-line (Euclidean) distance to this reference point is computed.
Fourth step, computing the development share. The distance is scaled against the typical distance measure in the data (the mean of the distances plus a safety margin added on top) and subtracted from one. The alternative closest to the reference takes a share nearest to 1; alternatives are ranked from largest to smallest on this share.
The formulas behind each step, the intermediate tables and the citation formats are given on the DecisionMind method page; this card carries no formulas.
How to Read the Output
The development share shows an alternative's relative position against the ideal reference. The closer the share is to 1, the more "developed" the alternative is considered, but this is not a percentage or a probability. The share rarely comes out exactly 0 or 1, because the reference scale is derived from the typical distance distribution in the data; a very extreme alternative could theoretically even take a negative share. The share is meaningful only for this particular group of alternatives; it cannot be directly compared with a share from a different group, because the mean and standard deviation are recalculated for every group.
For this reason:
"This region's development share is 0.44, that is, 44 per cent developed"
should be written instead as:
"This region's relative distance to the ideal reference corresponds to a development share of 0.44 against the typical distance scale within this group of regions; this share is meaningful only within this group"
Data Type and Inputs
TAXONOMY works with crisp data: one number per cell. DecisionMind holds TAXONOMY only in this base form; there is no separate data-type extension. You need alternatives in rows, criteria in columns, one number per cell, no empty cells; direction information (higher is better/lower is better) for every criterion. Because a standard score is computed, every criterion must have at least a few distinct values; a criterion taking the same value across every alternative sets its standard deviation to zero and drops that criterion from the calculation.
Note: DecisionMind's TAXONOMY engine currently treats all criteria with equal weight; even if a criterion weight is entered, it is not taken into account, and the DecisionMind team is reviewing this point. A minimum of two alternatives is required, and the comfortable range for the number of criteria is three to twelve.
When to Use It, When Not To
If your aim is to combine a group of units (regions, institutions, branches) on several indicators into a single development ranking, and you accept compensation between the indicators, TAXONOMY is a suitable choice. If the number of units in your group is very small (two or three units), the standard deviation stops being a reliable estimate and the z-scores can become meaningless. If you need to give the criteria different, justified weights, another method that supports weighting (TOPSIS, WSM) should be preferred, since DecisionMind's current TAXONOMY engine does not apply it.
Within-group development comparison, relative position matters → TAXONOMY
Very few units (two or three) → standard deviation unreliable, consider another method
Different weights need to be given to the criteria → weighting methods such as TOPSIS, WSM
Not a ranking but weights are needed → AHP, BWM, SWARA (subjective), or Entropy, CRITIC (objective)
Strengths
TAXONOMY's core advantage is that it makes indicators of different scale and unit directly comparable through the z-score; this combines a large-valued indicator such as income with a small-valued indicator such as a satisfaction score in the same table on a balanced footing. The method has a broad body of applied experience in official statistics and regional development literature going back to 1951 (Roszkowska, 2024). Its computational burden is small and the result compresses into a single, easily interpreted number.
Weaknesses
Its limitations stem from the logic of the standard score. First, the result is group-dependent: were the same alternative assessed within a different group, its mean and standard deviation would change, and so would its share. Second, the standard deviation is less reliable in small samples; working with only a few alternatives can leave the result fragile (Bielak and Kowerski, 2019). Third, DecisionMind's current engine does not apply criterion weights; this is a limitation for decisions that call for differentiated, justified weighting. Fourth, the method treats criteria as independent; for indicators that influence one another, this assumption can lead to implicit double counting.
Common Mistakes
The most common mistake is marking criterion direction wrongly. If a "lower is better" indicator such as an unemployment rate is marked "higher is better," the ideal reference is built from the unit with the highest unemployment, and the result becomes meaningless. A second mistake is applying TAXONOMY with very few units (two or three) and assuming the standard deviation is reliable. A third mistake is reading the development share as a percentage or a probability; the share shows only relative position within this group. A fourth mistake is assuming and reporting that different weights were given to the criteria; the current engine treats all criteria equally.
The governing principle is this:
A TAXONOMY result depends on the mean and standard deviation of the group the alternatives belong to; if the group changes, the share changes too, and the report must state this group dependence explicitly.
Cases
Each case opens with a decision table, describes in words what the method does to it, and shows how to read the result. The first case is DecisionMind's validation example. The remaining cases are illustrative constructions.
1. Regional Development: A development comparison of three regions
A development agency will compare three regions (A1, A2, A3) on three indicators and place them in a development order. The indicators are set as the per-capita income index (higher is better), the schooling-rate index (higher is better) and the unemployment-rate index (lower is better). No separate weight has been given to the criteria; all three are treated as equally important.
| Region | Income index | Schooling index | Unemployment index |
|---|---|---|---|
| A1 | 3.0 | 5.0 | 4.0 |
| A2 | 5.0 | 3.0 | 2.0 |
| A3 | 4.0 | 4.0 | 3.0 |
| Direction | higher is better | higher is better | lower is better |
The method standardises every indicator against its own mean and standard deviation, reverses the unemployment index's sign, builds a "most developed" reference region from the standardised table, and measures every region's distance to this reference.
| Region | Development share | Rank |
|---|---|---|
| A3 | 0.4449 | 1 |
| A2 | 0.3590 | 2 |
| A1 | 0.0935 | 3 |
The result reads as follows. A3 is not, on its own, best on any single indicator; it sits in the middle on income, schooling and unemployment alike. Even so, it comes out first because it shows no pronounced weakness on any indicator. A2 holds the best values on income and unemployment, but falls to second place because it takes the lowest value on schooling. A1, though best on schooling, finishes last because it has the lowest income and the worst unemployment value.
The agency hesitates here. An earlier calculation of this example had found a different order (A2 first, A3 second); a subsequent independent Python verification revealed that the earlier calculation did not fully match the manifest's steps, and the current order (A3 first) was obtained after that correction. This shows that TAXONOMY's z-score-based structure is sensitive to computational detail; a different implementation on the same data can produce a different order.
In the report: "Among the three regions, A3 holds the highest development share (0.4449); this result has been confirmed by an independent Python computation. The method's z-score-based structure is sensitive to computational detail, so small differences can be expected when it is recomputed with different software."
Source: Florek and colleagues (1951) established the method, and Hellwig (1968) developed it further in a study of country typology. The three-region table in this card does not come from the papers' own data but from a validation example prepared to test DecisionMind's TAXONOMY engine. Note: the DecisionMind team is separately reviewing single-criterion direction-test behaviour for this class of methods that use a data-dependent reference point.
2. Librarianship: A development ranking of district public libraries
A provincial cultural affairs office will compare the development level of four district public libraries. The indicators are set as books per capita (higher is better), the annual membership growth rate (higher is better) and weekly opening hours (higher is better).
The method standardises the four libraries, builds a "most developed" reference library from the highest values, and measures each library's distance to this reference. Suppose the result places the library with the highest number of books but low membership growth second, and places the library that is balanced across all three indicators first.
The office hesitates here: since only four districts are used, the standard deviation estimate may remain crude; including all of the province's districts (ten, say) could change the order. Explaining why the library with the highest number of books comes out second should also be handled separately, from a public-relations standpoint.
In the report: "With its balanced performance, the second district library holds the highest development share; since the sample covers only four districts, the ranking may change once all of the province's districts are included."
3. Care Homes: A care-quality development comparison
A health inspection body will compare the care quality of five care homes. The indicators are set as patients per nurse (lower is better), the patient satisfaction survey score (higher is better) and the annual fall/accident rate (lower is better).
The method standardises the five care homes, builds the reference point, and measures the distances. Suppose the result places a care home with a good nurse ratio but low satisfaction in the middle, and places a care home with high satisfaction and a low accident rate first.
The body hesitates here: one care home has reported a zero accident rate. This could indicate that it is genuinely safe, or that reporting is incomplete. TAXONOMY does not question the accuracy of the data, it only processes the figures given; the data from the care home reporting a zero accident rate should therefore be audited separately.
In the report: "The care home with the highest satisfaction score and the lowest accident rate comes out first; the accuracy of the record reporting a zero accident rate should be separately confirmed by the inspection team."
4. What Not to Do
Had the unemployment index been mistakenly marked "higher is better" in the regional development case, the region with the highest unemployment would count as the ideal reference and the order would become entirely meaningless. A second error is applying TAXONOMY with only two regions and assuming the standard deviation is a reliable estimate; with two observations the standard deviation stays very crude. A third error is reporting A3's share of 0.4449 as "region A3 is 44 per cent developed"; the share shows only relative position among these three regions, and is not a percentage.
Sources
For the formulas behind each step, the intermediate tables and citation formats (BibTeX, RIS, APA), see the DecisionMind method page: decisionmind.app/library/taxonomy
Florek, K., Łukaszewicz, J., Perkal, J., Steinhaus, H., & Zubrzycki, S. (1951). Taksonomia wrocławska. Przegląd Antropologiczny, 17, 193-211. (no DOI)
Hellwig, Z. (1968). Zastosowanie metody taksonomicznej do typologicznego podziału krajów. Przegląd Statystyczny, 15, 307-327. (no DOI)
Roszkowska, E. (2024). A Comprehensive Exploration of Hellwig's Taxonomic Measure of Development and Its Modifications: A Systematic Review of Algorithms and Applications. Applied Sciences, 14(21), 10029. DOI: 10.3390/app142110029
Bielak, J., & Kowerski, M. (2019). Dynamics of Economic Development Measure. Fiftieth Anniversary of Publication of the Article by Prof. Zdzisław Hellwig. Barometr Regionalny. Analizy i Prognozy, 16(4), 153-165. DOI: 10.56583/br.56
Pareto, A. (2023). Wroclaw Taxonomy Method in Composite Indicator Construction. In Encyclopedia of Quality of Life and Well-Being Research (pp. 7879-7880). Springer. DOI: 10.1007/978-3-031-17299-1_104637