Methods · Ranking
UTASTAR (Deriving an Additive Utility Function from a Reference Ranking)
UTASTAR takes a ranking the decision-maker has already given for a small group of alternatives, derives a utility function that reproduces that ranking with the least error, and applies that function to the whole list to rank every alternative.
Base method's data type: Classical
What Is the Method?
UTASTAR was developed for situations where the decision-maker struggles to state weights directly as numbers, but can say "I prefer this to that" over a small sample. It takes as input a decision matrix and a reference ranking the decision-maker has given for some of the alternatives in that matrix. Its output is an additive utility score between 0 and 1 for every alternative, and a complete ranking based on that score. The method does not ask the user for criterion weights; instead it works backwards from the ranking the user has already given. It was proposed by Siskos and Yannacopoulos in 1985 and is the most widely applied member of the UTA family; it is used in fields such as energy policy evaluation, portfolio selection and staff appraisal.
The Philosophy Behind It
Most ranking methods ask for weights first and produce the ranking afterwards. UTASTAR does the reverse. It takes a small ranking the decision-maker already knows and searches for the utility function that best explains it. This is called ordinal regression: the direction of travel is from the observed preference to the model, not from the model to the preference. This approach has a philosophical consequence: UTASTAR does not impose on the decision-maker a weight they never gave; it translates their intuition into a function and reports, through an error term, how well that function matches that intuition. Utility is additive, meaning compensatory: a weakness on one criterion can be offset by strength on another.
How It Works
The method proceeds through five steps.
First, the breakpoint grid. Every criterion is divided into several breakpoints running from its worst value to its best. The utility increase between each pair of consecutive breakpoints is defined as a variable that must be zero or positive. These variables are not yet known; the LP will find them.
Second, the error-tolerant form of the reference ranking. For every consecutive pair of alternatives in the small ranking the decision-maker gave, a difference is defined. Two error terms, one above and one below, are added to this difference to capture cases where the decision-maker's ranking does not exactly match the model.
Third, linear programming. The method finds the breakpoint increases that minimise total error, subject to constraints that preserve the order in the reference ranking. For a pair that is strictly preferred, the model requires the gap between them to exceed a small threshold; for a pair treated as indifferent, the gap is set to zero. The weights are constrained to sum to 1.
Fourth, testing the robustness of the solution. This linear program's solution may not be unique; more than one utility function can produce the same total error. UTASTAR therefore solves additional programs that separately find the lowest and highest contribution for every criterion, and averages these to produce a more stable function.
Fifth, generalisation. This stable utility function is applied not only to the alternatives in the reference set but to every alternative you hold, and the alternatives are ranked by decreasing utility score.
The formulas behind each step, the intermediate tables and the citation formats are given on the DecisionMind method page; this card carries no formulas.
How to Read the Output
The additive utility score states how good an alternative is according to the preference logic derived from the reference ranking. This score is meaningful only for this reference set and this breakpoint grid; it should not be compared against another analysis's score. The total error value (z*) carries separate information: if it is zero, the derived function reproduced the reference ranking without error; if it is greater than zero, the ranking the decision-maker gave does not exactly match an additive, compensatory model, and this is not a flaw but a finding. The method also produces a stability indicator: if the linear program has no unique solution, a criterion's contribution can range across a narrow or a wide interval; a wide interval means that criterion's weight is not well determined by this data.
Thus instead of writing:
"UTASTAR found the decision-maker's true weights"
the report should read:
"A utility function consistent with this reference ranking and this breakpoint grid was derived; some criteria's contributions remained undetermined across a wide interval"
Data Type and Inputs
UTASTAR works with crisp data, but a decision matrix on its own is not enough. You need: a decision matrix filled in rows and columns, directional information for every criterion, breakpoints (cut points) for every criterion, and, most importantly, a reference ranking the decision-maker has given for some of the alternatives. This reference ranking may also come in groups; alternatives within the same group are treated as indifferent (equally preferred). DecisionMind has no extension of UTASTAR; the base method runs on its own. The reference set should contain at least a few alternatives; too small a set leaves the linear program more flexible than it should be and makes the result fragile.
When to Use It, When Not To
UTASTAR is suitable where the decision-maker cannot give criterion weights directly as numbers but can give a reliable preference ranking over a small sample. It also works when you want to extend the implicit logic of past decisions (for example, last year's budget allocation, or a jury's shortlist) to new alternatives.
UTASTAR cannot be used if you have no reference ranking at all, or if no preference information can be obtained from the decision-maker; in that case you must derive weights directly from expert opinion (AHP, BWM, SWARA) or from the data itself (Entropy, CRITIC). UTASTAR is also unsuitable where the decision-maker will never compromise on one criterion, because additive utility is compensatory.
A small reference ranking exists, weights are not wanted → UTASTAR
No reference ranking, weights will come directly from an expert → AHP, BWM, SWARA
No reference ranking, weights will be derived from the data → Entropy, CRITIC
No compromise on one criterion → a non-compensatory elimination or dominance method
Strengths
UTASTAR's greatest strength is that it works without asking the decision-maker for an abstract weight figure; people find it hard to give numbers but are generally comfortable ranking. The method also measures its own consistency: the total error value shows how well the derived function explains the decision-maker's ranking. Its ability to learn from a small reference set and generalise to the whole list makes it a practical shortcut for groups of alternatives too large to rank by hand.
Weaknesses
How the breakpoint grid is chosen directly affects the result; too few breakpoints flattens the function, too many risks overfitting with too little data. The linear program's solution is often not unique; when Beuthe and Scannella (2001) compared different UTA variants, they showed that this solution ambiguity is a practically significant problem. If the reference set is small, the derived function can be an overfit specific to those few alternatives, and its generalisation to the wider list weakens. The method is additive and compensatory, and is therefore unsuitable for decisions where no compromise is acceptable on one criterion.
Common Mistakes
The most common mistake is choosing the number of breakpoints without justification; too few oversimplifies the function, too many overfits the data and generalises poorly to new alternatives. A second mistake is coding indifference relations in the reference ranking as though they were strict preferences; this needlessly constrains the linear program and artificially inflates the error. A third mistake is presenting a total error of zero as "a flawless model"; this only means full consistency for that reference set, not a guarantee that the same consistency will hold for other alternatives. A fourth mistake is generalising from a very small reference set to a very large list without ever questioning the gap between the two.
The governing principle is this:
The utility function UTASTAR produces is only a reflection of the reference ranking you gave and the breakpoints you chose; if the reference set is small or contested, the resulting ranking is contested too.
Cases
Each case opens with a decision table, describes in words what the method does to it, and shows how to read the result. The first case is drawn from the method's founding source. The remaining cases are illustrative constructions.
1. Transport: Choosing a Mode of Travel in Paris (Siskos and Yannacopoulos, 1985)
A transport study compares five modes for travelling from a central point to a destination: the express suburban train (RER), two different metro routes (METRO1, METRO2), the bus (BUS) and the taxi (TAXI). There are three criteria: fare (euros, lower is better), duration (minutes, lower is better) and comfort (a rank score from 0 to 3, higher is better). The decision-maker had already given the following ranking: RER best, METRO1 and METRO2 tied for second, BUS fourth, TAXI last.
| Mode | Fare (euros) | Duration (min) | Comfort |
|---|---|---|---|
| RER | 3 | 10 | 1 |
| METRO1 | 4 | 20 | 2 |
| METRO2 | 2 | 20 | 0 |
| BUS | 6 | 40 | 0 |
| TAXI | 30 | 30 | 3 |
The method searches for a utility function that reproduces the reference ranking given by these five modes with the least error. The linear program finds an error-free solution here: the total error is zero, meaning the derived function is fully consistent with the decision-maker's ranking. This function distributes marginal utility across all three criteria; the comfort criterion carries a substantial share of the total utility, while fare's share remains undetermined across a wide interval (between 0.05 and 0.94), because the linear program has no single solution for fare.
| Mode | Utility Score | Rank |
|---|---|---|
| RER | 0.920 | 1 |
| METRO1 | 0.870 | 2 |
| METRO2 | 0.870 | 2 |
| BUS | 0.511 | 4 |
| TAXI | 0.093 | 5 |
The result reads as follows. RER is neither the cheapest nor the most comfortable mode, but its balance of fare and duration puts it ahead. METRO1 and METRO2 receive equal scores in the derived function, showing that the indifference relation in the reference ranking (treating the two as equal) has been preserved. TAXI, despite its high comfort, comes last because it is the most expensive option.
The decision-maker hesitates here: the average stability indicator is a low 0.24; this shows that fare's contribution sits not on a single number in the linear program but across a wide interval, so the model cannot speak with certainty about fare's weight. This uncertainty is much narrower for the duration and comfort criteria.
In the report: "Under the given reference ranking, RER received the highest utility (0.920) and the total error is zero; however, because fare's contribution remains undetermined across a wide interval, presenting this share as a single number would be misleading."
Source: Siskos and Yannacopoulos (1985) and Siskos, Grigoroudis and Matsatsinis (2005), §2.3-2.4, Tables 7-3 and 7-6. The decision table, criteria and reference ranking are the source's own example. The utility scores are not the book's own tabulated values; they come from DecisionMind's engine re-solving the same problem through the same steps, and match the order the source reports exactly.
2. Librarianship: Budget Prioritisation at a Provincial Public Library
A provincial public library's directorate must allocate next year's acquisition budget across book categories. Three criteria have been set: annual loan count, waiting-list length, and the share met by donations. Where the share met by donations is high, less budget is allocated to purchasing; this criterion is therefore treated as "lower is better." To learn the implicit logic of last year's budget decision, the directorate gave a reference ranking over four categories: children's books the highest priority, fiction and self-help close together in a second group, and science fiction last.
The method derives the utility function that explains these four categories' reference ranking with the least error, and applies it to the ten categories the library tracks. Suppose the result gives children's books and young-adult books the highest shares, because the marginal utility placed on waiting-list length is higher than that placed on the other two criteria.
The directorate hesitates here: the reference ranking was given over only four categories, and the remaining six categories are generalised with values that fall within the range those four categories span. If a new category (say, graphic novels) does not sufficiently resemble any of the existing four, the utility assigned to it may be unreliable; the generalisation is safe only within the range the reference set covers.
In the report: "The utility function derived from last year's prioritisation of four categories has been applied to ten categories; this generalisation should be read cautiously for categories, such as graphic novels, that fall outside the reference set."
4. What Not to Do
Had the indifference relation between METRO1 and METRO2 in the same Paris table been coded as a strict preference, the linear program would have been needlessly tightened and the total error would have risen artificially. A second error is defining only two breakpoints (cheapest, most expensive) for the fare criterion; this flattens fare's marginal utility and loses the genuine differences between values. A third error is reporting a total error of zero as "this function always ranks correctly"; zero error only means full consistency with these five modes and this reference ranking, and is no guarantee that a new mode will be ranked with the same consistency.
Sources
For the formulas behind each step, the intermediate tables and citation formats (BibTeX, RIS, APA), see the DecisionMind method page: decisionmind.app/library/utastar
Siskos, Y., & Yannacopoulos, D. (1985). UTASTAR: An ordinal regression method for building additive value functions. Investigação Operacional, 5(1), 39-53. (no DOI)
Siskos, Y., Grigoroudis, E., & Matsatsinis, N. F. (2005). UTA Methods. In Multiple Criteria Decision Analysis: State of the Art Surveys (pp. 297-344). Springer. DOI: 10.1007/0-387-23081-5_8
Beuthe, M., & Scannella, G. (2001). Comparative analysis of UTA multicriteria methods. European Journal of Operational Research, 130(2), 246-262. DOI: 10.1016/s0377-2217(00)00042-4
Doumpos, M., & Zopounidis, C. (2014). The Robustness Concern in Preference Disaggregation Approaches for Decision Aiding: An Overview. In Optimization in Science and Engineering (pp. 157-173). Springer. DOI: 10.1007/978-1-4939-0808-0_8