Methods · Ranking
COMET (Characteristic Objects METhod)
Before any decision is made, COMET shows the expert not the real alternatives but every possible combination of criterion levels, the "characteristic objects"; once the expert has scored these fictional profiles, the real alternatives are placed onto this ready-made preference map.
Base method's data type: Classical
What Is the Method?
COMET is a ranking method that orders alternatives into a single sequence once you hold a decision table filled with numbers and an expert assessment. Its output is a preference score between 0 and 1 for every alternative, together with the rank that score produces. Wojciech Sałabun proposed it in 2015; the method's standout feature is that the ranking does not break down even when the number or values of the alternatives change, because the assessment is carried out not on the real alternatives but on a comparison grid built in advance. COMET does not generate criterion weights; weights are used as an input to the expert's preference function.
The Philosophy Behind It
Most ranking methods compare the real alternatives in hand directly. COMET follows a different path: a few reference levels (for instance low, medium, high) are first set for each criterion. Every possible combination of these levels forms fictional profiles called "characteristic objects"; with two criteria and three levels each, nine characteristic objects result. An expert scores these nine fictional profiles by preference order, without ever seeing the real alternatives.
This distinction has a consequence: the real alternatives are placed onto this ready-made profile map afterwards, so the profiles are not positioned relative to the alternatives, but the alternatives relative to the profiles. Adding a new alternative to the table, or removing one, therefore does not change the score of the others; this feature is presented as making COMET resistant to rank reversal. The cost is that the number of characteristic objects grows rapidly as the number of criteria increases.
How It Works
The method proceeds through four steps.
First, generating the characteristic objects. A few reference levels (for instance low, medium, high) are set for each criterion. Every possible combination of these levels across all criteria is listed individually; with two criteria and three levels each, nine combinations result, and with three criteria and three levels each, twenty-seven.
Second, expert assessment. An expert, or a predetermined preference function, gives every characteristic object a preference score between 0 and 1. This step is carried out without ever seeing the real alternatives; the expert only ranks fictional profiles such as "low-low" or "high-medium."
Third, degree of membership. Every real alternative's value on each criterion is assessed by where it falls among that criterion's reference levels. If an alternative's value sits between two reference levels, it is counted as a partial "member" of both; if it matches a level exactly, it is a member of that level alone.
Fourth, the total preference score. Every alternative's final score is found by summing the preference scores of all the characteristic objects, each weighted by the alternative's degree of membership in that object. Alternatives are ranked by this score from highest to lowest.
The formulas behind each step are given on the DecisionMind method page; this card carries no formulas.
How to Read the Output
The preference score shows an alternative's position on the preference map the expert built in advance, and nothing more. A score of 0.625 does not mean "62 per cent good"; it only expresses this alternative's relative position among the characteristic objects. The score stays fixed even when new alternatives are added, as long as the reference levels and the expert's scoring do not change; this is the feature that sets COMET apart from other ranking methods.
The quality of the result depends entirely on two things: whether the reference levels cover the real data's range, and whether the expert scored the characteristic objects in a consistent order. For this reason:
"COMET found the best alternative"
should be written as:
"With these reference levels and this expert assessment, the highest preference score falls on this alternative; the result depends on the scores the expert gave to the characteristic objects"
Data Type and Inputs
COMET works with crisp data: one number per cell. DecisionMind currently holds no extension of this method for another data type.
You need alternatives in rows, criteria in columns, one number per cell and no empty cells. You also need, for each criterion, a few reference levels (at least two, usually three) and either a preference function or an expert to score all combinations of these levels. COMET does not generate weights; weights are used as an input to the preference function. As the number of criteria grows, the number of characteristic objects grows exponentially: three criteria with three levels each give twenty-seven objects, and five criteria with five levels each give more than three thousand. For this reason it is recommended that the number of criteria not exceed four or five, and that the number of reference levels be kept small.
When to Use It, When Not To
COMET is a suitable choice if you have few criteria (between two and four) and want a result that is resistant to rank reversal. It works well in expert-driven, rule-based decision problems and in repeated assessments with a small number of criteria.
If the number of criteria exceeds five, the number of characteristic objects grows rapidly and it becomes hard for the expert to score them consistently; another ranking method should then be preferred. If the preference function is inconsistent with the criteria's natural order of importance (for instance if a low-low profile is scored higher than a high-high one), the result becomes meaningless.
Few criteria, a result resistant to rank reversal is wanted → COMET
Many criteria (more than five) → direct-comparison methods such as TOPSIS or VIKOR instead of COMET
No expert assessment, only data → objective weighting plus direct-ranking methods
Not a ranking but weights are needed → AHP, BWM, SWARA (subjective) · Entropy, CRITIC (objective)
Strengths
COMET's chief advantage is its resistance to rank reversal. The assessment is carried out not on the real alternatives but on a comparison grid built in advance, so adding a new alternative does not change the score of the others (Sałabun, Ziemba and Wątróbski, 2016). It allows expert knowledge to be built directly into the model and can express interaction between criteria more flexibly than a simple weighted sum. Comparative studies have shown COMET giving results consistent with methods such as TOPSIS, VIKOR, COPRAS and PROMETHEE II (Sałabun, Wątróbski and Shekhovtsov, 2020).
Weaknesses
Its limitations mainly concern computational load. As the number of criteria grows, the number of characteristic objects grows exponentially, which confines the method to problems with few criteria. The quality of the expert assessment lies outside the method itself; an inconsistent scoring, such as giving a worse profile a higher score, leads directly to a flawed result. The reference levels must also cover the range of the real data; if these levels are chosen too narrowly, real alternatives can fall outside the grid.
Common Mistakes
The most common mistake is choosing reference levels too narrowly to cover the range of the real data; alternatives then cling to the grid's extreme points and discrimination is lost. A second mistake is the expert scoring the characteristic objects in an inconsistent order; giving a low profile a higher score than a high profile invalidates the result. A third mistake is raising the number of criteria above five while ignoring the exponential growth in the number of characteristic objects. A fourth mistake is reading the preference score as a percentage; the score only shows relative position among the characteristic objects.
The governing principle is this:
COMET's preference score is a summary of the scores the expert gave to the characteristic objects and of the chosen reference levels; if these inputs change, the result changes too.
Cases
Each case opens with a decision table, describes in words what the method does to it, and shows how to read the result.
1. Education: An examination centre's assessment of mock-exam questions (illustrative example)
An examination centre will assess three mock-exam questions on two criteria: difficulty score and discrimination score (both on a 1–5 scale, higher is better). The centre first sets three reference levels for each criterion: low (3), medium (4) and high (5). Combining these three levels across both criteria produces nine fictional question profiles. An assessment specialist scores these nine profiles, without seeing the real questions, using a weighted preference function that gives difficulty a weight of 0.60 and discrimination 0.40.
| Question | Difficulty Score | Discrimination Score |
|---|---|---|
| Q1 | 3 | 5 |
| Q2 | 5 | 3 |
| Q3 | 4 | 4 |
| Weight | 0.60 | 0.40 |
The method first generates the nine fictional profiles and gives each a score through the expert's weighted preference function. Every real question is then placed among these nine profiles according to its value on the two criteria; the closer a question sits to a profile, the more it draws from that profile's score. Finally, these shares are summed to find the question's final preference score.
| Question | Preference Score | Rank |
|---|---|---|
| Q2 | 0.625 | 1 |
| Q3 | 0.500 | 2 |
| Q1 | 0.375 | 3 |
The result reads as follows. Q2 is the question with the highest difficulty score and the lowest discrimination score; because difficulty's weight (0.60) exceeds discrimination's (0.40), Q2 comes first. Q1 sits in exactly the opposite profile: it has the highest discrimination score and the lowest difficulty score, which is why it comes last. Q3 sits in the middle on both criteria and takes a score exactly in between.
The assessment specialist hesitates here. If the weights are reversed, giving difficulty 0.40 and discrimination 0.60, the ranking flips entirely: Q1 comes first, Q3 second, and Q2 last. If the weights are made equal (0.50–0.50), all three questions take exactly the same score. This shows that the result depends entirely on the choice of weights.
In the report: "With the 0.60 weight given to difficulty, Q2 has the highest preference score; the ranking reverses once discrimination is given more weight, so the choice of weights must be justified in the report."
Source: Sałabun (2015). The figures in this case are DecisionMind's own validation example; they were cross-checked against the pymcdm library's COMET implementation on 2026-05-20, and are not taken from an application in the paper.
2. Museum Curation: A museum directorate's acquisition prioritisation
A museum directorate will assess three artworks it could acquire on two criteria: historical-significance score and condition score (both on a 1–5 scale, higher is better). The directorate sets low, medium and high reference levels for each criterion, and a curator scores these nine fictional profiles with a weighting that favours historical significance.
The method places the three artworks among these nine profiles and calculates each work's final preference score. Suppose the work with the highest historical significance but poor condition comes first: because the weight on historical significance is high, its weakness in condition is outweighed.
The curator hesitates here: acquiring a work in poor condition brings additional restoration cost. This cost never enters the model as a separate criterion; the preference score weighs only historical significance and condition, not the budget.
In the report: "The highest preference score falls on the work that favours historical significance; this work's restoration cost should be assessed as a separate constraint."
4. What Not to Do
Had the reference levels in the same question table been chosen only within a narrow range such as (4,5), Q1, with a difficulty of 3, would have fallen outside the grid and its score would have become unreliable. A second error is reading Q2's score of 0.625 as "62.5 per cent correct question"; the score only shows relative position on the preference map the expert built. A third error is the expert scoring the nine profiles inconsistently; if, for instance, a low-low profile is given a higher score than a high-high one, the entire ranking loses its meaning.
Sources
For the formulas behind each step, the intermediate tables and citation formats, see the DecisionMind method page: decisionmind.app/library/comet
Sałabun, W. (2015). The Characteristic Objects Method: A New Distance-based Approach to Multicriteria Decision-making Problems. Journal of Multi-Criteria Decision Analysis, 22(1-2), 37–50. DOI: 10.1002/mcda.1525
Sałabun, W., Ziemba, P., & Wątróbski, J. (2016). The Rank Reversals Paradox in Management Decisions: The Comparison of the AHP and COMET Methods. Smart Innovation, Systems and Technologies. DOI: 10.1007/978-3-319-39630-9_15
Sałabun, W., Wątróbski, J., & Shekhovtsov, A. (2020). Are MCDA Methods Benchmarkable? A Comparative Study of TOPSIS, VIKOR, COPRAS, and PROMETHEE II Methods. Symmetry, 12(9), 1549. DOI: 10.3390/sym12091549