Methods · Subjective weighting
The Delphi Method
Delphi collects opinion from experts round by round, without them knowing one another's identity, and shows each round a summary of the previous one to track whether the views are converging on a consensus.
Base method's data type: Classical
What Is the Method?
Delphi is a method for measuring the extent to which a group of experts converges on a shared view of something such as criterion importance, level of risk or a forecast of the future, without bringing the experts together and without revealing their identities to one another. In a multi-criteria decision context, its output is a weight vector derived from the experts' average scores; it does not rank alternatives and does not evaluate them. Dalkey and Helmer (1963) developed it at the RAND Corporation to prevent experts from influencing one another while discussing face to face, as when a dominant individual's view comes to the fore; today it is used across a very wide range of fields, from criterion weighting to health policy, from technology forecasting to education planning.
The Philosophy Behind It
The idea behind Delphi is that a group's collective judgement emerges more reliably when its members think independently, without influencing one another. In a meeting room, the most senior or the loudest voice can overshadow everyone else's; Delphi prevents this by separating the experts and sharing only a statistical summary, the mean and the spread. In each round, experts see where the group is heading but not who said what; with this information they may change their view or hold their ground. The process repeats until the views converge sufficiently, or until a pre-set limit on the number of rounds is reached.
This carries a philosophical consequence: Delphi is a method of consensus, not a method of consistency. Methods such as AHP or BWM measure the internal consistency of a single expert's own scores; Delphi instead measures how closely several independent experts converge over time. If consensus is not reached, that is, if the views persistently stay scattered, Delphi's honest result is to say "expert opinion on this is not settled"; this is not a failure but valuable information about the decision itself.
How It Works
The method proceeds through a single cycle that unfolds round by round.
First, initial-round scoring. Each expert, independently of the others, gives an importance score to every criterion. At this stage experts do not see one another's scores.
Second, summary and feedback. For every criterion, the experts' mean and the extent to which their scores are dispersed (the coefficient of variation: the standard deviation as a proportion of the mean) are calculated. This summary is sent back to all experts while their identities remain hidden.
Third, checking convergence. If the coefficient of variation for a criterion falls below a pre-set threshold, that criterion is deemed to have reached consensus. Once the threshold is met for every criterion, the process stops.
Fourth, a new round. For criteria that have not met the threshold, experts are asked for a new score after seeing the group summary. An expert may hold to a previous score or update it towards the group's direction. The process returns to the second step.
Fifth, converting to weights. Once the process stops, the final round's mean scores are divided by their own total to convert them into criterion weights summing to one.
The formulas behind each step are given on the DecisionMind Delphi method page; this card carries no formulas.
How to Read the Output
A weight reflects the mean view the experts reached in the final round; it does not on its own show how solid the consensus behind that view is. A criterion's weight may be high while the coefficient of variation behind it still sits close to the threshold, meaning this weight rests on a "fragile consensus." In reading Delphi's output, it is necessary to look not only at the final weights but also at how many rounds it took and at the final round's spread; a criterion that converges quickly in few rounds and one that converges only with difficulty over many rounds may carry the same weight without being equally reliable.
Thus instead of writing:
"Delphi proved this criterion to be the most important"
the report should read:
"After three rounds, the experts gave this criterion the highest mean importance, and their views converged to a spread below the accepted threshold"
Data Type and Inputs
Delphi works with crisp (numerical) data: a score every expert gives to every criterion. DecisionMind holds no separately registered extension member within Delphi's own family. You need: at least two experts, at least two criteria, each expert's independent scores for every round, and a pre-set convergence threshold (typically a value between 0.10 and 0.30 for the coefficient of variation). Delphi produces weights; it does not ask for weights from outside. Between three and twelve criteria, and at least five to fifteen experts (so that the statistical summary is meaningful), are typical; with too few experts the mean and spread are not reliable.
When to Use It, When Not To
Delphi is a fitting choice where bringing the experts together is not possible, where even if they were brought together one person's view risks dominating the others, and where opinion can be gathered over several rounds without a pressing time constraint. Where experts can already make pairwise comparisons in a single sitting and there is no time to spare for a round-by-round feedback process, single-round methods such as AHP, BWM or FUCOM are more practical. Where a strong consensus among experts is already expected from the outset, on a technical standard that is already well settled, say, Delphi's multi-round structure wastes time needlessly.
Experts are dispersed, there is a risk of pressure, a process spread over time is workable → Delphi
Experts can make pairwise comparisons in a single round → AHP, BWM, FUCOM
What is really needed is not group consensus but the data's own variability → Entropy, CRITIC
Interaction between criteria must be measured → DEMATEL, DANP
Strengths
Delphi's most important strength is that, by concealing identity, it reduces group pressure, whether from a dominant individual, a hierarchy, or in-group reticence. It is easily applied where experts are geographically dispersed and cannot come together at the same time. Its round-by-round progression gives experts the chance to reconsider their own view after seeing others', an opportunity for learning that a one-off survey does not offer.
Weaknesses
The method's greatest limitation is its time cost: several rounds mean waiting for a response from experts at every round, and this process can take weeks; expert fatigue and the risk of participants dropping out mid-process reduce the number of participants and hence the reliability of the result (Hasson, Keeney and McKenna, 2000). Second, the choice of convergence threshold is subjective; too strict a threshold demands needlessly many rounds, too loose a threshold may count a meaningless agreement as "reached." Third, showing only the mean and the spread in feedback can push a minority view, one that might in fact be correct, towards the majority in the next round; Rowe and Wright (1999) describe this as a form of group pressure seeping into Delphi, a scaled-down version of the very problem the method sets out to remove.
Common Mistakes
The most frequent mistake is setting the convergence threshold not before the analysis begins but afterwards, by looking at the results and deciding "this many rounds is enough"; the threshold must be fixed in advance. A second mistake is running Delphi with too few experts (three, say) and assuming the statistical summary, the mean and the coefficient of variation, is reliable; measures of spread are unstable with so few experts. A third is forcibly halting the process when a round fails to reach consensus and presenting the final round's mean as "the result"; the weight of a criterion that has not converged must be flagged clearly as uncertain in the report. A fourth mistake is skipping the concealment of identity when giving feedback to experts and showing who said what; this defeats the method's whole purpose of unpressured, independent opinion.
The governing principle is this:
Delphi's weight is the final mean the experts reached; how solid a consensus that mean rests on is separate information, and it must be shown alongside it in the report.
Cases
Each case begins with a decision table, describes in words what the method does to it, and shows how to read the result. The first case is an illustrative validation example; the rest are constructed.
1. Method Validation: three experts, three criteria, a single round (DecisionMind's validation example)
Three experts score three criteria on a 0-10 scale. The scores are taken to have already converged in the first round, so a second round is not needed.
| Expert | C1 | C2 | C3 |
|---|---|---|---|
| U1 | 8 | 6 | 4 |
| U2 | 7 | 7 | 5 |
| U3 | 9 | 5 | 3 |
The method takes the column mean for every criterion: for C1, (8+7+9)/3 = 8; for C2, (6+7+5)/3 = 6; for C3, (4+5+3)/3 = 4. These means are converted into weights by dividing by their total (18).
| Criterion | Mean score | Weight |
|---|---|---|
| C1 | 8.00 | 0.444 |
| C2 | 6.00 | 0.333 |
| C3 | 4.00 | 0.222 |
The result reads as follows. C1 is the criterion all three experts scored highest, and it receives the highest weight. The three experts' scores also sit fairly close to one another (between 7 and 9 for C1, between 5 and 7 for C2, between 3 and 5 for C3), which shows that sufficient consensus was already reached in a single round.
The team hesitates here: calculating a coefficient of variation from only three experts is statistically weak; adding a fourth expert could change the means, and hence the weights. The report should therefore state the small number of experts and note that the result may be sensitive to the addition of a new expert.
In the report: "The average of three experts' scores in a single round gave C1 a weight of 0.444, C2 0.333 and C3 0.222; because the number of experts is small, these weights are sensitive to the addition of a new expert."
Source: These figures rest on the method of Dalkey and Helmer (1963); this case is an illustrative validation example, not the numerical example of the book or paper itself; it is DecisionMind's own internal consistency check.
3. Fire Service: a metropolitan fire service weights its equipment-renewal criteria
A metropolitan fire service is to decide which equipment category to prioritise on a limited budget. There are four criteria: effect on shortening response time, contribution to personnel safety, maintenance cost, and the age of existing equipment. Eighteen firefighters from different stations, their identities concealed, give scores over two rounds.
In the first round there is almost complete consensus on the personnel-safety-contribution criterion; on the response-time criterion, however, a marked disagreement emerges between stations, because different stations face different traffic conditions. By the end of the second round, the spread on the response-time criterion has narrowed but does not fall fully below the threshold.
The fire service management hesitates here: whether the response-time criterion's weight should be reported as "uncertain" or whether a third round should be run is debated. The management accepts that the different stations genuinely operate under different conditions, that this difference is not an error but a real regional variation, and decides to weight the criterion separately by region instead.
In the report: "The personnel-safety criterion converged over two rounds and received the highest weight; the response-time criterion did not converge, owing to differing conditions between stations, and has been assessed separately by region."
4. What Not to Do
In the first case's three-expert example, declaring "consensus reached" after a single round without ever calculating the coefficient of variation is wrong; even a small spread must be tested against a threshold. A second error is showing experts, in the second-round feedback, who gave which score; this undermines Delphi's whole purpose of unpressured, independent opinion. A third error is reporting a criterion that has not converged with the same reliability as one that has, using its final-round mean; criteria that have not converged must be flagged separately.
Sources
For the formulas behind each step, the intermediate tables and citation formats, see the DecisionMind method page: decisionmind.app/library/delphi
Dalkey, N., & Helmer, O. (1963). An experimental application of the Delphi method to the use of experts. Management Science, 9, 458-467. DOI: 10.1287/mnsc.9.3.458
Rowe, G., & Wright, G. (1999). The Delphi technique as a forecasting tool: issues and analysis. International Journal of Forecasting, 15, 353-375. DOI: 10.1016/S0169-2070(99)00018-7
Hasson, F., Keeney, S., & McKenna, H. (2000). Research guidelines for the Delphi survey technique. Journal of Advanced Nursing, 32, 1008-1015. DOI: 10.1046/j.1365-2648.2000.t01-1-01567.x
Okoli, C., & Pawlowski, S. D. (2004). The Delphi method as a research tool: an example, design considerations and applications. Information & Management, 42, 15-29. DOI: 10.1016/j.im.2003.11.002