Methods · Efficiency
DEA (Data Envelopment Analysis)
DEA is a benchmarking method that compares multiple units of the same kind by how well they convert inputs into outputs. It gives every unit an efficiency score between 0 and 1; it is not a preference-ranking method.
Base method's data type: Classical
What Is the Method?
DEA answers not "which is best" but "which sits at the frontier of what is possible with the resources it has". Similar units such as branches, hospitals, universities or farms (decision-making units, DMUs) convert several inputs (staff, budget, land) into several outputs (transaction volume, number of patients treated, quantity produced). DEA measures the relative efficiency of this conversion. Its output is an efficiency score (theta) for every unit and whether that unit sits on the efficient frontier. It does not produce a "best alternative" or an "importance weight" the way TOPSIS or AHP does. Charnes, Cooper and Rhodes proposed the method in 1978 (the CCR model). DEA is used across nearly every field where multi-input, multi-output service production is compared, banking, healthcare, education and agriculture among them.
The Philosophy Behind It
The idea behind DEA is this: rather than constructing a hypothetical ideal, weave a frontier (an envelope) out of the units that genuinely exist and perform best. If, among the units being compared, no unit produces more output with the same input or the same output with less input, then that unit sits above this frontier. It is, in other words, efficient. For every unit that falls below the frontier, DEA builds a concrete answer: which of its genuinely existing peers, in what combination, could have produced the same output with less input?
This idea carries a philosophical consequence: DEA is not a preference or importance weighting, it is a benchmarking tool. Weights for every unit, meaning its input/output multipliers, are not taken from outside. The method searches, separately for every unit, for the weights that would show that unit to the greatest possible advantage. There is only one constraint: with these weights, no unit may exceed one hundred per cent. Every unit tries the best-case scenario in its own favour, in other words, but must stay honest while doing so. The result is a distinction not between "good" and "bad" but between "at the frontier" and "below the frontier"; among units at the frontier, DEA does not itself produce a ranking of superiority.
How It Works
The method proceeds through four steps.
First, separating inputs from outputs. The method classifies every measure as either an input (a resource consumed: staff, budget, land; here, less is better) or an output (a result produced: transaction volume, number of patients treated; here, more is better). This should not be confused with the "direction" label ("higher is better" / "lower is better") used in TOPSIS. In DEA, the first question is whether a measure is a resource or a result. Resource measures are then automatically treated as "lower is better", result measures as "higher is better".
Second, a separate optimisation for every unit. DEA solves a separate linear-programming problem for every unit under examination. In this problem, it searches for the input/output multipliers that would show that unit's efficiency score at its highest; these are internal, unit-specific weights. There is only one constraint: with the same multipliers, no unit's score may exceed 1, that is, one hundred per cent efficiency. This is the fundamental difference separating DEA from methods such as AHP or Entropy. In those methods, weights come from outside as a single fixed set; in DEA, every unit produces its own weights. Because the same constraint applies to everyone, no one can be accused of having "cheated" with these weights.
Third, determining the efficient frontier. Units whose score (theta) comes out at 1 sit on the efficient frontier. No weighting can make these units look more "efficient". Units below 1 are not efficient; this score shows how much of the input, theoretically, would have been needed for the same output, that is, the radial reduction ratio.
Fourth, the reference set and targets. For every inefficient unit, DEA computes which efficient units, that is, its reference or peer set, in what combination (peer weights) construct the target performance that unit could reach. It also separately reports any slack not closed by the radial reduction. Units can be ranked from largest to smallest theta, but this is an efficiency ranking, not a preference ranking. Units at the frontier (theta=1) are equally efficient as far as DEA is concerned; DEA makes no distinction among them.
The linear-programming model behind each step, the intermediate tables and the citation formats are given on the DecisionMind method page; this card carries no formulas.
How to Read the Output
A theta score of 1 does not mean "perfect". Its meaning is this: within this set of units and with this definition of inputs and outputs, no weighting could show this unit to be more efficient. This is not a fixed fact but a result tied to the set. When a new, highly efficient unit is added to the set, a previously efficient (theta=1) unit can lose its efficiency, because the new unit imposes an additional constraint on the old ones; scores can only fall, never rise. A theta below 1, say 0.833, means: this unit could have produced the same output with roughly 83 per cent of the input, relative to its peers in the reference set. The remaining gap shows that unit's potential saving.
This score is NOT several things. It is not a quality or success percentage; a small-scale or single-product unit can come out "efficient" purely from a lack of diversity in the comparison set. Nor is it a figure comparable across different DEA studies, because different studies use different unit sets and different input/output definitions. Finally, it is not a discriminating tool that answers "which is better" among several units that share theta=1.
Thus instead of writing:
"A2 is the most successful/best-managed branch"
the report should read:
"A2 sits on the efficient frontier (θ=1) with these three branches and this input-output definition; the other branches would reach this score if they could produce the same output with less input"
Data Type and Inputs
DEA works with crisp, non-negative numerical data; there must be no empty cells. Your data should contain: units (DMUs) in rows, measures in columns, a clear input/output classification for every measure, and units comparable within the same context or sector. A single measure cannot be both an input and an output. DEA does not ask for weights. It takes no importance weight from the user; it produces, separately for every unit, the input/output multipliers that show that unit to the greatest advantage. Supplying a fixed external weight set runs against the method's own logic. Alongside the base CCR model, DecisionMind offers eleven DEA members: a variable-returns-to-scale model, super-efficiency, cross-efficiency, a range-based measure, network and dynamic-network models, and an environmental model with undesirable outputs, among others. Which one fits depends on the units' scale structure (similar size or widely different) and the type of output (whether an undesirable or negative output exists). A minimum of two units and two measures, that is, one input and one output, is required. Three to twelve measures work comfortably. If the number of units is very small relative to the combined number of inputs and outputs, say three units against five measures, discriminating power is lost and nearly every unit comes out efficient. This is a data-adequacy problem, not a fault of the method.
When to Use It, When Not To
If you have comparable units of the same kind with several inputs and outputs, and the question is "which is relatively more efficient", DEA is a suitable choice. Its typical territory includes branch and bank efficiency, hospital and clinic performance, university and school efficiency, and agricultural-enterprise or supply-chain benchmarking.
The situation where DEA should not be used is when the goal is preference, not efficiency. If the question "which one should be chosen" is to be answered with subjective importance weights, DEA is the wrong tool, because DEA takes no weights and expresses no preference. If units are at very different scales, a large hospital against a small clinic, for instance, the base model, which assumes constant returns to scale, can be misleading; a variable-returns-to-scale model is needed in that case. If the number of units is very small relative to the number of measures, the result becomes unreliable.
Comparable units of the same kind, a clear input/output split, the goal is relative efficiency → DEA
The goal is a preference ranking, a choice made with subjective weights → TOPSIS, VIKOR and similar ranking methods
Units at very different scales (small and large mixed) → DEA's variable-returns-to-scale (BCC) member
The number of units is very small relative to the combined input and output count → the result is unreliable; the data or measure count should be reviewed first
What is needed is not producing or taking weights but measuring resource-to-output efficiency → DEA
Strengths
DEA's most important strength is that it reduces several inputs and outputs at once to a single score without taking a subjective weight from outside. This removes the risk of the analyst carrying personal bias into the weights. For every inefficient unit, the method gives a concrete, actionable target: it shows which peers, in what combination, should be emulated. It is independent of the units of measurement and assumes no pre-specified functional form for the production relationship, linear or exponential, for instance. It builds the frontier directly from the best performances observed in the data.
Weaknesses
Its limitations stem from the same flexibility. First, because every unit chooses its own weights, if the number of units is small relative to the number of measures, nearly every unit can find some weight combination that makes it look advantageous. In that case the unit comes out efficient and discriminating power is lost (Dyson et al., 2001). Second, the base CCR model assumes constant returns to scale. Mixing units of very different sizes produces a misleading result under this model. Banker, Charnes and Cooper's (1984) variable-scale (BCC) model solves this problem, but it requires choosing a different model. Third, a single outlier unit can determine the frontier on its own and affect every other unit's score. Fourth, DEA is deterministic; it has no statistical error term for measurement error or noise, so a data error translates directly into a score error (Cook and Seiford, 2009). Fifth, it measures only relative (set-dependent) efficiency; it says nothing about absolute or theoretical maximum efficiency.
Common Mistakes
The most common mistake is confusing inputs with outputs. Trying to maximise a cost measure as though it were an output renders the result meaningless. A second mistake is trying to feed DEA a set of importance weights for the measures from outside, say a 0.40/0.35/0.25 set taken from another method. DEA does not take weights; linear programming produces the weights itself, separately for every unit. A third mistake is using too few units with too many inputs/outputs; discriminating power drops to zero in that case and almost everyone comes out efficient. A fourth mistake is comparing units of very different size or context, a small clinic against a large hospital, for instance, using the constant-scale base model. A fifth mistake is declaring a unit with theta=1 "the best" and ignoring the other co-efficient units at the frontier. Part of this mistake is also presenting the score as a fixed fact and failing to say that it can change once a new unit is added to the set.
The governing principle is this:
A DEA score is a relative efficiency measure against the set of peers a unit is currently compared with; as units enter and leave the set the score changes too, and theta=1 means "not yet surpassed in this set", not "perfect".
Cases
Each case opens with a decision table, describes in words what the method does to it, and shows how to read the result. The first case is DecisionMind's validation example; the figures are illustrative. The other cases are illustrative constructions.
1. Banking: an efficiency comparison across three branches
A bank will compare the relative efficiency of three of its branches (A1, A2, A3). The bank has set one input (branch operating expense, "lower is better") and two outputs (transaction volume and a customer-satisfaction score, both "higher is better"). DEA does not ask for weights; the method finds, for itself, which output-to-input ratio is most advantageous for each branch.
| Branch | Transaction volume (output) | Satisfaction score (output) | Operating expense (input) |
|---|---|---|---|
| A1 | 3 | 5 | 4 |
| A2 | 5 | 3 | 2 |
| A3 | 4 | 4 | 3 |
| Input/Output | Output | Output | Input |
(DEA has no "Weight" row: the method computes, separately for every branch and within itself, the output/input multipliers that would show that branch to the greatest advantage; it takes no fixed importance weight from outside.)
The method solves a separate optimisation for every branch: it searches for the multipliers that would show that branch as efficient as possible, but with these multipliers no branch may exceed one hundred per cent. A2 is the only branch that cannot, under any such multipliers, be shown as "more efficient", and so it sits on the efficient frontier.
| Branch | Efficiency score (θ) |
|---|---|
| A2 | 1.000 |
| A3 | 0.889 |
| A1 | 0.833 |
The result reads as follows: A2 sits on the efficient frontier; it delivers the highest transaction volume at the lowest expense. A3, relative to the same peer (A2), could have produced the same output using roughly 89 per cent of its resources; for A1 this ratio is roughly 83 per cent. This does not mean A1 and A3 are "poorly managed"; it means only that, with these three branches and this input-output definition, they have not reached the resource-to-output ratio A2 shows.
The bank's hesitation is this: what happens if a fourth branch, producing much higher output with much lower expense, is added to the set? This scenario was recomputed using the same linear-programming model in Python. The result confirms that adding the fourth branch drops A2's score from 1.000 to 0.500. The reason is that the new branch displays an output-to-input ratio A2 had not previously reached; A2 is therefore no longer on the efficient frontier. This example shows that efficiency in DEA is set-dependent. When a new unit enters, even a previously efficient unit can fall below the frontier.
In the report: "Among the current three branches, A2 sits on the efficient frontier (θ=1.000); A3 (0.889) and A1 (0.833) would reach this score if they could produce the same output with less input. This ranking depends on the set of branches being compared; adding a new branch to the set can change A2's efficiency status."
Source: This is DecisionMind's DEA (CCR) engine validation example; the figures rest on the classical CCR formulation introduced by Charnes, Cooper and Rhodes (1978), but are not a table taken from the book's own page. This is an illustrative example.
2. Healthcare: a provincial health directorate's hospital efficiency comparison
A provincial health directorate will compare the efficiency of three state hospitals. The directorate has set two inputs (number of beds and annual staff expense) and two outputs (number of patients treated and patient-satisfaction score).
The method finds, for every hospital, the multipliers that show that hospital to the greatest advantage. Suppose the result shows that the hospital with the most beds does not come out efficient; a medium-sized hospital comes out on the efficient frontier instead. The reason is that this hospital uses its resources proportionally less while treating a similar number of patients.
The directorate's hesitation is this: is it fair to compare a large hospital with a small district hospital under the same constant-scale model? If the scale gap is large, say one has ten times the bed capacity of the other, the base CCR model can unfairly disadvantage the large hospital. In that case the DEA member assuming variable returns to scale should be preferred. If the directorate declares "the large hospital is inefficient" without doing so, it has ignored the scale gap.
In the report: "In the current comparison the medium-sized hospital sits on the efficient frontier; however, given the large scale gap between the hospitals, the result should be re-checked with a model assuming variable returns to scale."
3. Agriculture: a cooperative's farm efficiency assessment
An agricultural cooperative will compare the efficiency of five member farms and prioritise support distribution accordingly. The cooperative has set two inputs (land size and fertiliser-pesticide expense) and one output (annual crop yield).
The method solves a separate optimisation for every farm. Suppose the smallest-landed farm comes out on the efficient frontier because it reaches the highest yield ratio. The largest-landed farm receives a low score because it produced proportionally less crop.
The cooperative's hesitation is this: if a factor outside the farm's control, drought or pests, for instance, depressed one year's yield, how much of the gap DEA reads as "inefficiency" does this explain? DEA is deterministic; it cannot distinguish noise arising from measurement or an external shock. A single year of poor weather can make even a genuinely well-managed farm show a low score. The cooperative should look at the average across several years before cutting support based on a single year's data.
In the report: "Based on this year's data, three farms fall below the efficient frontier; however, because a single year's data may have been affected by external conditions (drought, pests), the support decision should not be made without confirmation from a multi-year average."
4. What Not to Do
The first mistake is treating operating expense in the banking example as an output rather than an input. In that case the branch that spends the most looks "most efficient" and the result becomes meaningless. A second mistake is trying to feed DEA a weight set such as 0.40/0.35/0.25, whether or not it was taken from another method, as an "importance weight". DEA does not accept such an input; it produces the weights itself, separately for every branch. A third mistake is reporting A2's score of 1.000 as "A2 is the best-run branch, the others are poorly managed". This reporting fails to state that the score is valid only for these three branches and this input-output definition. The score is a relative result that can change once a fourth branch is added.
Extensions: for different data types
DEA has 10 extensions in the library. Same decision logic, different data type: if your data is not a classical number, read the relevant data type card, then open that member.
Classical9
- DEA-BCC - Data Envelopment Analysis (BCC / VRS model)Academy card →
- DEA Cross-Efficiency - peer appraisal using cross-evaluation matrixAcademy card →
- DEA-DYNAMIC-NETWORK - Dynamic-Network DEA with Carryovers and Bad OutputsAcademy card →
- DEA-ENV - Environmental DEA with Undesirable Outputs (EEI model)Academy card →
- DEA-NETWORK - Two-Stage Network DEA with Undesirable OutputsAcademy card →
- DEA-NETWORK-SBM - Network Slacks-Based Measure DEA with Window AnalysisAcademy card →
- DEA-RAM - Range-Adjusted Measure of InefficiencyAcademy card →
- DEA-SBM - Slack-Based Measure Data Envelopment AnalysisAcademy card →
- DEA-SUPEREFF - Super-Efficiency Data Envelopment AnalysisAcademy card →
Sources
For the linear-programming model, the intermediate tables and citation formats, see the DecisionMind method page: decisionmind.app/library/dea
Charnes, A., Cooper, W. W., & Rhodes, E. (1978). Measuring the efficiency of decision making units. European Journal of Operational Research, 2(6), 429–444. DOI: 10.1016/0377-2217(78)90138-8
Banker, R. D., Charnes, A., & Cooper, W. W. (1984). Some models for estimating technical and scale inefficiencies in data envelopment analysis. Management Science, 30(9), 1078–1092. DOI: 10.1287/mnsc.30.9.1078
Dyson, R. G., Allen, R., Camanho, A. S., Podinovski, V. V., Sarrico, C. S., & Shale, E. A. (2001). Pitfalls and protocols in DEA. European Journal of Operational Research, 132(2), 245–259. DOI: 10.1016/S0377-2217(00)00149-1
Cook, W. D., & Seiford, L. M. (2009). Data envelopment analysis (DEA) – Thirty years on. European Journal of Operational Research, 192(1), 1–17. DOI: 10.1016/j.ejor.2008.01.032