Data types
Linguistic
This is the data structure in which an assessment is made not with a number but with a term drawn from a predefined, ordered set of words, and in which the calculation runs on those terms without ever converting them into numbers.
Example cell: s₄, 0
What Is It?
A linguistic data structure expresses an alternative's standing on a criterion as a word chosen from an ordered term set, such as "very low, low, medium, high, very high". Each term occupies a position: in a five-term set, "medium" is third. The calculation works on these positions; the word is never converted into a triangular fuzzy number or any other figure, and the result is read back as a term: "on quality, this alternative sits between high and very high, closer to high".
What distinguishes this structure is that it preserves natural evaluative language exactly as given. When an expert says "good", nobody assumes which number that corresponds to; it is enough that "good" sits one place above "medium" and one below "very good", and the calculation uses only this ordering.
When to Use It
Linguistic data suits criteria whose natural measure is a word: service quality, user satisfaction, design aesthetics, institutional reputation, degree of fit. Experts struggle to give a number for such criteria but not to say "high" or "medium". It is valuable in group decisions where several experts share one term set, wherever the result should reach the decision-maker as a word, and wherever the scale assumption behind converting a word into a number is unwanted.
Conversely, if a criterion is genuinely measurable (price, duration, count), converting it into a word discards information; writing a measured value as "high" throws away the detail the figure carried.
Can Linguistic Data Be Built from Crisp Data?
Yes, but it takes two steps, and the second one costs information.
The first is declaring the term set and each term's boundaries before the analysis: how many terms, in what order, and which range of the measured value falls into which term. These boundaries must come from thresholds already accepted within the criterion's domain, not from arbitrary choice.
The second is assigning the measured value to a term against those boundaries. A price of 2,450 TL might fall, against declared budget thresholds, into "affordable". But this assignment erases the difference between 2,450 and 2,550; both become "affordable". A measurable criterion should therefore stay measured, and the linguistic structure should be reserved for criteria that cannot be.
What must not be done is defining the term set during or after the analysis by looking at the result. The number and order of terms must be identical for every expert and alternative, and this must be stated in the report.
A Linguistic Assessment Is Not the Same as a Word Converted into a Fuzzy Number
Words also appear in the fuzzy data structure: "good" is converted, via a declared scale, into a triangular number such as (7, 9, 10), and the calculation then works with triangles. In the linguistic structure, by contrast, "good" is never converted into any number; the calculation runs on its position in the set and its ordinal relation to neighbouring terms.
The difference is which assumption is made. Converting to a triangle assumes "this is the numerical equivalent of 'good'", and the result depends on that assumption. The linguistic approach says only "good sits above medium and below very good"; it makes fewer assumptions but, in return, treats the distances between terms as equal.
For this reason:
"The experts assessed with words, so the words should be converted into numbers and a fuzzy method applied"
should give way to:
"The experts assessed with words; the words stay within the term set, and the result is read back as a term"
This is the correct use of the linguistic structure. Both approaches are legitimate; the distinction is whether one wishes to make a scale assumption at all.
Strengths
The chief advantage of the linguistic structure is that it places no translation layer between what the expert says and what the calculation uses. Converting a word into a number requires a scale, and a scale is itself an assumption; different scales give different results, and the linguistic structure removes that assumption altogether.
Its second advantage is readability. Rather than "alternative A scored 0.73 on quality", the decision-maker sees "alternative A sits between high and very high on quality, closer to high" — the very language in which the assessment was made.
Limitations
The linguistic structure treats the distance between terms as equal: the gap between "low" and "medium" is taken to equal that between "high" and "very high". This does not hold for every criterion; the gap between the extreme terms may in reality be larger or smaller.
How many terms are used affects the result. A five-term set is coarse, a nine-term set fine-grained, and whether experts can consistently distinguish nine terms is a separate question. Where different experts use different term sets, those sets must be mapped onto one another, and that mapping is itself an assumption.
In practice, the term set is inseparable from the analysis. The uploaded file must carry a scale sheet stating which terms exist in which order; without it, the cells cannot be read. This is not a shortcoming but a consequence of the fact that a label such as "s3" cannot be understood, outside the file, as a position in some particular set.
When Are 2-Tuple and Probabilistic Linguistic Data Used?
When a calculation runs over terms, the result often falls between two of them: somewhere between "medium" and "high". The classical structure rounds this to the nearest term and loses information. The 2-tuple linguistic structure instead preserves the result as a term plus a symbolic shift: ("high", −0.3) means "a little below high", so aggregation and ranking proceed without loss. A 2-tuple cell holds a term position and a shift value.
Where an expert gives several terms, each with a probability ("60 per cent high, 40 per cent very high"), the probabilistic linguistic structure is used. It preserves the expert's opinion as distributed across terms with their probabilities, which suits aggregating different experts' distributions in group decisions.
Where an expert uses an expression such as "at least medium" or "between medium and high" without a probability, that is the domain of the hesitant linguistic structure, covered under the hesitant data type.
Common Mistakes
The most frequent mistake is converting a measurable criterion into a word. Price, duration and count should stay as numbers; writing "affordable / not affordable" instead discards the detail.
The second is failing to declare the term set in advance, or letting it vary from expert to expert. If one expert's "good" is fourth in a five-term set and another's is fifth in a seven-term set, the two cannot be compared directly.
The third is treating term positions as plain numbers: saying "medium = 3, high = 4" and averaging smuggles back, through the rear door, the very scale assumption the linguistic structure was built to avoid. Discarding the shift when rounding to a single term, rather than keeping it as a 2-tuple, likewise destroys the discriminating power a group decision needs.
The governing principle is this:
The term set is declared before the analysis and held identical across every expert and alternative; a word is never converted into a number, and the result is read back as a word.
Examples
Each example opens with a familiar, single precise figure and shows the conditions under which, and the steps by which, that same figure moves into linguistic form.
1. HR: A Performance Score of 4/5
A precise figure. On an annual appraisal form, an employee's "teamwork" score is 4/5. It looks like a number, but the 4 on the form is the numerical form of the term "good"; nothing has actually been measured.
Step 1: Is the criterion genuinely assessed in words? The decision ranks three candidates for promotion; the criterion is "teamwork". The appraiser thought of this not as a number but as "good"; the linguistic structure undoes the conversion and keeps the word as given.
Step 2: Declare the term set, then select the term. A five-term set: very poor, poor, medium, good, very good. The term set is written into the scale sheet of the upload file; the system requires this sheet for linguistic data types and rejects a file without one. The appraiser chose "good"; its position in the set is fourth.
Linguistic form. The employee stands at good (s₄) on the "teamwork" criterion. The calculation works on this position; "good" is never converted into a triangular number.
Same figure, different situation. If three appraisers say medium, good and good, the aggregated result does not fall exactly on one term; it is preserved in 2-tuple form as (good, −0.33), meaning "a little below good". Rounding to the nearest term would give "good" outright, losing the fact that one appraiser said "medium".
2. Tourism: A Hotel Rating of 4.3/5
A precise figure. On a booking platform, a hotel's average rating is 4.3. This is an average of hundreds of star ratings; it is a measured figure and stays as a number in the "overall rating" criterion.
Step 1: Is the criterion genuinely assessed in words? The decision compares three hotels; the criterion is "staff attentiveness". Guest reviews express this criterion in words: "very good", "good", "medium". Rather than converting these into numbers, they are kept as terms.
Step 2: Declare the term set, extract the distribution. A five-term set is declared. Of two hundred reviews, 45 per cent say "very good", 35 per cent say "good", and 20 per cent say "medium". Rather than a single term, the probability distribution over terms is used.
Linguistic form. The hotel stands, in probabilistic linguistic form, at {very good (0.45), good (0.35), medium (0.20)} on the "staff attentiveness" criterion. The distribution is not collapsed to a single term; two hotels reaching the same "good" average through different distributions remain visible to the decision.
Same figure, different situation. If, instead of reviews, there is a single mystery-guest assessment and the assessor said "good", the classical linguistic form suffices: good (s₄). The figure 4.3, whatever the situation, is never converted into a word; the measured criterion stays as a number.
3. Education: A Rubric Score of 2/3
A precise figure. On the "quality of argument" criterion of a project report, the rubric score is 2/3. On the rubric, 1 means "insufficient", 2 means "developing", and 3 means "sufficient"; the number is in fact the label of a term.
Step 1: Is the criterion genuinely assessed in words? The decision ranks three projects for an award; the criterion is "quality of argument". The jury does not measure this; it selects one of the three terms on the rubric. The linguistic structure preserves the term: developing (s₂).
Step 2: Declare the term set, then aggregate. A three-term set comes from the rubric and is written into the scale sheet. Three jurors rate the same project developing, sufficient, developing.
Linguistic form. The aggregated result, in 2-tuple form, is (developing, +0.33), meaning "a little above developing". The decision-maker reads the result in the rubric's own language.
Same figure, different situation. If the same project's "presentation time" criterion is measured at 12 minutes, it stays as a number; it is not converted into a term such as "long" or "short". Since a single decision matrix must hold one data type, when measured and linguistic criteria are to be used together, the choice of which structure to adopt is settled before the analysis.
4. Business: A Call Centre Response Time of 45 Seconds (A Natural Fit for the Linguistic Structure)
A precise figure. A call centre's average response time is 45 seconds. This is a measurement and stays as a number on the "response time" criterion.
Step 1. The decision compares three call-centre service providers; the criterion is "agent courtesy". This criterion is not measured; mystery-caller assessors score it in words.
Step 2. A seven-term set is declared. Four assessors rate the same provider good, very good, good, slightly good.
Linguistic form. The aggregated result, in 2-tuple form, is (good, +0.25), meaning "a little above good". The decision-maker reads the result in the word used for the assessment, not as a number. Response time stays as a number on its own, separate criterion, at 45 seconds.
5. What Not to Do
Writing a rating of 4.3 as "good", 45 seconds as "fast", or 12 minutes as "long". Each of these three measured values loses the detail it carried as a number, and the threshold used to convert it into a word has no justification. The reverse is equally wrong: treating the 4/5 on the form or the 2/3 on the rubric as though it were a plain number and averaging over it uses ordinal information as though it were scale information. The linguistic structure is for criteria that cannot be measured; the term set is declared before the analysis, and no upload is accepted without a scale sheet.
The numbers in the examples are fictional; they are not real data.
Short decision rule
A criterion that can be measured → Crisp; do not convert it into a word
The criterion's natural measure is a word, a single term → Linguistic
Aggregation falls between terms and the loss is unwanted → 2-tuple linguistic
An expert gives several terms with probabilities → Probabilistic linguistic
An expert uses a range expression such as "at least medium" without a probability → Hesitant linguistic (hesitant card)
One wishes to convert a word into a triangular number using a declared scale → Not linguistic, but fuzzy (fuzzy card)
Key sources
Zadeh, L. A. (1975). The concept of a linguistic variable and its application to approximate reasoning—I. Information Sciences, 8(3), 199–249. DOI: 10.1016/0020-0255(75)90036-5
Herrera, F., & Martínez, L. (2000). A 2-tuple fuzzy linguistic representation model for computing with words. IEEE Transactions on Fuzzy Systems, 8(6), 746–752. DOI: 10.1109/91.890332
Wei, G., Lin, R., Zhao, X., & Wang, H. (2010). Models for multiple attribute group decision making with 2-tuple linguistic assessment information. International Journal of Computational Intelligence Systems, 3(3), 315–324. DOI: 10.1080/18756891.2010.9727702
Rodríguez, R. M., Martínez, L., & Herrera, F. (2012). Hesitant fuzzy linguistic term sets for decision making. IEEE Transactions on Fuzzy Systems, 20(1), 109–119. DOI: 10.1109/TFUZZ.2011.2170076
Pang, Q., Wang, H., & Xu, Z. (2016). Probabilistic linguistic term sets in multi-attribute group decision making. Information Sciences, 369, 128–143. DOI: 10.1016/j.ins.2016.06.021