Measure of Textual Lexical Diversity (MTLD)
Also known as: MTLD, Measure of Textual Lexical Diversity, Sequential TTR Factor Measure
The Measure of Textual Lexical Diversity (MTLD) is a length-robust index of vocabulary richness introduced by Philip McCarthy in 2005 and validated by McCarthy and Jarvis in 2010. Rather than computing a single ratio over the whole text, MTLD reads the text word by word, tracking a running type-token ratio, and counts how many sequential word runs are needed before the ratio repeatedly falls to a criterion value of 0.720. The mean length of those runs, computed forward and backward and averaged, is the MTLD score — and in validation studies it was the one common index that did not vary systematically with text length.
Key highlights
- Empirically the most length-stable of the common lexical diversity indices: McCarthy and Jarvis found it did not vary systematically with text length across a wide range.
- Captures the sequential unfolding of repetition rather than a single static ratio, reflecting how diversity is actually distributed through a text.
- Bidirectional averaging cancels end-of-text artifacts, giving a stable, reproducible score with no random-sampling variance.
- Conceptually transparent and deterministic — the same text always yields the same value — unlike sampling-based measures such as the original vocd-D.
Intuition
This section is available to Pro members. Upgrade to Pro
How it works
This section is available to Pro members. Upgrade to Pro
When to use it
Use MTLD when you need to compare the lexical diversity of texts that differ in length, which is the situation where raw TTR fails. It is a strong default for learner-corpus research, clinical language samples, stylometry, and any setting where samples cannot be trimmed to equal length. McCarthy and Jarvis recommend texts of at least roughly 100 tokens, below which the measure becomes unstable because too few factors are formed. For very short or highly fixed-length material, simpler indices or the hypergeometric HD-D may be preferable, and reporting MTLD alongside vocd-D or HD-D is good practice since the measures capture overlapping but not identical aspects of diversity.
Strengths & limitations
- Empirically the most length-stable of the common lexical diversity indices: McCarthy and Jarvis found it did not vary systematically with text length across a wide range.
- Captures the sequential unfolding of repetition rather than a single static ratio, reflecting how diversity is actually distributed through a text.
- Bidirectional averaging cancels end-of-text artifacts, giving a stable, reproducible score with no random-sampling variance.
- Conceptually transparent and deterministic — the same text always yields the same value — unlike sampling-based measures such as the original vocd-D.
- Becomes unstable for short texts (well under about 100 tokens) because too few complete factors are formed to estimate a reliable mean.
- The criterion value of 0.720 is an empirical default; results shift if it is altered, and its optimality is specific to the validation corpora.
- Like all type-based measures it ignores the frequency distribution of repeated words, treating a word used twice the same as a word used twenty times within a factor.
- Sensitive to tokenization, lemmatization, and the treatment of proper nouns and numerals, which must be standardized for comparisons to be valid.
Common pitfalls
This section is available to Pro members. Upgrade to Pro
Applications
This section is available to Pro members. Upgrade to Pro
Frequently asked
Why is the criterion TTR set to 0.720?
McCarthy chose 0.720 empirically as the value at which running TTR curves tend to stabilize across many texts — roughly the point where the sequential ratio settles after its steep initial decline. Setting the factor boundary there makes each factor represent a comparable amount of lexical diversity wherever it occurs in the text. The value is a tuned default rather than a theoretical constant, so it should be kept fixed at 0.720 whenever scores are to be compared across studies.
Why does MTLD compute the score forward and backward?
Processing a text from start to end almost always leaves an incomplete final factor, which is estimated by interpolation. Reading the same text from end to start places that partial factor at the opposite end and generally produces a slightly different factor count. By computing the directional MTLD in both directions and averaging, these end effects largely cancel, yielding a more stable and symmetric value that does not depend on which end of the text the analysis happens to start from.
How is MTLD different from vocd-D?
Both are length-robust diversity measures, but they work differently. vocd-D fits the empirical TTR-versus-sample-size curve to a probabilistic model using repeated random sampling, so it carries some sampling variance. MTLD instead reads the text sequentially and counts how many runs maintain a criterion TTR, making it deterministic with no sampling noise. McCarthy and Jarvis found MTLD the most length-stable of the common indices and recommend reporting it together with the hypergeometric HD-D, the sampling-free counterpart of vocd-D.
Sources
- 1.McCarthy, P. M., & Jarvis, S. (2010). MTLD, vocd-D, and HD-D: A validation study of sophisticated approaches to lexical diversity assessment. Behavior Research Methods, 42(2), 381–392.
- 2.McCarthy, P. M. (2005). An assessment of the range and usefulness of lexical diversity measures and the potential of the measure of textual, lexical diversity (MTLD) (Doctoral dissertation). University of Memphis.
You have read it. What now?
Cite this page
ScholarGate. (2026, June 22). Measure of Textual Lexical Diversity (MTLD). ScholarGate. https://scholargate.app/linguistics/mtld-lexical-diversity