Lexical Diversity — Measuring Vocabulary Richness
Also known as: lexical richness, vocabulary richness, Sözcüksel Çeşitlilik Analizi
Lexical diversity analysis quantifies how varied the vocabulary of a text is — how rich an author's word choice is — using measures such as the type-token ratio (TTR), MTLD, vocd-D, and Yule's K. The MTLD and vocd-D measures were validated by McCarthy and Jarvis (2010), building on earlier work by Tweedie and Baayen (1998) on the stability of lexical-richness measures.
Key highlights
- Turns a text's vocabulary richness into a single comparable number using well-validated measures (TTR, MTLD, vocd-D, Yule's K).
- MTLD and vocd-D are designed to be more robust to text length than the raw type-token ratio.
- Introductory to apply and useful for both descriptive profiling and comparison across authors or groups.
Intuition
This section is available to Pro members. Upgrade to Pro
How it works
This section is available to Pro members. Upgrade to Pro
When to use it
Use lexical diversity when you have tokenised text and want to describe or compare the richness of its vocabulary — across authors, documents, or groups. It suits descriptive and comparative goals and needs only a modest amount of text (around ten tokens at minimum), but for fair comparison the texts being compared must be put on equal length footing and matched for language and genre.
Strengths & limitations
- Turns a text's vocabulary richness into a single comparable number using well-validated measures (TTR, MTLD, vocd-D, Yule's K).
- MTLD and vocd-D are designed to be more robust to text length than the raw type-token ratio.
- Introductory to apply and useful for both descriptive profiling and comparison across authors or groups.
- Text length affects the TTR calculation; raw TTR is unreliable when comparing texts of different sizes.
- Comparative analysis requires equalised text lengths (standard snippets or rarefaction) to be valid.
- Language and genre shift diversity scores, so uncontrolled comparisons can be misleading.
Common pitfalls
This section is available to Pro members. Upgrade to Pro
Applications
This section is available to Pro members. Upgrade to Pro
Frequently asked
Why not just use the type-token ratio (TTR)?
TTR is intuitive but shrinks as a text grows, because every text eventually reuses common words. That makes raw TTR unreliable for comparing texts of different lengths. MTLD and vocd-D are designed to be far less sensitive to text length, which is why they are preferred for comparison.
How much text do I need?
The method runs on a modest amount of text — around ten tokens at minimum — but more text yields more stable estimates. For comparison, what matters most is that the texts are placed on equal length footing rather than that any single text is very long.
Can I compare the diversity of two different texts directly?
Only after putting them on equal footing. Equalise their lengths with standard-sized snippets or rarefaction, and control for language and genre, because all three factors shift diversity scores independently of vocabulary richness.
Which measure should I report?
MTLD and vocd-D are the validated, length-robust choices and are usually preferred. Yule's K describes how concentrated repeated words are, and TTR remains useful as a simple descriptive figure when texts are length-matched. Reporting more than one measure gives a fuller picture.
Sources
- 1.McCarthy, P. M. & Jarvis, S. (2010). MTLD, vocd-D, and HD-D: A validation study of sophisticated approaches to lexical diversity assessment. Behavior Research Methods, 42(2), 381-392.
- 2.Tweedie, F. J. & Baayen, R. H. (1998). How Variable May a Constant Be? Measures of Lexical Richness in Perspective. Computers and the Humanities, 32(5), 323-352.
You have read it. What now?
Cite this page
ScholarGate. (2026, June 1). Lexical Diversity. ScholarGate. https://scholargate.app/text-mining/lexical-diversity