Lexical Diversity — Measuring Vocabulary Richness
Lexical Diversity Analysis · Also known as: lexical richness, vocabulary richness, Sözcüksel Çeşitlilik Analizi
Lexical diversity analysis quantifies how varied the vocabulary of a text is — how rich an author's word choice is — using measures such as the type-token ratio (TTR), MTLD, vocd-D, and Yule's K. The MTLD and vocd-D measures were validated by McCarthy and Jarvis (2010), building on earlier work by Tweedie and Baayen (1998) on the stability of lexical-richness measures.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Use lexical diversity when you have tokenised text and want to describe or compare the richness of its vocabulary — across authors, documents, or groups. It suits descriptive and comparative goals and needs only a modest amount of text (around ten tokens at minimum), but for fair comparison the texts being compared must be put on equal length footing and matched for language and genre.
Strengths & limitations
- Turns a text's vocabulary richness into a single comparable number using well-validated measures (TTR, MTLD, vocd-D, Yule's K).
- MTLD and vocd-D are designed to be more robust to text length than the raw type-token ratio.
- Introductory to apply and useful for both descriptive profiling and comparison across authors or groups.
- Text length affects the TTR calculation; raw TTR is unreliable when comparing texts of different sizes.
- Comparative analysis requires equalised text lengths (standard snippets or rarefaction) to be valid.
- Language and genre shift diversity scores, so uncontrolled comparisons can be misleading.
Frequently asked
Why not just use the type-token ratio (TTR)?
TTR is intuitive but shrinks as a text grows, because every text eventually reuses common words. That makes raw TTR unreliable for comparing texts of different lengths. MTLD and vocd-D are designed to be far less sensitive to text length, which is why they are preferred for comparison.
How much text do I need?
The method runs on a modest amount of text — around ten tokens at minimum — but more text yields more stable estimates. For comparison, what matters most is that the texts are placed on equal length footing rather than that any single text is very long.
Can I compare the diversity of two different texts directly?
Only after putting them on equal footing. Equalise their lengths with standard-sized snippets or rarefaction, and control for language and genre, because all three factors shift diversity scores independently of vocabulary richness.
Which measure should I report?
MTLD and vocd-D are the validated, length-robust choices and are usually preferred. Yule's K describes how concentrated repeated words are, and TTR remains useful as a simple descriptive figure when texts are length-matched. Reporting more than one measure gives a fuller picture.
Sources
- McCarthy, P. M. & Jarvis, S. (2010). MTLD, vocd-D, and HD-D: A validation study of sophisticated approaches to lexical diversity assessment. Behavior Research Methods, 42(2), 381-392. DOI: 10.3758/BRM.42.2.381 ↗
- Tweedie, F. J. & Baayen, R. H. (1998). How Variable May a Constant Be? Measures of Lexical Richness in Perspective. Computers and the Humanities, 32(5), 323-352. DOI: 10.1023/A:1001749303137 ↗
How to cite this page
ScholarGate. (2026, June 1). Lexical Diversity Analysis. ScholarGate. https://scholargate.app/en/text-mining/lexical-diversity
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Sentiment AnalysisText mining↔ compare
- TF-IDFText mining↔ compare
- Topic ModelingDeep learning↔ compare