Process / pipelineText miningPipeline

Lexical Diversity — Measuring Vocabulary Richness

Also known as: lexical richness, vocabulary richness, Sözcüksel Çeşitlilik Analizi

Sources2Related methods8

Lexical diversity analysis quantifies how varied the vocabulary of a text is — how rich an author's word choice is — using measures such as the type-token ratio (TTR), MTLD, vocd-D, and Yule's K. The MTLD and vocd-D measures were validated by McCarthy and Jarvis (2010), building on earlier work by Tweedie and Baayen (1998) on the stability of lexical-richness measures.

Key highlights

  • Turns a text's vocabulary richness into a single comparable number using well-validated measures (TTR, MTLD, vocd-D, Yule's K).
  • MTLD and vocd-D are designed to be more robust to text length than the raw type-token ratio.
  • Introductory to apply and useful for both descriptive profiling and comparison across authors or groups.

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Use lexical diversity when you have tokenised text and want to describe or compare the richness of its vocabulary — across authors, documents, or groups. It suits descriptive and comparative goals and needs only a modest amount of text (around ten tokens at minimum), but for fair comparison the texts being compared must be put on equal length footing and matched for language and genre.

Strengths & limitations

Strengths
  • Turns a text's vocabulary richness into a single comparable number using well-validated measures (TTR, MTLD, vocd-D, Yule's K).
  • MTLD and vocd-D are designed to be more robust to text length than the raw type-token ratio.
  • Introductory to apply and useful for both descriptive profiling and comparison across authors or groups.
Limitations
  • Text length affects the TTR calculation; raw TTR is unreliable when comparing texts of different sizes.
  • Comparative analysis requires equalised text lengths (standard snippets or rarefaction) to be valid.
  • Language and genre shift diversity scores, so uncontrolled comparisons can be misleading.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

Why not just use the type-token ratio (TTR)?

TTR is intuitive but shrinks as a text grows, because every text eventually reuses common words. That makes raw TTR unreliable for comparing texts of different lengths. MTLD and vocd-D are designed to be far less sensitive to text length, which is why they are preferred for comparison.

How much text do I need?

The method runs on a modest amount of text — around ten tokens at minimum — but more text yields more stable estimates. For comparison, what matters most is that the texts are placed on equal length footing rather than that any single text is very long.

Can I compare the diversity of two different texts directly?

Only after putting them on equal footing. Equalise their lengths with standard-sized snippets or rarefaction, and control for language and genre, because all three factors shift diversity scores independently of vocabulary richness.

Which measure should I report?

MTLD and vocd-D are the validated, length-robust choices and are usually preferred. Yule's K describes how concentrated repeated words are, and TTR remains useful as a simple descriptive figure when texts are length-matched. Reporting more than one measure gives a fuller picture.

Sources

  1. 1.
    McCarthy, P. M. & Jarvis, S. (2010). MTLD, vocd-D, and HD-D: A validation study of sophisticated approaches to lexical diversity assessment. Behavior Research Methods, 42(2), 381-392.
  2. 2.
    Tweedie, F. J. & Baayen, R. H. (1998). How Variable May a Constant Be? Measures of Lexical Richness in Perspective. Computers and the Humanities, 32(5), 323-352.

You have read it. What now?

Cite this page

ScholarGate. (2026, June 1). Lexical Diversity. ScholarGate. https://scholargate.app/text-mining/lexical-diversity