Process / pipelineLinguisticsQuantitative historical linguisticsPipeline

Lexicostatistics

Also known as: Lexical Statistics, Basic Vocabulary Comparison, Cognate Percentage Method

OriginatorMorris SwadeshYear1952Sources2Related methods7

Lexicostatistics is a quantitative method in historical linguistics that gauges how closely two or more languages are genealogically related by measuring the percentage of cognates they share within a fixed list of basic, culture-neutral vocabulary — classically Morris Swadesh's 100- or 200-word list. By converting word comparisons into similarity percentages, it produces a matrix of pairwise scores from which subgroupings within a language family can be inferred. It is the statistical core that underlies glottochronology, but on its own it makes no claim about absolute dates — it speaks only to degree of relatedness.

Key highlights

  • Produces an explicit, reproducible number for relatedness, making classifications transparent and comparable across analysts and families.
  • Requires only a short, standardized word list rather than full grammars, so it scales to large surveys and under-documented languages.
  • Focuses on basic vocabulary, which is replaced slowly and resists borrowing, giving a relatively stable signal of genealogical depth.
  • Provides a ready-made distance matrix that feeds directly into clustering and computational phylogenetic methods.

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Use lexicostatistics when you have comparable basic-vocabulary lists for a set of related languages and want a quick, reproducible, quantitative first pass at their internal classification — which languages cluster as closer relatives. It is valuable for surveying large or poorly documented families, for generating subgrouping hypotheses to be tested by the comparative method, and as the input layer for distance-based phylogenetic analyses. It is not appropriate for proving that languages are related in the first place (that requires the comparative method), for languages with heavy mutual borrowing, or when reliable cognacy judgments cannot be made. Unlike glottochronology, it deliberately stops short of estimating dates.

Strengths & limitations

Strengths
  • Produces an explicit, reproducible number for relatedness, making classifications transparent and comparable across analysts and families.
  • Requires only a short, standardized word list rather than full grammars, so it scales to large surveys and under-documented languages.
  • Focuses on basic vocabulary, which is replaced slowly and resists borrowing, giving a relatively stable signal of genealogical depth.
  • Provides a ready-made distance matrix that feeds directly into clustering and computational phylogenetic methods.
Limitations
  • The result is only as good as the cognacy judgments; surface look-alikes, chance resemblances, and undetected loanwords inflate or distort percentages.
  • It measures degree of similarity, not absolute relationship, and cannot by itself demonstrate that two languages are genetically related.
  • A single percentage collapses the rich, structured evidence of regular sound correspondences into one lossy number.
  • List choice, semantic matching decisions, and how synonyms are scored all affect the percentage, so different analysts can reach different figures.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

How is lexicostatistics different from glottochronology?

Lexicostatistics is the broader practice of measuring relatedness by the percentage of shared basic-vocabulary cognates; it yields a similarity score and a classification but no dates. Glottochronology is a specific extension that assumes basic vocabulary is replaced at a roughly constant rate and uses the cognate proportion to estimate the time depth of separation. In short, all glottochronology rests on lexicostatistics, but lexicostatistics can be done without ever making a temporal claim.

Why use a fixed list like the Swadesh list instead of the whole vocabulary?

Most of a language's vocabulary is cultural and unstable — words for technology, food, and social institutions are borrowed and replaced rapidly, swamping the genealogical signal. Swadesh selected a short list of basic, universal concepts that are present in all cultures, change slowly, and resist borrowing, so that the percentage of shared items reflects inherited common ancestry rather than contact or cultural fashion. The fixed list also makes comparisons standardized and reproducible.

Can lexicostatistics prove that two languages are related?

No. A high cognate percentage is consistent with relatedness but does not establish it, because resemblances can arise from borrowing, chance, or onomatopoeia. Demonstrating genetic relationship requires the comparative method, which shows systematic, recurring sound correspondences across many forms. Lexicostatistics presupposes that cognacy can be judged reliably and is best used to subgroup languages already known to be related, not to discover relationships from scratch.

Sources

  1. 1.
    Swadesh, M. (1952). Lexico-statistic dating of prehistoric ethnic contacts. Proceedings of the American Philosophical Society, 96(4), 452–463.
  2. 2.
    Campbell, L. (2013). Historical Linguistics: An Introduction (3rd ed.). Edinburgh University Press.
    ISBN 9780748675593

You have read it. What now?

Cite this page

ScholarGate. (2026, June 22). Lexicostatistics. ScholarGate. https://scholargate.app/linguistics/lexicostatistics