Process / pipelineLinguisticsCorpus linguisticsPipeline

Keyness Analysis

Also known as: Keyword Analysis, Corpus Keyness, Keyness Statistics

OriginatorMike ScottYear1997Sources3Related methods8

Keyness analysis identifies the words that are characteristically frequent (or infrequent) in a target corpus relative to a reference corpus, using statistical tests to measure how unexpected each word's frequency is. Introduced by Mike Scott in 1997, it answers the question 'what is this text or collection distinctively about?' and is a central technique in corpus linguistics and corpus-assisted discourse analysis for surfacing the salient vocabulary of a genre, period, author, or social group.

Key highlights

  • Provides an empirical, replicable way to identify the salient vocabulary of a corpus, reducing analyst cherry-picking.
  • Surfaces both expected and unexpected keywords, often pointing to themes the analyst would not have anticipated.
  • Scales to very large corpora and integrates naturally with concordance analysis for follow-up interpretation.
  • Supports comparison across time, source, or social group, making it powerful for discourse and diachronic studies.

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Use keyness analysis when you want a data-driven answer to what is distinctive about a text or corpus compared with a norm — to characterize a genre, track a topic across time, compare how different sources cover an issue, or find a starting point for discourse analysis. It requires a well-chosen reference corpus and is most informative for content words. It is less suitable when no appropriate reference exists, when the corpora are too small for stable frequency estimates, or when the question concerns how words are used rather than which words are salient.

Strengths & limitations

Strengths
  • Provides an empirical, replicable way to identify the salient vocabulary of a corpus, reducing analyst cherry-picking.
  • Surfaces both expected and unexpected keywords, often pointing to themes the analyst would not have anticipated.
  • Scales to very large corpora and integrates naturally with concordance analysis for follow-up interpretation.
  • Supports comparison across time, source, or social group, making it powerful for discourse and diachronic studies.
Limitations
  • Results depend heavily on the reference corpus chosen; a different benchmark yields a different keyword list.
  • With large corpora, significance tests flag trivially small differences as significant unless effect sizes are reported.
  • Keyness is computed per word type, so it can fragment multiword expressions and miss patterns above the single-word level.
  • Keywords indicate salience, not meaning; interpretation still requires returning to context via concordances.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

What statistic is used to compute keyness?

The most common is the log-likelihood ratio (G²), which compares observed and expected frequencies across the two corpora and is robust to the skewed, low-count distributions typical of word frequencies. Chi-squared was used historically but is less reliable for rare items. Increasingly, analysts report log-likelihood for significance alongside an effect-size metric such as %DIFF or log ratio to rank keywords by how large, not just how reliable, the difference is.

How do I choose a reference corpus?

The reference corpus defines 'normal', so it should be larger than the target and appropriate to the comparison. A general balanced corpus (such as a national reference corpus) suits questions about what is distinctive relative to language at large, while a contrasting subcorpus suits questions about what differs between two specific varieties. A mismatched reference produces keywords that merely reflect the mismatch, so the choice must be justified.

Why report effect size as well as significance?

In large corpora, the significance test has so much power that even negligible frequency differences become 'significant', producing inflated keyword lists dominated by high-frequency function words. Effect-size measures such as %DIFF or log ratio quantify how much more frequent a word is, allowing analysts to rank by substantive importance rather than statistical detectability, which Gabrielatos and others argue is essential for meaningful keyness analysis.

Sources

  1. 1.
    Scott, M. (1997). PC analysis of key words — and key key words. System, 25(2), 233–245.
  2. 2.
    Baker, P. (2006). Using Corpora in Discourse Analysis. Continuum.
    ISBN 9780826477248
  3. 3.
    Gabrielatos, C. (2018). Keyness analysis: Nature, metrics and techniques. In C. Taylor & A. Marchi (Eds.), Corpus Approaches to Discourse: A Critical Review (pp. 225–258). Routledge.
    ISBN 9781138895157

You have read it. What now?

Cite this page

ScholarGate. (2026, June 22). Keyness Analysis. ScholarGate. https://scholargate.app/linguistics/keyness-analysis