Process / pipelineLinguisticsCorpus linguistics / discourse analysisPipeline

Corpus-Assisted Discourse Studies

Also known as: CADS, Corpus-Assisted Discourse Analysis, Corpus-Based Discourse Analysis

OriginatorAlan Partington and colleaguesYear2004Sources3Related methods6

Corpus-Assisted Discourse Studies (CADS) is a mixed-methods approach that combines the quantitative power of corpus linguistics with the interpretive depth of discourse analysis to investigate how meanings, evaluations, and ideologies are constructed across large collections of text. Pioneered by Alan Partington and colleagues, CADS uses corpus techniques such as keyness, collocation, and concordancing to identify patterns no analyst could find by reading alone, then 'shunts' back to close qualitative reading to interpret what those patterns mean in their discursive context.

Key highlights

  • Disciplines analyst intuition: corpus statistics show whether a noticed pattern is typical or anecdotal across the whole dataset.
  • Surfaces patterns invisible to manual reading, such as pervasive collocations and keyness, that signal where ideology is concentrated.
  • Combines quantitative replicability with qualitative interpretive depth, strengthening the rigour and falsifiability of discourse claims.
  • Scales discourse analysis to large, comparative datasets across outlets, communities, or historical periods.

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Use CADS when you have a discourse-analytic question — about representation, evaluation, ideology, or change over time — and a body of text large enough that reading alone would risk missing or over-claiming patterns. It is especially valuable for comparing discourses (across outlets, periods, or communities) and for adding empirical weight and falsifiability to critical discourse analysis. It is less suited to very small text sets where manual analysis suffices, to questions about interaction and multimodality that lie outside the textual co-text, or to projects without a clear comparison that would make corpus statistics meaningful.

Strengths & limitations

Strengths
  • Disciplines analyst intuition: corpus statistics show whether a noticed pattern is typical or anecdotal across the whole dataset.
  • Surfaces patterns invisible to manual reading, such as pervasive collocations and keyness, that signal where ideology is concentrated.
  • Combines quantitative replicability with qualitative interpretive depth, strengthening the rigour and falsifiability of discourse claims.
  • Scales discourse analysis to large, comparative datasets across outlets, communities, or historical periods.
Limitations
  • Corpus measures locate salient lexis but cannot themselves interpret meaning; the qualitative shunt remains essential and analyst-dependent.
  • Findings are bounded by corpus design: an unbalanced or unrepresentative corpus or a poorly chosen reference corpus distorts keyness and collocation.
  • The lexical and co-text focus can miss discourse features that operate above the sentence, across modes, or in interaction.
  • Integrating two epistemologies is demanding: naive number-crunching or impressionistic reading each undermines the synergy CADS depends on.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

How is CADS different from ordinary corpus linguistics or from critical discourse analysis?

Plain corpus linguistics often stops at describing linguistic patterns (frequencies, collocations, keywords) without a sustained interpretive account of social meaning, while traditional critical discourse analysis interprets meaning richly but on small text samples that risk cherry-picking. CADS deliberately integrates the two: it uses corpus statistics to find and warrant patterns systematically, then applies discourse-analytic interpretation to explain them, shunting between the two until the evidence converges.

What does Partington mean by 'shunting'?

Shunting is the iterative back-and-forth movement between quantitative corpus work and qualitative close reading. A keyness or collocation result points the analyst to a pattern; the analyst reads the concordance lines and co-text closely to interpret it; that interpretation raises a new question, which sends them back to the corpus for another measure. CADS is therefore a cycle, not a one-way pipeline, with each pass refining and constraining the other.

Why is a reference corpus so important in CADS?

Keyness — the core technique for finding what is distinctive about a text or set of texts — is always relative: a word is 'key' only by comparison with some norm. The reference (or comparison) corpus supplies that norm. If it is poorly matched to the study corpus in register or genre, the resulting keywords reflect those mismatches rather than the discourse of interest, so choosing and justifying the reference corpus is one of the most consequential design decisions in CADS.

Sources

  1. 1.
    Partington, A., Duguid, A., & Taylor, C. (2013). Patterns and Meanings in Discourse: Theory and Practice in Corpus-Assisted Discourse Studies (CADS). John Benjamins.
    ISBN 9789027203885
  2. 2.
    Baker, P., Gabrielatos, C., KhosraviNik, M., Krzyżanowski, M., McEnery, T., & Wodak, R. (2008). A useful methodological synergy? Combining critical discourse analysis and corpus linguistics to examine discourses of refugees and asylum seekers in the UK press. Discourse & Society, 19(3), 273–306.
  3. 3.
    Partington, A. (2004). Corpora and discourse, a most congruous beast. In A. Partington, J. Morley, & L. Haarman (Eds.), Corpora and Discourse (pp. 11–20). Peter Lang.
    ISBN 9783039102488

You have read it. What now?

Cite this page

ScholarGate. (2026, June 22). Corpus-Assisted Discourse Studies. ScholarGate. https://scholargate.app/linguistics/corpus-assisted-discourse-studies