Process / pipelineQualitativePipeline

Content Analysis — Systematic Coding of Text and Media

Also known as: İçerik Analizi, systematic content coding, quantitative content analysis

OriginatorKlaus Krippendorff (systematic formulation); roots in early 20th-century communications researchYearSystematised through Krippendorff's methodology work; 4th edition 2018Sources1Related methods82

Content analysis is a systematic research technique for reducing text, visual, or media material into coded categories so that patterns can be counted, compared, and interpreted. Formalised by Klaus Krippendorff in his widely cited methodology textbook (latest edition 2018), the method sits at the boundary of qualitative and quantitative inquiry: it imposes structured, replicable coding on inherently meaning-laden material.

Key highlights

  • Provides a transparent, auditable, and replicable process for analysing large bodies of text or media.
  • Intercoder reliability statistics give a quantitative quality check that is absent from purely interpretive methods.
  • Bridges qualitative richness and quantitative summation — categories can be counted, compared, and entered into further statistical analyses.
  • Applicable across domains and content types: news articles, interview transcripts, social media posts, policy documents, images, and broadcasts.

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Content analysis is appropriate when research questions concern the presence, frequency, or pattern of specific features in text or media, and when systematic, replicable coding across a corpus is required. It suits descriptive and exploratory purposes across health, social science, education, communication, and business research. The minimum viable corpus is approximately ten documents; below that, reliable theme production is not possible and thematic analysis of a smaller, richer dataset is preferable. The method does not require normality assumptions and works on purely textual variables.

Strengths & limitations

Strengths
  • Provides a transparent, auditable, and replicable process for analysing large bodies of text or media.
  • Intercoder reliability statistics give a quantitative quality check that is absent from purely interpretive methods.
  • Bridges qualitative richness and quantitative summation — categories can be counted, compared, and entered into further statistical analyses.
  • Applicable across domains and content types: news articles, interview transcripts, social media posts, policy documents, images, and broadcasts.
Limitations
  • Developing a reliable codebook is labour-intensive; it typically requires multiple pilot rounds and revisions.
  • Reliability above 0.70 Kappa is required — with poorly defined categories or complex content, achieving this threshold is difficult.
  • The method captures what is explicitly present in content but is less suited to latent meaning, irony, or context that lies outside the text.
  • Corpus must contain at least ten documents to sustain meaningful category frequencies; very small corpora should use thematic analysis instead.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

What is the difference between content analysis and thematic analysis?

Content analysis emphasises systematic, rule-based coding that can be checked for intercoder reliability and produces quantifiable frequencies. Thematic analysis is more interpretive, seeks latent patterns across a smaller dataset, and does not require multiple coders or a formal reliability check. When you have a large corpus and need replicable counts, content analysis is the right choice; when you have a smaller, richer dataset and want to explore meaning in depth, thematic analysis is preferable.

How do I know my coding is reliable enough?

Compute Cohen's Kappa or Krippendorff's Alpha between at least two independent coders after each pilot round. A Kappa above 0.70 is the conventional minimum for acceptability. If reliability falls below this threshold, revise the codebook — typically by adding decision rules and exemplars for the categories where coders disagreed — and recode before accepting results.

Can content analysis be done by a single researcher?

A single coder can apply the codebook, but without a second independent coder there is no way to compute intercoder reliability. Single-coder studies are therefore less credible in peer review. Where a second human coder is unavailable, some researchers use automated coding (rule-based or machine-learning) as a comparison point and report agreement with the automated system.

How large does my corpus need to be?

The minimum for meaningful content analysis is approximately ten documents. Below that, frequency-based findings are unstable and a thematic analysis of the available material is more appropriate. There is no hard upper limit — large corpora can be sampled — but every document in the sample must be coded according to the codebook.

Sources

  1. 1.
    Krippendorff, K. (2018). Content Analysis: An Introduction to Its Methodology (4th ed.). Sage.
    ISBN 978-1506395661

You have read it. What now?

Cite this page

ScholarGate. (2026, June 1). Content Analysis. ScholarGate. https://scholargate.app/qualitative/content-analysis