Process / pipelineQualitative ResearchSystematic-categorizationPipeline

Qualitative Content Analysis

Also known as: Content Analysis, Categorical Content Analysis

OriginatorKlaus Krippendorff; refined by Margrit SchreierYear1980Sources3Related methods6

Qualitative Content Analysis (QCA) is a systematic, inductive method for analyzing textual or visual data by identifying and categorizing meaning units into content categories. Developed and formalized by Klaus Krippendorff (1980), QCA can be purely qualitative (inductive, exploratory) or combined with quantitative counting; it analyzes manifest content (explicit, surface meanings) and latent content (underlying, interpretive meanings).

Key highlights

  • Systematic and transparent: explicit coding scheme and rules enable reproducibility and inter-rater reliability assessment; ideal for rigorous analysis of large datasets.
  • Combines rigor with flexibility: can incorporate both manifest (objective) and latent (interpretive) content, bridging quantitative and qualitative approaches.
  • Suitable for diverse data types: text, images, video, social media, documents—applicable across many research contexts.
  • Code frequency quantification provides summary of emphasis and distribution, complementing qualitative interpretation with counts and proportions.

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Use QCA when you have large volumes of text or visual data and need systematic, transparent categorization. Ideal for analyzing documents (policies, clinical guidelines, media), interview transcripts, social media content, or any communicative artifact. QCA works well for research questions asking 'What themes or topics are prominent in these documents?' or 'How is a particular concept represented in media or discourse?' QCA is less suited for research seeking to generate new theory (use Grounded Theory) or explore lived experience (use phenomenology). QCA excels in applied research, policy analysis, media analysis, and organizational research where systematic documentation of content patterns informs decisions.

Strengths & limitations

Strengths
  • Systematic and transparent: explicit coding scheme and rules enable reproducibility and inter-rater reliability assessment; ideal for rigorous analysis of large datasets.
  • Combines rigor with flexibility: can incorporate both manifest (objective) and latent (interpretive) content, bridging quantitative and qualitative approaches.
  • Suitable for diverse data types: text, images, video, social media, documents—applicable across many research contexts.
  • Code frequency quantification provides summary of emphasis and distribution, complementing qualitative interpretation with counts and proportions.
Limitations
  • Risk of over-interpretation in latent content analysis; assigning meaning to underlying messages is inherently subjective and may reflect coder bias.
  • Creating and refining coding schemes is time-consuming; changing categories mid-analysis undermines consistency.
  • Does not generate theory or deep contextual understanding; content analysis describes what is present, not why or how it came to be that way.
  • Inter-rater reliability, while valuable, can be artificially inflated if coders simply follow rigid rules without understanding meaning; agreement ≠ validity.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

What is the difference between manifest and latent content analysis?

Manifest content analysis focuses on explicit, surface-level meanings: what is directly stated or shown. For example, counting the frequency of certain words or explicit statements (e.g., 'the word pain appears 47 times'). Latent content analysis interprets underlying or implicit meanings—what the content suggests, implies, or reveals about values, attitudes, or worldviews. For example, analyzing how media portrayals of obesity implicitly convey moral judgment or shame. Manifest analysis is more objective; latent analysis requires interpretive judgment. Most QCA combines both: manifest categories (what is said) with latent interpretation (what it means).

How is Qualitative Content Analysis different from Thematic Analysis?

Both identify patterns in qualitative data, but differ in structure and purpose. Thematic analysis identifies broad themes or patterns without necessarily creating a rigid coding scheme; it is more inductive and exploratory. Content analysis creates an explicit, systematically applied coding scheme with defined categories and decision rules; it is more structured and allows frequency counting. Content analysis is better for large datasets requiring consistency; thematic analysis is often faster for exploratory research. Some consider TA a type of content analysis; others view them as distinct. Use TA for exploratory pattern-finding; use QCA for systematic, transparent categorization with frequency data.

How do I ensure my coding scheme is reliable and valid?

Develop a clear definition for each category with inclusion/exclusion criteria. Have at least one second coder independently code a subset of data (10–20% minimum). Calculate inter-rater reliability: percent agreement for simple schemes, Cohen's kappa or Krippendorff's alpha for more complex ones (target >0.70). Resolve disagreements through discussion, refining category definitions if needed. Regular coder meetings throughout analysis maintain consistency. Document your coding decisions and reasoning (audit trail). Validity: ensure your categories logically relate to your research question and provide meaningful interpretation of patterns. Pilot-test your scheme on a small sample before full analysis.

Can I use QCA with both deductive and inductive coding?

Yes. Purely inductive QCA develops categories from data. Purely deductive QCA applies predetermined categories from theory or prior research. Many studies use hybrid approaches: start with deductive categories grounded in theory, then add inductive codes that emerge from the data. This can strengthen analysis by combining theoretical rigor with data-responsiveness. Be explicit about which approach (or combination) you used and justify your choice.

How many codes are typical, and when should I merge or split categories?

Initial inductive coding often produces 30–100+ codes from 15–20 interviews or documents, which are then grouped into 10–25 categories. The goal is not to minimize codes but to create a coherent hierarchy: specific codes nested within broader categories. Merge codes when they represent the same concept or produce redundancy (no purpose in keeping them separate). Split categories if they are too broad and contain heterogeneous meanings (e.g., 'barriers' might split into 'structural barriers' and 'individual barriers'). Audit your scheme: if a code appears in only one document, question whether it deserves a code or is an outlier.

Sources

You have read it. What now?

Cite this page

ScholarGate. (2026, June 4). Qualitative Content Analysis. ScholarGate. https://scholargate.app/qualitative-research/content-analysis-qualitative

Qualitative Content Analysis | ScholarGate