Skip to contentScholarGate
LibraryBookshelfDeskReview StudioAssistant
Sign in
On this page
IntuitionHow it worksWhen to use itStrengths & limitationsCommon pitfallsApplicationsFrequently asked🔒 Read the full methodSourcesRelated methods
Cite this pageSpotted an issue on this page? Report or suggest a fix →
Home›Statistics›Cohen's Kappa Coefficient
Hypothesis test

Cohen's Kappa Coefficient

Cohen's Kappa Coefficient of Inter-Rater Agreement · Also known as: kappa coefficient, kappa statistic, Cohen's Kappa (Değerlendiriciler Arası Uyum)

Cohen's kappa (κ) is a statistical measure of inter-rater reliability for categorical classifications, introduced by Jacob Cohen in 1960. Unlike simple percent agreement, kappa corrects for the level of agreement that would be expected purely by chance, making it the standard metric when two raters independently assign observations to the same set of mutually exclusive categories.

ScholarGate
  1. Hypothesis test
  2. v1
  3. 2 Sources
  4. PUBLISHED
Cite this page →
Tools & resources
Download slides
Learn & explore

Read the full method

Members only

Sign in with a free account to read this section.

Sign in

Method map

The neighbourhood of related methods — select a node to explore.

Cohen's Kappa
Chi-square testFleiss' KappaMcNemar's testBland-Altman AnalysisInterrater ReliabilityIntraclass Correlation C…

When to use it

Use Cohen's kappa when exactly two raters (human coders, diagnostic instruments, or classification algorithms) have independently categorised the same set of items into the same fixed, mutually exclusive, and exhaustive categories, and you need a reliability coefficient that adjusts for chance agreement. The categories may be nominal or ordinal; for ordinal categories, weighted kappa is preferred. The minimum recommended sample size is approximately 20 items. If three or more raters are involved, use Fleiss' kappa instead. Be alert to the prevalence-and-bias paradox: when one category is far more prevalent than others, kappa can appear low even when percent agreement is high.

Strengths & limitations

Strengths
  • Corrects for chance agreement, making it more informative than simple percent agreement.
  • Applicable to nominal and (in weighted form) ordinal categorical data.
  • Widely accepted across health sciences, social sciences, and computational linguistics as the standard reliability metric.
  • Simple to interpret via the Landis & Koch benchmark scale.
Limitations
  • Designed for exactly two raters; requires Fleiss' kappa or Krippendorff's alpha for three or more.
  • Sensitive to the prevalence-and-bias paradox: high percent agreement can coexist with a low kappa when categories are very unequally distributed.
  • The Landis & Koch benchmarks are widely used but were not derived from empirical data and should not be applied mechanically.
  • Does not model rater-specific bias or systematic disagreement patterns.

Frequently asked

What is the difference between percent agreement and kappa?

Percent agreement counts all cases where the two raters chose the same category. Kappa subtracts the level of agreement expected purely by chance (given the marginal distributions) and normalises against the maximum improvement beyond chance. Two raters who agree 80% of the time may have a kappa of only 0.50 if they both have a strong tendency to choose the same dominant category.

When should I use weighted kappa instead of unweighted kappa?

Use weighted kappa when the categories have a natural order (e.g. severity grades, Likert-type ratings) and you want near misses to count as partial agreement. The quadratic weighting scheme is most common; it penalises larger disagreements more heavily. For nominal categories with no inherent order, use unweighted kappa.

My kappa is low but my percent agreement is high — what is happening?

This is the prevalence-and-bias paradox. When one category is much more frequent than others, both raters will assign it often simply because of its prevalence, inflating chance agreement and deflating kappa. Report both statistics together and consider whether the category distribution is informative about the rating challenge.

How do I handle more than two raters?

Cohen's kappa is defined for exactly two raters. For three or more raters rating the same items, use Fleiss' kappa. If raters differ across items or you need a single reliability estimate that works with any number of raters and scale types, Krippendorff's alpha is a more general alternative.

Sources

  1. Cohen, J. (1960). A Coefficient of Agreement for Nominal Scales. Educational and Psychological Measurement, 20(1), 37–46. DOI: 10.1177/001316446002000104 ↗
  2. Landis, J.R. & Koch, G.G. (1977). The Measurement of Observer Agreement for Categorical Data. Biometrics, 33(1), 159–174. DOI: 10.2307/2529310 ↗

How to cite this page

ScholarGate. (2026, June 1). Cohen's Kappa Coefficient of Inter-Rater Agreement. ScholarGate. https://scholargate.app/en/statistics/cohens-kappa

Related methods

Chi-square testFleiss' KappaMcNemar's test

Which method?

Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.

  • Chi-square testStatistics↔ compare
  • Fleiss' KappaStatistics↔ compare
  • McNemar's testStatistics↔ compare
Compare side by side →

Referenced by

Bland-Altman AnalysisFleiss' KappaInterrater ReliabilityIntraclass Correlation Coefficient

Similar methods

Fleiss' KappaInterrater ReliabilityKrippendorff's AlphaIntercoder ReliabilityIntraclass Correlation CoefficientGwet's AC1Scott's PiInter-Indexer Consistency

Related reference concepts

Interrater ReliabilityMeasurement Validity and ReliabilityCorrelation and CovarianceChi-Squared and Fisher Exact TestsEvaluation and AnnotationSensitivity

Spotted an issue on this page? Report or suggest a fix →

ScholarGate — Cohen's Kappa (Cohen's Kappa Coefficient of Inter-Rater Agreement). Retrieved 2026-07-21 from https://scholargate.app/en/statistics/cohens-kappa · Dataset: https://doi.org/10.5281/zenodo.20539026
Quick facts
Originator
Jacob Cohen
Year
1960
Family
Hypothesis test
Type
Inter-rater reliability coefficient
Raters
2
Outcome
categorical / ordinal
Parametric
No
ChanceCorrection
Yes
InterpretationScale
Landis & Koch (1977)
Related methods
Chi-square testFleiss' KappaMcNemar's test
ScholarGate

A content-first reference library for research methods — what each one is, how it works, and where it comes from.

Open data (CC-BY)

Explore

  • Library
  • Search the library…
  • Browse by field
  • Fields
  • Journey
  • Compare
  • Which method?

Reference

  • Subjects
  • Atlas
  • Glossary
  • Methodology
  • Philosophy

Your tools

  • Bookshelf
  • Desk
  • Chat

Company

  • About
  • Pricing
  • Contact
  • Suggest a method

Entries are compiled from published sources for reference. Verifying the accuracy and suitability of any information for your own use remains your responsibility.

© 2026 ScholarGate · A research-method reference library
  • Privacy
  • Cookies
  • Terms
  • Delete account