Cohen's Kappa Coefficient
Cohen's Kappa Coefficient of Inter-Rater Agreement · Also known as: kappa coefficient, kappa statistic, Cohen's Kappa (Değerlendiriciler Arası Uyum)
Cohen's kappa (κ) is a statistical measure of inter-rater reliability for categorical classifications, introduced by Jacob Cohen in 1960. Unlike simple percent agreement, kappa corrects for the level of agreement that would be expected purely by chance, making it the standard metric when two raters independently assign observations to the same set of mutually exclusive categories.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Use Cohen's kappa when exactly two raters (human coders, diagnostic instruments, or classification algorithms) have independently categorised the same set of items into the same fixed, mutually exclusive, and exhaustive categories, and you need a reliability coefficient that adjusts for chance agreement. The categories may be nominal or ordinal; for ordinal categories, weighted kappa is preferred. The minimum recommended sample size is approximately 20 items. If three or more raters are involved, use Fleiss' kappa instead. Be alert to the prevalence-and-bias paradox: when one category is far more prevalent than others, kappa can appear low even when percent agreement is high.
Strengths & limitations
- Corrects for chance agreement, making it more informative than simple percent agreement.
- Applicable to nominal and (in weighted form) ordinal categorical data.
- Widely accepted across health sciences, social sciences, and computational linguistics as the standard reliability metric.
- Simple to interpret via the Landis & Koch benchmark scale.
- Designed for exactly two raters; requires Fleiss' kappa or Krippendorff's alpha for three or more.
- Sensitive to the prevalence-and-bias paradox: high percent agreement can coexist with a low kappa when categories are very unequally distributed.
- The Landis & Koch benchmarks are widely used but were not derived from empirical data and should not be applied mechanically.
- Does not model rater-specific bias or systematic disagreement patterns.
Frequently asked
What is the difference between percent agreement and kappa?
Percent agreement counts all cases where the two raters chose the same category. Kappa subtracts the level of agreement expected purely by chance (given the marginal distributions) and normalises against the maximum improvement beyond chance. Two raters who agree 80% of the time may have a kappa of only 0.50 if they both have a strong tendency to choose the same dominant category.
When should I use weighted kappa instead of unweighted kappa?
Use weighted kappa when the categories have a natural order (e.g. severity grades, Likert-type ratings) and you want near misses to count as partial agreement. The quadratic weighting scheme is most common; it penalises larger disagreements more heavily. For nominal categories with no inherent order, use unweighted kappa.
My kappa is low but my percent agreement is high — what is happening?
This is the prevalence-and-bias paradox. When one category is much more frequent than others, both raters will assign it often simply because of its prevalence, inflating chance agreement and deflating kappa. Report both statistics together and consider whether the category distribution is informative about the rating challenge.
How do I handle more than two raters?
Cohen's kappa is defined for exactly two raters. For three or more raters rating the same items, use Fleiss' kappa. If raters differ across items or you need a single reliability estimate that works with any number of raters and scale types, Krippendorff's alpha is a more general alternative.
Sources
- Cohen, J. (1960). A Coefficient of Agreement for Nominal Scales. Educational and Psychological Measurement, 20(1), 37–46. DOI: 10.1177/001316446002000104 ↗
- Landis, J.R. & Koch, G.G. (1977). The Measurement of Observer Agreement for Categorical Data. Biometrics, 33(1), 159–174. DOI: 10.2307/2529310 ↗
How to cite this page
ScholarGate. (2026, June 1). Cohen's Kappa Coefficient of Inter-Rater Agreement. ScholarGate. https://scholargate.app/en/statistics/cohens-kappa
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Chi-square testStatistics↔ compare
- Fleiss' KappaStatistics↔ compare
- McNemar's testStatistics↔ compare