Skip to contentScholarGate
LibraryBookshelfDeskReview StudioAssistant
Sign in
On this page
IntuitionHow it worksWhen to use itStrengths & limitationsCommon pitfallsApplicationsFrequently asked🔒 Read the full methodSourcesRelated methods
Cite this pageSpotted an issue on this page? Report or suggest a fix →
Home›Psychometrics›Interrater Reliability (Cohen's κ and ICC)
Latent structure

Interrater Reliability (Cohen's κ and ICC)

Also known as: inter-rater reliability, interrater agreement, rater agreement, Değerlendiriciler Arası Güvenilirlik (Cohen's κ, ICC), Cohen's kappa and ICC

Interrater reliability quantifies the degree to which two or more independent raters produce consistent scores when evaluating the same individuals or products. The family encompasses Cohen's kappa, introduced in 1960 for categorical judgments, and the Intraclass Correlation Coefficient (ICC) for continuous ratings, together spanning most measurement scenarios encountered in behavioral, health, and educational research.

ScholarGate
  1. Latent structure
  2. v1
  3. 2 Sources
  4. PUBLISHED
Cite this page →
Tools & resources
Download slides
Learn & explore

Read the full method

Members only

Sign in with a free account to read this section.

Sign in

Method map

The neighbourhood of related methods — select a node to explore.

Interrater Reliability
Bland-Altman AnalysisCohen's KappaCronbach's AlphaFleiss' KappaG-TheoryIntraclass Correlation C…Robust Test-Retest Relia…Test-Retest Reliability

When to use it

Interrater reliability methods are appropriate whenever two or more raters independently score the same units and you need to demonstrate that the ratings are reproducible. Cohen's kappa is used when the outcome is categorical or ordinal (diagnoses, severity grades, qualitative codes in content analysis). ICC is used when ratings are continuous or finely graded (physical measurements, scale scores, observer ratings on Likert-type scales). A minimum of roughly 20 subjects is needed for a stable estimate, though power analyses for ICC typically recommend 30 or more. The design must ensure independent rating: raters should not confer, consult each other's scores, or rate in a fixed order that induces carryover effects.

Strengths & limitations

Strengths
  • Directly addresses the reproducibility of measurement, a prerequisite for valid inference in any study that relies on human-generated scores.
  • Kappa adjusts for chance agreement, giving a more conservative and honest index than raw percent agreement.
  • ICC models the full variance structure of rater-by-subject designs, distinguishing consistency (covariation) from absolute agreement (systematic bias), so researchers can choose the index that matches their substantive question.
  • Both statistics are widely reported and understood across clinical, educational, and behavioral research communities.
Limitations
  • Kappa is sensitive to the prevalence of categories: when one category is very common or very rare the marginal totals constrain the maximum achievable kappa, making comparisons across studies with different base rates misleading.
  • ICC requires specification of the correct variance-component model (random vs. mixed, consistency vs. absolute agreement); an incorrect model choice produces a statistic that does not match the research question.
  • Both statistics describe agreement in the sample; they do not diagnose which rater is more accurate or why disagreements occur.

Frequently asked

When should I use kappa rather than percent agreement?

Always, unless you have a specific reason to ignore chance. Percent agreement can be misleadingly high when one category dominates: if 90 % of cases are in category A, two raters randomly assigning A to everything will agree 81 % of the time with no skill. Kappa adjusts for this and gives the proportion of agreement above what chance predicts.

Which ICC model should I choose?

The choice depends on your design. If your raters are a random sample from a larger pool of possible raters and you want results to generalise to other raters, use a two-way random-effects model. If the same fixed raters will always be used in practice, use a two-way mixed-effects model. Then decide between consistency (acceptable when a constant rater offset does not matter) and absolute agreement (required when systematic differences between raters would affect decisions).

How many raters and subjects do I need?

For a stable ICC estimate, a common recommendation is at least 30 subjects and at least two raters, though more raters and subjects improve precision. Published power-analysis methods for ICC (e.g., Bonett 2002) can guide sample-size planning for a target confidence-interval width.

Can I combine ICC and kappa in the same study?

Yes, and it is sometimes appropriate. If your instrument has both continuous subscale scores and categorical overall classifications, report ICC for the continuous scores and kappa (or weighted kappa) for the categorical classifications. Report each statistic with its confidence interval and the model or weighting scheme used.

Sources

  1. Cohen, J. (1960). A Coefficient of Agreement for Nominal Scales. Educational and Psychological Measurement, 20(1), 37–46. DOI: 10.1177/001316446002000104 ↗
  2. Koo, T.K. & Li, M.Y. (2016). A Guideline of Selecting and Reporting Intraclass Correlation Coefficients for Reliability Research. Journal of Chiropractic Medicine, 15(2), 155–163. DOI: 10.1016/j.jcm.2016.02.012 ↗

How to cite this page

ScholarGate. (2026, June 1). Interrater Reliability (Cohen's κ and ICC). ScholarGate. https://scholargate.app/en/psychometrics/interrater-reliability

Related methods

Bland-Altman AnalysisCohen's KappaCronbach's AlphaFleiss' KappaG-TheoryIntraclass Correlation Coefficient

Which method?

Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.

  • Bland-Altman AnalysisStatistics↔ compare
  • Cohen's KappaStatistics↔ compare
  • Cronbach's AlphaStatistics↔ compare
  • Fleiss' KappaStatistics↔ compare
  • G-TheoryPsychometrics↔ compare
  • Intraclass Correlation CoefficientStatistics↔ compare
Compare side by side →

Referenced by

G-TheoryRobust Test-Retest ReliabilityTest-Retest Reliability

Similar methods

Intraclass Correlation CoefficientCohen's KappaIntercoder ReliabilityFleiss' KappaKrippendorff's AlphaMultilevel Test-Retest ReliabilityScott's PiOrdinal Test-Retest Reliability

Related reference concepts

Interrater ReliabilityMeasurement Validity and ReliabilityCorrelation and CovarianceStatistical Power and Sample SizeHeterogeneity in Meta-AnalysisPsychometrics & Statistics & Methodology

Spotted an issue on this page? Report or suggest a fix →

ScholarGate — Interrater Reliability (Interrater Reliability (Cohen's κ and ICC)). Retrieved 2026-07-21 from https://scholargate.app/en/psychometrics/interrater-reliability · Dataset: https://doi.org/10.5281/zenodo.20539026
Quick facts
Originator
Cohen (kappa, 1960); Shrout & Fleiss (ICC, 1979)
Year
1960 (kappa); 1979 (ICC)
Type
Reliability / agreement analysis
Outcome
Kappa coefficient (categorical) or ICC (continuous)
Data
Categorical or continuous ratings from two or more raters
Min Sample
20
Normality Required
No
Kappa Benchmarks
< 0.40 poor, 0.40–0.60 moderate, 0.60–0.80 good, > 0.80 excellent
Related methods
Bland-Altman AnalysisCohen's KappaCronbach's AlphaFleiss' KappaG-TheoryIntraclass Correlation Coefficient
ScholarGate

A content-first reference library for research methods — what each one is, how it works, and where it comes from.

Open data (CC-BY)

Explore

  • Library
  • Search the library…
  • Browse by field
  • Fields
  • Journey
  • Compare
  • Which method?

Reference

  • Subjects
  • Atlas
  • Glossary
  • Methodology
  • Philosophy

Your tools

  • Bookshelf
  • Desk
  • Chat

Company

  • About
  • Pricing
  • Contact
  • Suggest a method

Entries are compiled from published sources for reference. Verifying the accuracy and suitability of any information for your own use remains your responsibility.

© 2026 ScholarGate · A research-method reference library
  • Privacy
  • Cookies
  • Terms
  • Delete account