Intraclass Correlation Coefficient (ICC)
Also known as: ICC, intraclass correlation, rater reliability coefficient, Sınıf İçi Korelasyon Katsayısı (ICC)
The Intraclass Correlation Coefficient (ICC) is a parametric reliability statistic that quantifies the degree of agreement or consistency among repeated measurements or multiple raters on a continuous outcome. The modern six-form taxonomy was established by Shrout and Fleiss in 1979 and remains the standard framework for selecting and reporting ICC in inter-rater reliability, test-retest repeatability, and multilevel-data analyses.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Use ICC whenever multiple raters or repeated occasions measure the same subjects on a continuous scale and you need to quantify how reproducible those measurements are. It is the standard tool for inter-rater reliability in clinical, educational, and behavioural research, and for test-retest reliability in instrument validation studies. Key requirements are: the outcome must be continuous and approximately normally distributed, the study design must be identified (one-way random, two-way random, or two-way mixed), and the distinction between absolute agreement and consistency must be decided before analysis. ICC is not appropriate for categorical or binary ratings — use Cohen's kappa or Fleiss' kappa instead.
Strengths & limitations
- Covers six distinct model forms, accommodating a wide range of study designs from fixed to random raters.
- Simultaneously tests statistical significance and provides an interpretable 0-to-1 reliability magnitude.
- Accounts for both systematic rater bias (absolute agreement) and relative consistency depending on the chosen model.
- Confidence intervals provide information about estimation precision, which is especially important in small samples.
- The choice among the six forms requires explicit design decisions; selecting the wrong model produces a misleading ICC value.
- Sensitive to the range of the target variable: a restricted range (homogeneous subjects) artificially deflates the ICC even when raters agree well.
- Assumes continuous, approximately normally distributed ratings; ordinal data with few categories require alternative approaches.
- Confidence intervals can be very wide with small sample sizes (below 30 subjects), making point estimates unreliable.
Frequently asked
Which of the six ICC forms should I use?
The choice depends on two decisions: (1) the rater model — one-way random if each subject is rated by a different random subset of raters, two-way random if all subjects are rated by the same raters who are a random sample from a larger pool, two-way mixed if the same fixed raters will always be used; (2) the unit — single measures if you will apply a single rater's score in practice, average measures if you will average across raters. Shrout and Fleiss (1979) and Koo and Li (2016) both provide decision flowcharts.
What is the difference between agreement and consistency?
Consistency (ICC3) captures whether raters rank subjects in the same order, tolerating a constant offset between raters. Absolute agreement (ICC1 and ICC2) additionally requires that raters assign numerically similar values. For clinical substitutability of measurement methods, absolute agreement is the correct choice; for psychometric scale validation where systematic bias can be calibrated out, consistency may suffice.
My sample has only 15 subjects. Can I still report ICC?
You can compute an ICC, but the 95% confidence interval will be very wide, making the point estimate unstable. As a rule of thumb, at least 30 subjects are recommended for a reasonably precise estimate. With very small samples, report the full confidence interval prominently and interpret the findings with caution.
How is ICC different from Pearson's correlation when comparing two raters?
Pearson's r measures only whether two variables move together linearly and is insensitive to systematic additive or proportional differences between raters. ICC for absolute agreement captures both the linear relationship and whether the actual numerical values match, making it the appropriate choice when you need to establish that raters are interchangeable rather than merely correlated.
Sources
- Shrout, P.E. & Fleiss, J.L. (1979). Intraclass Correlations: Uses in Assessing Rater Reliability. Psychological Bulletin, 86(2), 420–428. DOI: 10.1037/0033-2909.86.2.420 ↗
- Koo, T.K. & Li, M.Y. (2016). A Guideline of Selecting and Reporting Intraclass Correlation Coefficients for Reliability Research. Journal of Chiropractic Medicine, 15(2), 155–163. DOI: 10.1016/j.jcm.2016.02.012 ↗
How to cite this page
ScholarGate. (2026, June 1). Intraclass Correlation Coefficient (ICC). ScholarGate. https://scholargate.app/en/statistics/icc-intraclass-correlation
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Bland-Altman AnalysisStatistics↔ compare
- Cohen's KappaStatistics↔ compare
- Cronbach's AlphaStatistics↔ compare
- Fleiss' KappaStatistics↔ compare
- Pearson CorrelationStatistics↔ compare