Process / pipelineCommunicationIntercoder reliability coefficientsPipeline

Gwet's AC1

Also known as: Gwet AC1, AC1 coefficient, Gwet's first-order agreement coefficient, Gwet AC1 Katsayısı

OriginatorKilem L. GwetYear2008Sources3Related methods4

Gwet's AC1 is a chance-corrected agreement coefficient introduced by Kilem Gwet in 2008 as a robust alternative to Cohen's and Fleiss' kappa. It targets the kappa paradox — the unsettling result where coders agree on the vast majority of units yet kappa is near zero because one category dominates — by estimating chance agreement in a way that does not collapse when category prevalence is extreme.

Key highlights

  • Resistant to the prevalence/kappa paradox, so it does not collapse to near zero under skewed category distributions despite high observed agreement.
  • Comes with a closed-form variance estimator, enabling confidence intervals and hypothesis tests without bootstrapping.
  • Generalizes to multiple raters and, via AC2, to ordinal categories with user-specified weights.
  • Conceptually motivated chance model that distinguishes genuinely ambiguous units from easy ones, giving fairer estimates for rare-event coding.

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Use Gwet's AC1 when you expect — or observe — that one category dominates the coding, a situation where Cohen's or Fleiss' kappa can be paradoxically low despite excellent observed agreement. It is well suited to rare-event coding (e.g., flagging the small fraction of messages containing hate speech or misinformation) and to any categorical reliability task with skewed prevalence. AC1 assumes nominal (or, in its AC2 extension, ordinal/weighted) categories and interchangeable raters. It is less necessary when categories are roughly balanced, where kappa and AC1 largely agree, and it does not replace Krippendorff's alpha for designs with missing data or interval-level variables — though AC1 and alpha often serve as complementary robustness checks.

Strengths & limitations

Strengths
  • Resistant to the prevalence/kappa paradox, so it does not collapse to near zero under skewed category distributions despite high observed agreement.
  • Comes with a closed-form variance estimator, enabling confidence intervals and hypothesis tests without bootstrapping.
  • Generalizes to multiple raters and, via AC2, to ordinal categories with user-specified weights.
  • Conceptually motivated chance model that distinguishes genuinely ambiguous units from easy ones, giving fairer estimates for rare-event coding.
Limitations
  • Its chance model rests on assumptions about random rating that, like kappa's, are not directly observable and can be debated.
  • Less widely reported than kappa and alpha, so reviewers and readers may be less familiar with its interpretation.
  • Benchmark thresholds borrowed from kappa are conventions and may not transfer perfectly to AC1's scale.
  • For balanced categories it offers little advantage over kappa, adding a coefficient without changing conclusions.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

Why does Gwet's AC1 avoid the kappa paradox?

The paradox arises because kappa's chance term, built from the product of category prevalences, becomes very large when one category dominates, so it subtracts away almost all observed agreement. AC1's chance term instead weights each category by π_c(1 − π_c), which shrinks toward zero as a category's prevalence approaches one. So when coders genuinely agree because one category is obviously common, AC1 attributes that agreement to real classification rather than to chance, and stays high.

What is the difference between AC1 and AC2?

AC1 treats categories as unordered (nominal), counting any mismatch equally. AC2 extends the coefficient to ordinal or interval data by applying weights — typically quadratic or linear — so that disagreements between adjacent categories count less than disagreements between distant ones. Use AC1 for nominal coding and AC2 when your categories have a meaningful order, analogous to choosing the right metric in Krippendorff's alpha.

Should I report AC1 instead of kappa?

Report whatever your field expects, but disclose the category distribution and ideally more than one coefficient. If prevalence is skewed and kappa is paradoxically low while observed agreement is high, AC1 gives a more defensible estimate and should be reported with an explanation. When categories are balanced, AC1 and kappa agree closely and either suffices. Transparency about prevalence is more important than the choice of a single coefficient.

Sources

  1. 1.
    Gwet, K. L. (2008). Computing inter-rater reliability and its variance in the presence of high agreement. British Journal of Mathematical and Statistical Psychology, 61(1), 29–48.
  2. 2.
    Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1), 37–46.
  3. 3.
    Scott, W. A. (1955). Reliability of content analysis: The case of nominal scale coding. Public Opinion Quarterly, 19(3), 321–325.

You have read it. What now?

Cite this page

ScholarGate. (2026, June 22). Gwet's AC1. ScholarGate. https://scholargate.app/communication/gwet-ac1-reliability

Gwet's AC1 — Gwet's AC1 Agreement Coefficient | ScholarGate