Computerized Adaptive Test Generalizability Theory
Also known as: CAT G-theory, adaptive test generalizability, G-theory in CAT, computerized adaptive generalizability analysis
Generalizability theory (G-theory) applied to computerized adaptive testing (CAT) evaluates the dependability of adaptive test scores by decomposing score variance across measurement facets such as persons, items, and occasions. Unlike classical test theory, G-theory quantifies multiple simultaneous sources of measurement error, offering a richer reliability picture for adaptively administered assessments.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Apply CAT generalizability theory when you need to report the dependability of adaptive test scores beyond a single IRT-based standard error, or when your assessment supports both relative (norm-referenced) and absolute (criterion-referenced) score interpretations simultaneously. It is also useful when multiple facets — items, raters, occasions, or forms — contribute to measurement error in the adaptive system. Do not use it as a replacement for IRT parameter estimation; G-theory assumes a classical or variance-component framework that complements but does not replace the IRT ability estimates that drive item selection. Avoid applying standard G-theory designs naively to CAT data without accounting for the non-random, ability-matched nature of item selection, as this will distort variance component estimates.
Strengths & limitations
- Decomposes multiple simultaneous sources of measurement error that classical reliability indices cannot separate.
- Provides both relative (E-rho-squared) and absolute (Phi) dependability coefficients relevant to different score interpretations.
- D-study projections allow planners to determine optimal item bank size and minimum test length for a target reliability.
- Accommodates complex measurement designs with crossed and nested facets, suitable for large-scale adaptive systems.
- Bridges classical psychometric reporting standards with the more technically demanding IRT framework expected in CAT documentation.
- The non-random item selection in CAT violates the random sampling assumption underlying standard G-study designs, requiring methodological adaptations.
- Variance component estimation requires large samples and broad item banks to be stable; small examinee pools yield imprecise G-coefficients.
- Integrating G-theory with IRT scoring is methodologically complex and not yet fully standardised in software.
- G-theory does not model the conditional measurement error that IRT provides at each ability level; precision varies across the ability continuum in ways G-theory averages over.
Frequently asked
Can I use Cronbach's alpha to assess the reliability of CAT scores?
No. Cronbach's alpha assumes all examinees respond to the same items, but in a CAT each examinee receives a unique subset. Applying alpha to pooled CAT responses will produce a misleading coefficient. G-theory or IRT-based conditional standard errors should be used instead.
How do I handle the non-random item selection when running a G-study on CAT data?
The standard approach is either to use marginal maximum likelihood estimation of variance components that treats item selection as ignorable given the ability estimate, or to pool person scores from IRT and then treat those scores in a G-study framework. Some researchers simulate random administration from the bank using bootstrap resampling to obtain design-consistent variance components.
What is the difference between the G-coefficient and the dependability coefficient Phi?
The G-coefficient (E-rho-squared) is relevant when you care about the relative ordering of examinees, such as selecting the top scorers. Phi is relevant for absolute decisions, such as passing everyone who scores above a cut score, because it includes item mean differences as part of the error. Phi is always equal to or lower than E-rho-squared.
Is G-theory a replacement for IRT in CAT?
No. IRT drives item selection and provides ability estimates with conditional standard errors that vary by ability level. G-theory provides an overall dependability summary useful for program-level reporting and for comparing design alternatives. They serve complementary purposes and are best reported together.
How large a sample is needed for stable G-study variance components in CAT?
As a rough guide, at least 200 to 300 examinees and at least 50 to 100 items in the bank are needed for stable variance component estimates in a simple two-facet design. More complex designs with additional facets require proportionally larger samples.
Sources
- Brennan, R. L. (2001). Generalizability Theory. Springer. ISBN: 978-0387952826
- Van der Linden, W. J., & Glas, C. A. W. (2000). Computerized adaptive testing: Theory and practice. Kluwer Academic Publishers. link ↗
How to cite this page
ScholarGate. (2026, June 3). Computerized Adaptive Test Generalizability Theory. ScholarGate. https://scholargate.app/en/psychometrics/computerized-adaptive-test-generalizability-theory
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Computerized adaptive test item response theoryPsychometrics↔ compare
- Computerized adaptive test reliability analysisPsychometrics↔ compare
- Generalizability TheoryPsychometrics↔ compare
- Item Response TheoryPsychometrics↔ compare
- Multilevel Reliability AnalysisPsychometrics↔ compare
- Test-Retest ReliabilityPsychometrics↔ compare