Multilevel Generalizability Theory
Also known as: multilevel G-theory, ML-GT, hierarchical generalizability theory, multilevel G-study
Multilevel generalizability theory extends classical G-theory to measurement designs where observations are nested within higher-level units — for example, items nested within raters, or students nested within classrooms. It decomposes score variance into components attributable to persons, facets, and their interactions across hierarchical levels, enabling precise estimation of measurement precision in complex, real-world assessment settings.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Use multilevel generalizability theory when your measurement design includes a genuine hierarchical structure — students nested within classrooms, items nested within test booklets, raters nested within scoring sites — and you need accurate estimates of measurement reliability or the sources of measurement error. It is especially valuable in large-scale educational assessment, workplace evaluation with multiple nested raters, and clinical measurement with clustered data. Do not use it when the design is fully crossed and balanced (standard G-theory suffices), when sample sizes at higher levels are very small (variance component estimates become unstable), or when the research question is about prediction or latent structure rather than measurement precision.
Strengths & limitations
- Correctly partitions variance when data have a hierarchical structure, avoiding the biased estimates that arise when nesting is ignored.
- Provides actionable guidance through D-studies: shows which facets drive measurement error and how design changes would improve reliability.
- Accommodates complex real-world assessment designs where full crossing of all facets is impractical or impossible.
- Distinguishes relative error (relevant to norm-referenced decisions) from absolute error (relevant to criterion-referenced decisions) through the E-rho and phi coefficients.
- Can be integrated with multilevel modelling software (e.g., lme4 in R) to leverage well-tested estimation algorithms.
- Requires large samples at each level of the hierarchy; small cluster sizes yield unstable variance component estimates, sometimes including negative estimates.
- Fully nested designs cannot disentangle certain confounded variance components, limiting interpretability compared with crossed designs.
- The framework assumes a linear additive model for variance; it does not naturally accommodate nonlinear interactions or non-normal response distributions without extensions.
- Software support is less plug-and-play than for classical reliability indices; analysts must specify the hierarchical model carefully and verify that the nesting structure is correctly coded.
Frequently asked
How does multilevel G-theory differ from standard generalizability theory?
Standard G-theory assumes a flat design where all facets (raters, items, occasions) are either fully crossed with or fully nested within persons. Multilevel G-theory handles the additional complication that facets or persons themselves are clustered within higher-level units — such as classrooms or scoring sites — and partitions variance at each level of the hierarchy rather than collapsing across levels.
Can I run multilevel G-theory in standard statistical software?
Yes, by fitting a sequence of variance-components models using multilevel modelling routines (e.g., lme4 in R or PROC MIXED in SAS). The analyst specifies random effects corresponding to the persons, facets, and nesting units, then extracts and combines the resulting variance component estimates to compute G and phi coefficients.
What sample size is needed at the higher level of the hierarchy?
A common guideline is at least 30 units at the highest level for stable variance component estimates, though more is always better. With fewer than 10 clusters the between-cluster variance component can be severely biased, and negative estimates become common.
When should I use the phi coefficient instead of the G-coefficient?
Use the phi (absolute) coefficient whenever your decision is criterion-referenced — that is, when you are classifying individuals against a fixed standard (pass/fail, meets-standard) rather than ranking them relative to each other. Phi is more stringent because it includes absolute error variance in the denominator, not just the relative error that affects rank ordering.
Is multilevel G-theory the same as multilevel confirmatory factor analysis?
No. Both account for hierarchical data, but they answer different questions. Multilevel CFA estimates latent factor loadings and tests structural relationships between constructs at each level. Multilevel G-theory is specifically focused on quantifying measurement error, estimating reliability, and planning efficient assessment designs — it does not model latent factors or test structural hypotheses.
Sources
- Briggs, D. C. & Wilson, M. (2003). An introduction to multidimensional measurement using Rasch models and generalizability theory. Journal of Applied Measurement, 4(1), 1–19. link ↗
- Webb, N. M., Shavelson, R. J. & Haertel, E. H. (2006). Reliability coefficients and generalizability theory. Handbook of Statistics, 26, 81–124. DOI: 10.1016/S0169-7161(06)26004-8 ↗
How to cite this page
ScholarGate. (2026, June 3). Multilevel Generalizability Theory. ScholarGate. https://scholargate.app/en/psychometrics/multilevel-generalizability-theory
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Confirmatory factor analysisPsychometrics↔ compare
- Generalizability TheoryPsychometrics↔ compare
- Item Response TheoryPsychometrics↔ compare
- Multilevel ModelingResearch Statistics↔ compare