Multi-group Generalizability Theory
Also known as: MG G-theory, multi-group G-theory, generalizability theory across groups, cross-group G-study
Multi-group generalizability theory (MG G-theory) extends classical generalizability theory to estimate and compare variance components — attributable to persons, items, raters, occasions, and their interactions — simultaneously across two or more defined groups. It reveals whether a measurement procedure is equally reliable and generalizable for every group studied, supporting fair and equitable score interpretation.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Use multi-group generalizability theory when you need to evaluate and compare the reliability and generalizability of measurement across two or more defined groups — for instance, gender, language background, clinical diagnosis, or cultural group — and when the measurement design involves multiple facets such as raters, items, or occasions. It is especially valuable in educational and clinical assessment where fairness across groups must be documented. Do not use MG G-theory when only one group is present (classical G-theory suffices), when the measurement design involves only one facet (simpler reliability indices apply), or when sample sizes within any group are too small to obtain stable variance-component estimates — a rough minimum is 30 persons per group, more when the design is complex.
Strengths & limitations
- Simultaneously estimates multiple sources of measurement error (raters, items, occasions) within each group, providing a richer picture than a single reliability coefficient.
- Supports direct comparison of reliability and error-source profiles across groups, enabling evidence-based fairness evaluations.
- D-study projections allow researchers to plan how many items or raters are needed to reach target reliability separately for each group.
- Handles crossed and nested measurement designs, accommodating a wide range of real-world data collection situations.
- Provides both relative (norm-referenced, E-rho-squared) and absolute (criterion-referenced, Phi) reliability coefficients for each group.
- Requires adequate sample sizes within each group; sparse cells in a complex design lead to negative or unstable variance-component estimates.
- Assumes linear additive decomposition of variance; interactions beyond two-way terms are typically confounded with residual error.
- Group must be operationally defined before analysis; MG G-theory does not identify groups or test their existence.
- Software support is less widespread than for CFA-based measurement invariance approaches; specialized programs (mGENOVA, urGENOVA, or custom R code) are usually required.
Frequently asked
How does multi-group G-theory differ from ordinary G-theory?
Ordinary G-theory analyses a single group and partitions score variance into sources such as persons, items, raters, and their interactions. Multi-group G-theory runs the same decomposition separately within each group and then compares the resulting variance components and reliability coefficients to determine whether the measurement procedure operates with equal consistency across populations.
How does multi-group G-theory relate to CFA-based measurement invariance testing?
Both evaluate whether a measurement instrument behaves equivalently across groups, but they approach the question differently. MG CFA tests whether factor loadings and item intercepts are equal across groups, focusing on the latent-variable model. MG G-theory estimates and compares variance components attributable to facets of the measurement procedure (raters, items, occasions), focusing on sources of error. They are complementary: G-theory is particularly informative when raters or occasions are salient facets, whereas CFA is preferred when the goal is construct-level invariance.
What sample size is needed within each group?
There is no single universal rule, but stable variance-component estimation generally requires at least 30 persons per group for simple designs. Complex designs with many crossed facets may require substantially larger samples. Simulation studies suggest that negative variance-component estimates and wide confidence intervals are symptoms of insufficient within-group sample size.
Which software can fit multi-group G-theory models?
Brennan's mGENOVA program was purpose-built for multi-group G-studies. urGENOVA handles unbalanced designs. In R, the gtheory and lme4 packages can be adapted for variance-component estimation within each group, though they require manual comparison steps. Dedicated G-theory software remains the most straightforward option for standard crossed or nested designs.
Can multi-group G-theory handle missing data?
Standard ANOVA-based G-theory assumes a balanced design with no missing observations. Imbalance or missing data require restricted maximum likelihood (REML) estimation, available through mixed-model software such as lme4 in R. Researchers should verify that missingness is not group-differential before comparing variance components, as systematic missing data within one group can bias its estimated error components.
Sources
- Brennan, R. L. (2001). Generalizability Theory. Springer. ISBN: 978-0387952826
- Shavelson, R. J. & Webb, N. M. (1989). Generalizability theory: 1973–1988. British Journal of Mathematical and Statistical Psychology, 42(1), 3–27. DOI: 10.1037/0003-066x.44.6.922 ↗
How to cite this page
ScholarGate. (2026, June 3). Multi-group Generalizability Theory. ScholarGate. https://scholargate.app/en/psychometrics/multi-group-generalizability-theory
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Generalizability TheoryPsychometrics↔ compare
- Multi-group confirmatory factor analysisPsychometrics↔ compare
- Multi-group Cronbach's alphaPsychometrics↔ compare
- Multi-group measurement invariancePsychometrics↔ compare
- Multi-group Rasch modelPsychometrics↔ compare
- Multi-group Reliability AnalysisPsychometrics↔ compare