Longitudinal Generalizability Theory
Also known as: longitudinal G-theory, longitudinal GT, repeated-measures generalizability theory, G-theory for longitudinal designs
Longitudinal generalizability theory extends classical G-theory to repeated-measures and longitudinal designs, decomposing score variance across persons, measurement occasions, raters, and items simultaneously. It quantifies how reliably scores can be generalized across time points, evaluators, and conditions — information that is invisible to cross-sectional reliability indices.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Use longitudinal G-theory when scores are collected at multiple time points and you need to know whether the measure is consistent across occasions, raters, items, or all three simultaneously. It is especially valuable in educational assessment tracking student growth, in clinical studies monitoring symptom change, and in organizational research examining performance ratings over time. Do not use it as a substitute for latent growth curve modeling when your goal is to estimate individual growth trajectories or test predictors of change — G-theory answers reliability questions, not structural ones. It is also unsuitable when each person is measured only once (cross-sectional data) or when you have fewer than two levels on any facet you wish to estimate.
Strengths & limitations
- Simultaneously estimates multiple sources of measurement error — persons, occasions, raters, items — that cross-sectional reliability indices confound.
- The D-study framework lets researchers optimize future measurement designs before data collection, balancing cost and reliability.
- Produces both relative (E-rho-squared) and absolute (Phi) reliability estimates, covering both ranking and mastery-type decisions.
- Formally models person-by-occasion interactions, capturing whether individuals differ in their trajectories — a key longitudinal insight.
- Applicable to both fully-crossed and partially-nested longitudinal designs, matching real-world data collection constraints.
- Requires sufficiently large, balanced or near-balanced samples for stable ANOVA-based variance component estimation; heavily unbalanced designs complicate standard estimation.
- Assumes variance components are homogeneous across time points; if measurement properties shift substantially (e.g., scale restructuring), a single G-study may be misleading.
- Does not model nonlinear growth trajectories or test predictors of change — those questions require latent growth curve or multilevel modeling.
- Interpretation of interaction variance components (especially three-way interactions) can be ambiguous without careful design planning.
Frequently asked
How does longitudinal G-theory differ from standard (cross-sectional) G-theory?
Standard G-theory treats all facets as potentially interchangeable sources of error. Longitudinal G-theory explicitly includes occasions as a substantive facet, allowing the analyst to separate true score change over time (a meaningful longitudinal signal) from random measurement error. It also enables D-studies that optimize the number of occasions, not just raters or items.
Can I use longitudinal G-theory with unbalanced designs?
Standard expected-mean-squares estimation assumes balanced or near-balanced data. For substantially unbalanced longitudinal data (e.g., dropout, missing occasions), restricted maximum likelihood (REML) or Bayesian estimation of variance components is recommended, implemented in software such as GENOVA, urGENOVA, or the R package gtheory.
What software runs longitudinal G-theory analyses?
GENOVA and urGENOVA (by Brennan and colleagues) handle a wide range of G-theory designs. The R packages gtheory and lme4 (for REML-based variance component estimation) are common alternatives. Longitudinal designs with nesting require careful syntax to define the crossed-versus-nested structure.
Is a high generalizability coefficient sufficient evidence of good longitudinal measurement?
Not alone. A high E-rho-squared shows scores rank-order persons consistently, but it does not guarantee measurement invariance across occasions. You should also test configural and metric invariance with longitudinal confirmatory factor analysis to ensure items measure the same construct at each time point.
What sample size is needed?
There is no single rule, because the required sample depends on the number of facets and levels. As a rough guideline, at least 30–50 persons and at least two levels on every facet are needed for stable variance component estimates. Simulation studies (e.g., using the bootG function in R) can assess estimate precision for a specific planned design.
Sources
- Webb, N. M., Shavelson, R. J., & Harrigan, E. H. (2007). Generalizability theory: Overview. In C. R. Rao & S. Sinharay (Eds.), Handbook of Statistics, Vol. 26: Psychometrics (pp. 1–43). Elsevier. link ↗
- Brennan, R. L. (2001). Generalizability Theory. Springer. ISBN: 978-0387952826
How to cite this page
ScholarGate. (2026, June 3). Longitudinal Generalizability Theory. ScholarGate. https://scholargate.app/en/psychometrics/longitudinal-generalizability-theory
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Confirmatory factor analysisPsychometrics↔ compare
- EFAStatistics↔ compare
- Generalizability TheoryPsychometrics↔ compare
- Multilevel ModelingResearch Statistics↔ compare