Structural Equation Modeling
Structural Equation Modeling (SEM) · Also known as: SEM, path analysis, latent variable modeling, causal modeling
Structural equation modeling (SEM) is a comprehensive statistical framework combining path analysis (Sewall Wright, 1921) and confirmatory factor analysis to test complex causal models linking observed and latent variables. Formalized by Jöreskog (1973) with LISREL software, SEM enables simultaneous estimation of measurement relationships (how variables measure latent constructs) and structural relationships (how constructs influence outcomes), making it powerful for theory testing in psychology, epidemiology, organizational research, and health sciences where complex mediation, moderation, and latent processes require integrated analysis.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
+40 more
When to use it
Use SEM when testing complex multivariate theories involving mediation, moderation, latent constructs, and reciprocal causation. Psychology: testing whether cognitive coping (latent) mediates the effect of stress on depression, or how personality affects behavior through motivation. Epidemiology: modeling disease risk pathways (e.g., socioeconomic status → health behaviors → cardiovascular outcomes). Organizational research: testing whether leadership (measured via multiple items) affects performance through employee engagement (latent). Education: modeling effects of school climate (latent construct from teacher and student reports) on student achievement. SEM handles latent variables naturally, making it ideal when constructs are abstract or measured via multiple indicators with error.
Strengths & limitations
- Integrated measurement and structural components: simultaneously estimates how variables measure constructs (validity) and how constructs relate (theory testing).
- Accounts for measurement error: path coefficients corrected for unreliability in measurement, yielding more valid structural estimates than regression.
- Tests complex relationships: mediation, moderation, indirect effects, and reciprocal causation in one coherent framework.
- Flexible for missing data: full-information maximum likelihood enables analysis with incomplete data under MCAR assumption.
- Comprehensive fit diagnostics: multiple indices (CFI, RMSEA, residuals) enable thorough model evaluation and identification of specific misfit sources.
- Requires large samples and balanced design: parameter estimation unstable with small n or extreme sample characteristics; minimum n ≥ 100–200 with multiple latent variables.
- Assumes correct model specification a priori: SEM tests specified model; incorrect specification (omitted paths, wrong causal direction) may fit well despite being wrong (equivalent models).
- Multiple equivalent models possible: different path structures can reproduce the same covariance matrix; SEM cannot distinguish among them without additional theory or constraints.
- Non-experimental data do not imply causality: path coefficients represent regression weights, not causal effects; confounding, reverse causality, and selection bias still threaten validity.
- Complex interactions and nonlinear relationships are difficult to specify and estimate in SEM; typically assumes linear relationships.
Frequently asked
What is the difference between a direct effect and an indirect effect (mediation)?
A direct effect is the path from X to Y, quantifying how X influences Y directly. An indirect effect is the path X→M→Y, quantifying how X influences Y through mediator M. The indirect effect equals the product of the X→M path times the M→Y path. Total effect = direct + indirect. Mediation occurs when: (1) X→Y direct effect decreases after controlling for M, and (2) X→M→Y indirect effect is significant (bootstrap CI excludes zero). Full mediation: direct effect becomes nonsignificant after including mediator. Partial mediation: both direct and indirect effects significant, meaning X affects Y through multiple pathways.
How do I know if my model fits well? What do CFI and RMSEA mean?
Model fit is judged by multiple indices (Hu & Bentler, 1999 criteria): CFI (Comparative Fit Index) compares your model to a null model (all variables uncorrelated); CFI > 0.95 indicates good fit (ranges 0–1). RMSEA (Root Mean Square Error of Approximation) quantifies discrepancy between model and data per degree of freedom; RMSEA < 0.06 excellent, 0.06–0.08 good, >0.10 poor (close to zero better). TLI > 0.95 and SRMR < 0.08 also recommended. Meeting all criteria simultaneously indicates good fit. However, good overall fit does not guarantee all pathways fit; examine standardized residuals (should be < |2|) and modification indices.
What is a latent variable, and how do I measure it?
A latent variable is an unobserved construct (e.g., depression, intelligence, organizational culture) measured indirectly through multiple observed indicators (questionnaire items, test scores). Each indicator reflects the latent construct plus measurement error. In SEM, you specify which observed variables measure which latent variables via factor loadings. For example, the latent variable 'depression' might be measured by items: sadness, hopelessness, sleep disturbance (observed indicators). Each loading should be ≥0.5; if an indicator loads weakly, consider removing it or reconceptualizing the latent construct.
Can I use SEM to establish causality from observational data?
No. SEM estimates associations and tests whether a hypothesized causal model reproduces the correlation structure in data. Path coefficients represent regression weights (conditional associations), not causal effects. Causality requires: (1) temporal precedence (X measured before Y), (2) covariation (X and Y correlate), (3) no plausible alternatives (confounding, reverse causality). Observational SEM violates criterion 3: confounding variables and unmeasured causes threaten validity. SEM's statistical fit alone cannot prove causality. Use SEM with theory, qualitative evidence, and study design (randomization, instrumental variables) to strengthen causal inference. Consider sensitivity analyses testing alternative models.
Sources
- Jöreskog, K. G., & Sörbom, D. (1973). LISREL: A general computer program for estimating a linear structural equation system. Research Bulletin 73-5. University of Stockholm. link ↗
- Hu, L. T., & Bentler, P. M. (1999). Cutoff criteria for fit indices in covariance structure analysis: Conventional criteria versus new alternatives. Structural Equation Modeling, 6(1), 1–55. DOI: 10.1080/10705519909540118 ↗
- Wright, S. (1921). Correlation and causation. Journal of Agricultural Research, 20(7), 557–585. link ↗
How to cite this page
ScholarGate. (2026, June 4). Structural Equation Modeling (SEM). ScholarGate. https://scholargate.app/en/research-statistics/structural-equation-modeling
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Factor AnalysisResearch Statistics↔ compare
- Multilevel ModelingResearch Statistics↔ compare
- Multiple Regression AnalysisResearch Statistics↔ compare