Process / pipelineResearch StatisticsMultivariate-modelingPipeline

Structural Equation Modeling

Also known as: SEM, path analysis, latent variable modeling, causal modeling

OriginatorSewall WrightYear1921Sources3Related methods53

Structural equation modeling (SEM) is a comprehensive statistical framework combining path analysis (Sewall Wright, 1921) and confirmatory factor analysis to test complex causal models linking observed and latent variables. Formalized by Jöreskog (1973) with LISREL software, SEM enables simultaneous estimation of measurement relationships (how variables measure latent constructs) and structural relationships (how constructs influence outcomes), making it powerful for theory testing in psychology, epidemiology, organizational research, and health sciences where complex mediation, moderation, and latent processes require integrated analysis.

Key highlights

  • Integrated measurement and structural components: simultaneously estimates how variables measure constructs (validity) and how constructs relate (theory testing).
  • Accounts for measurement error: path coefficients corrected for unreliability in measurement, yielding more valid structural estimates than regression.
  • Tests complex relationships: mediation, moderation, indirect effects, and reciprocal causation in one coherent framework.
  • Flexible for missing data: full-information maximum likelihood enables analysis with incomplete data under MCAR assumption.
  • Comprehensive fit diagnostics: multiple indices (CFI, RMSEA, residuals) enable thorough model evaluation and identification of specific misfit sources.

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Use SEM when testing complex multivariate theories involving mediation, moderation, latent constructs, and reciprocal causation. Psychology: testing whether cognitive coping (latent) mediates the effect of stress on depression, or how personality affects behavior through motivation. Epidemiology: modeling disease risk pathways (e.g., socioeconomic status → health behaviors → cardiovascular outcomes). Organizational research: testing whether leadership (measured via multiple items) affects performance through employee engagement (latent). Education: modeling effects of school climate (latent construct from teacher and student reports) on student achievement. SEM handles latent variables naturally, making it ideal when constructs are abstract or measured via multiple indicators with error.

Strengths & limitations

Strengths
  • Integrated measurement and structural components: simultaneously estimates how variables measure constructs (validity) and how constructs relate (theory testing).
  • Accounts for measurement error: path coefficients corrected for unreliability in measurement, yielding more valid structural estimates than regression.
  • Tests complex relationships: mediation, moderation, indirect effects, and reciprocal causation in one coherent framework.
  • Flexible for missing data: full-information maximum likelihood enables analysis with incomplete data under MCAR assumption.
  • Comprehensive fit diagnostics: multiple indices (CFI, RMSEA, residuals) enable thorough model evaluation and identification of specific misfit sources.
Limitations
  • Requires large samples and balanced design: parameter estimation unstable with small n or extreme sample characteristics; minimum n ≥ 100–200 with multiple latent variables.
  • Assumes correct model specification a priori: SEM tests specified model; incorrect specification (omitted paths, wrong causal direction) may fit well despite being wrong (equivalent models).
  • Multiple equivalent models possible: different path structures can reproduce the same covariance matrix; SEM cannot distinguish among them without additional theory or constraints.
  • Non-experimental data do not imply causality: path coefficients represent regression weights, not causal effects; confounding, reverse causality, and selection bias still threaten validity.
  • Complex interactions and nonlinear relationships are difficult to specify and estimate in SEM; typically assumes linear relationships.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

What is the difference between a direct effect and an indirect effect (mediation)?

A direct effect is the path from X to Y, quantifying how X influences Y directly. An indirect effect is the path X→M→Y, quantifying how X influences Y through mediator M. The indirect effect equals the product of the X→M path times the M→Y path. Total effect = direct + indirect. Mediation occurs when: (1) X→Y direct effect decreases after controlling for M, and (2) X→M→Y indirect effect is significant (bootstrap CI excludes zero). Full mediation: direct effect becomes nonsignificant after including mediator. Partial mediation: both direct and indirect effects significant, meaning X affects Y through multiple pathways.

How do I know if my model fits well? What do CFI and RMSEA mean?

Model fit is judged by multiple indices (Hu & Bentler, 1999 criteria): CFI (Comparative Fit Index) compares your model to a null model (all variables uncorrelated); CFI > 0.95 indicates good fit (ranges 0–1). RMSEA (Root Mean Square Error of Approximation) quantifies discrepancy between model and data per degree of freedom; RMSEA < 0.06 excellent, 0.06–0.08 good, >0.10 poor (close to zero better). TLI > 0.95 and SRMR < 0.08 also recommended. Meeting all criteria simultaneously indicates good fit. However, good overall fit does not guarantee all pathways fit; examine standardized residuals (should be < |2|) and modification indices.

What is a latent variable, and how do I measure it?

A latent variable is an unobserved construct (e.g., depression, intelligence, organizational culture) measured indirectly through multiple observed indicators (questionnaire items, test scores). Each indicator reflects the latent construct plus measurement error. In SEM, you specify which observed variables measure which latent variables via factor loadings. For example, the latent variable 'depression' might be measured by items: sadness, hopelessness, sleep disturbance (observed indicators). Each loading should be ≥0.5; if an indicator loads weakly, consider removing it or reconceptualizing the latent construct.

Can I use SEM to establish causality from observational data?

No. SEM estimates associations and tests whether a hypothesized causal model reproduces the correlation structure in data. Path coefficients represent regression weights (conditional associations), not causal effects. Causality requires: (1) temporal precedence (X measured before Y), (2) covariation (X and Y correlate), (3) no plausible alternatives (confounding, reverse causality). Observational SEM violates criterion 3: confounding variables and unmeasured causes threaten validity. SEM's statistical fit alone cannot prove causality. Use SEM with theory, qualitative evidence, and study design (randomization, instrumental variables) to strengthen causal inference. Consider sensitivity analyses testing alternative models.

Sources

  1. 1.
    Jöreskog, K. G., & Sörbom, D. (1973). LISREL: A general computer program for estimating a linear structural equation system. Research Bulletin 73-5. University of Stockholm.
  2. 2.
    Hu, L. T., & Bentler, P. M. (1999). Cutoff criteria for fit indices in covariance structure analysis: Conventional criteria versus new alternatives. Structural Equation Modeling, 6(1), 1–55.
  3. 3.
    Wright, S. (1921). Correlation and causation. Journal of Agricultural Research, 20(7), 557–585.

You have read it. What now?

Cite this page

ScholarGate. (2026, June 4). Structural Equation Modeling. ScholarGate. https://scholargate.app/research-statistics/structural-equation-modeling