Process / pipelineResearch StatisticsCausal-reasoningPipeline

Correlation vs Causation

Also known as: correlation and causation, causal inference, spurious correlation, confounding

OriginatorMultiple sources (Bradford Hill, Judea Pearl, Donald Rubin)Year1965Sources3Related methods4

Correlation measures the strength and direction of association between two variables; causation implies that changes in one variable directly produce changes in another. A strong correlation (e.g., r = 0.9) does not prove causation. Classic examples abound: shoe size and reading ability are correlated in children (confounded by age), but shoe size does not cause reading ability. Understanding when correlation implies causation requires evaluating study design, confounding variables, temporal precedence, and mechanism. Randomized experiments offer the strongest causal evidence; observational studies must carefully control for confounders.

Key highlights

  • Prevents incorrect causal interpretations of correlational data, protecting against flawed policy and clinical decisions.
  • Randomized experiments offer definitive causal evidence by breaking confounding through random assignment.
  • Structured frameworks (Hill's criteria, DAGs) provide systematic approaches to evaluating causation in non-experimental studies.
  • Highlights the importance of study design and confounding control, promoting more rigorous research planning.

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Always ask: 'Does this study design support causal inference?' When reading observational research claiming causation without controlling for confounders, be skeptical. When designing a study where causal inference is important, use randomization if possible. When randomization is infeasible (e.g., studying effects of smoking on health), use observational designs with strong confounding control: measure potential confounders, adjust statistically, and discuss residual confounding limitations. Use causal graphs (Directed Acyclic Graphs, DAGs) to visualize confounding and guide adjustment strategies.

Strengths & limitations

Strengths
  • Prevents incorrect causal interpretations of correlational data, protecting against flawed policy and clinical decisions.
  • Randomized experiments offer definitive causal evidence by breaking confounding through random assignment.
  • Structured frameworks (Hill's criteria, DAGs) provide systematic approaches to evaluating causation in non-experimental studies.
  • Highlights the importance of study design and confounding control, promoting more rigorous research planning.
Limitations
  • Randomized trials are often infeasible, unethical, or expensive (e.g., studying effects of smoking requires observational designs).
  • Unmeasured confounders in observational studies can bias causal estimates; statistical adjustment cannot control for variables not measured.
  • Causal inference from observational data requires strong assumptions (e.g., no unmeasured confounders, correct functional form) that are often unverifiable.
  • Complex scenarios with feedback loops, selection bias, or time-varying confounders require advanced causal methods (instrumental variables, marginal structural models) that are less familiar to most researchers.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

Does a high correlation coefficient (e.g., r = 0.95) prove causation?

No. A very strong correlation is necessary but not sufficient for causation. Example: the number of firefighters at a fire is strongly correlated with fire damage (r might be 0.8+), but firefighters do not cause damage—they respond to large fires, which cause damage. Confounding by fire size explains the correlation. Always investigate alternative explanations before claiming causation.

What is confounding and how do I identify it?

Confounding occurs when a third variable (confounder) is associated with both the cause and effect, creating a spurious association. Example: age confounds the relation between shoe size (cause) and reading ability (effect) because age is associated with both. To identify confounders: (1) Think logically about what affects both variables. (2) Check if the association weakens after adjusting for the confounder statistically. (3) Use causal diagrams (DAGs) to visualize the causal structure.

Are randomized experiments always better than observational studies?

For establishing causation, yes—RCTs are the gold standard. But RCTs have limitations: they are expensive, may be unethical, and may not reflect real-world conditions (external validity). Observational studies are necessary for studying long-term effects, rare outcomes, and exposures that cannot be randomly assigned. Use RCTs for causal hypothesis testing; use well-designed observational studies for hypothesis generation and real-world effectiveness.

What is reverse causation and how do I avoid it?

Reverse causation occurs when you incorrectly infer the direction of causality. Example: does depression cause insomnia or does insomnia cause depression? Avoid reverse causation by: (1) Establishing temporal precedence with longitudinal designs (measure cause before outcome). (2) Using theory and prior research to guide causal direction. (3) Considering bidirectional relationships; some causal pathways are mutually reinforcing. Cross-sectional data cannot establish direction; use prospective cohort studies.

What are unmeasured confounders and can I account for them?

Unmeasured confounders are variables that confound the association but were not measured in the study. You cannot statistically control for variables you did not measure. Address unmeasured confounding by: (1) Measuring and adjusting for likely confounders in the study design. (2) Sensitivity analyses: assume different amounts of unmeasured confounding and see if conclusions change. (3) Using instrumental variables or natural experiments, which can estimate causal effects even with unmeasured confounding. (4) Discussing limitations honestly in the paper.

Sources

  1. 1.
    Pearl, J. (2009). Causality: Models, Reasoning, and Inference (2nd ed.). Cambridge University Press.
    ISBN 978-0-521-89560-6
  2. 2.
    Rubin, D. B. (1974). Estimating causal effects of treatments in randomized and nonrandomized studies. Journal of Educational Psychology, 66(5), 688–701.
  3. 3.
    Hill, A. B. (1965). The Environment and Disease: Association or Causation? Proceedings of the Royal Society of Medicine, 58(5), 295–300.

You have read it. What now?

Cite this page

ScholarGate. (2026, June 3). Correlation vs Causation. ScholarGate. https://scholargate.app/research-statistics/correlation-vs-causation