Correlation vs Causation
Understanding the Distinction Between Correlation and Causation in Research · Also known as: correlation and causation, causal inference, spurious correlation, confounding
Correlation measures the strength and direction of association between two variables; causation implies that changes in one variable directly produce changes in another. A strong correlation (e.g., r = 0.9) does not prove causation. Classic examples abound: shoe size and reading ability are correlated in children (confounded by age), but shoe size does not cause reading ability. Understanding when correlation implies causation requires evaluating study design, confounding variables, temporal precedence, and mechanism. Randomized experiments offer the strongest causal evidence; observational studies must carefully control for confounders.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Always ask: 'Does this study design support causal inference?' When reading observational research claiming causation without controlling for confounders, be skeptical. When designing a study where causal inference is important, use randomization if possible. When randomization is infeasible (e.g., studying effects of smoking on health), use observational designs with strong confounding control: measure potential confounders, adjust statistically, and discuss residual confounding limitations. Use causal graphs (Directed Acyclic Graphs, DAGs) to visualize confounding and guide adjustment strategies.
Strengths & limitations
- Prevents incorrect causal interpretations of correlational data, protecting against flawed policy and clinical decisions.
- Randomized experiments offer definitive causal evidence by breaking confounding through random assignment.
- Structured frameworks (Hill's criteria, DAGs) provide systematic approaches to evaluating causation in non-experimental studies.
- Highlights the importance of study design and confounding control, promoting more rigorous research planning.
- Randomized trials are often infeasible, unethical, or expensive (e.g., studying effects of smoking requires observational designs).
- Unmeasured confounders in observational studies can bias causal estimates; statistical adjustment cannot control for variables not measured.
- Causal inference from observational data requires strong assumptions (e.g., no unmeasured confounders, correct functional form) that are often unverifiable.
- Complex scenarios with feedback loops, selection bias, or time-varying confounders require advanced causal methods (instrumental variables, marginal structural models) that are less familiar to most researchers.
Frequently asked
Does a high correlation coefficient (e.g., r = 0.95) prove causation?
No. A very strong correlation is necessary but not sufficient for causation. Example: the number of firefighters at a fire is strongly correlated with fire damage (r might be 0.8+), but firefighters do not cause damage—they respond to large fires, which cause damage. Confounding by fire size explains the correlation. Always investigate alternative explanations before claiming causation.
What is confounding and how do I identify it?
Confounding occurs when a third variable (confounder) is associated with both the cause and effect, creating a spurious association. Example: age confounds the relation between shoe size (cause) and reading ability (effect) because age is associated with both. To identify confounders: (1) Think logically about what affects both variables. (2) Check if the association weakens after adjusting for the confounder statistically. (3) Use causal diagrams (DAGs) to visualize the causal structure.
Are randomized experiments always better than observational studies?
For establishing causation, yes—RCTs are the gold standard. But RCTs have limitations: they are expensive, may be unethical, and may not reflect real-world conditions (external validity). Observational studies are necessary for studying long-term effects, rare outcomes, and exposures that cannot be randomly assigned. Use RCTs for causal hypothesis testing; use well-designed observational studies for hypothesis generation and real-world effectiveness.
What is reverse causation and how do I avoid it?
Reverse causation occurs when you incorrectly infer the direction of causality. Example: does depression cause insomnia or does insomnia cause depression? Avoid reverse causation by: (1) Establishing temporal precedence with longitudinal designs (measure cause before outcome). (2) Using theory and prior research to guide causal direction. (3) Considering bidirectional relationships; some causal pathways are mutually reinforcing. Cross-sectional data cannot establish direction; use prospective cohort studies.
What are unmeasured confounders and can I account for them?
Unmeasured confounders are variables that confound the association but were not measured in the study. You cannot statistically control for variables you did not measure. Address unmeasured confounding by: (1) Measuring and adjusting for likely confounders in the study design. (2) Sensitivity analyses: assume different amounts of unmeasured confounding and see if conclusions change. (3) Using instrumental variables or natural experiments, which can estimate causal effects even with unmeasured confounding. (4) Discussing limitations honestly in the paper.
Sources
- Pearl, J. (2009). Causality: Models, Reasoning, and Inference (2nd ed.). Cambridge University Press. ISBN: 978-0-521-89560-6
- Rubin, D. B. (1974). Estimating causal effects of treatments in randomized and nonrandomized studies. Journal of Educational Psychology, 66(5), 688–701. DOI: 10.1037/h0037350 ↗
- Hill, A. B. (1965). The Environment and Disease: Association or Causation? Proceedings of the Royal Society of Medicine, 58(5), 295–300. DOI: 10.1177/003591576505800503 ↗
How to cite this page
ScholarGate. (2026, June 3). Understanding the Distinction Between Correlation and Causation in Research. ScholarGate. https://scholargate.app/en/research-statistics/correlation-vs-causation
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Effect SizeResearch Statistics↔ compare
- Multiple Comparisons ProblemResearch Statistics↔ compare
- Null Hypothesis TestingResearch Statistics↔ compare
- P-Value and Statistical SignificanceResearch Statistics↔ compare