Effect Size
Effect Size: Quantifying the Magnitude of Research Findings · Also known as: ES, Cohen's d, standardized effect, practical significance
Effect size quantifies the magnitude of a research finding independent of sample size. While a p-value tells you whether a result is statistically significant, an effect size tells you how big the result is. Jacob Cohen formalized effect size measurement in behavioral sciences (1988), establishing standard benchmarks (small = 0.2, medium = 0.5, large = 0.8 for Cohen's d). Effect sizes are essential for meta-analysis, power analysis, and communicating the practical importance of research findings.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Report effect sizes in every empirical research paper, alongside p-values and confidence intervals. Mandatory for meta-analysis (effect sizes are pooled across studies). Use effect sizes to plan sample sizes (power analysis requires an expected effect size). Essential in fields where practical significance differs sharply from statistical significance (e.g., medicine, psychology, education). When the goal is to understand relationships, not just test hypotheses, effect sizes are more informative than p-values.
Strengths & limitations
- Sample-size independent: effect size is not inflated by large sample sizes, unlike p-values.
- Enables meta-analysis: effect sizes from multiple studies can be pooled using standard statistical techniques.
- Facilitates power analysis: researchers can determine required sample size based on the expected effect size.
- Improves interpretation: effect size directly communicates the practical or clinical magnitude of the finding.
- Standardized across studies: common metrics (Cohen's d, Pearson r) allow comparison of results across different research contexts.
- Benchmarks (small/medium/large) are somewhat arbitrary; they originated in psychology and may not apply equally to all disciplines.
- Interpretation requires domain knowledge: a 'small' effect size may be important in medicine but negligible in education.
- Assumes the outcome variable is measured on a comparable scale; effect sizes for ordinal or categorical variables are less intuitive.
- For non-normal data or outlier-heavy distributions, standardized effect sizes may not accurately reflect practical importance.
- Different effect size metrics (d, r, OR, RR) apply to different designs; researchers must choose the correct metric.
Frequently asked
Is effect size the same as practical significance?
Not exactly. Effect size is a quantitative measure of magnitude; practical significance is a judgment about whether that magnitude matters in real-world contexts. A large effect size (d = 0.8) is usually practically significant, but context matters. Example: a new medication that reduces pain by 1 point on a 100-point scale has a tiny effect size but may be practically significant if combined with other treatments. Always report effect size and let stakeholders judge practical importance.
What are Cohen's d benchmarks and do they apply to all fields?
Cohen (1988) suggested: d = 0.2 (small), 0.5 (medium), 0.8 (large). These were derived from behavioral sciences. Many fields report that Cohen's conventions do not fit their context. Medical research often considers small effects important; large-scale education studies may find small effect sizes to be policy-relevant. Always contextualize against your field's standards and prior literature, not blindly to Cohen's benchmarks.
Can I calculate effect size from a p-value?
Not precisely, but you can approximate it. For a t-test, you can convert t-statistic and degrees of freedom to Cohen's d: d = t × √(2/n). Conversely, if you know effect size and sample size, you can estimate the t-statistic and approximate the p-value. However, this is inferior to reporting the actual observed effect size from your data. Always calculate effect size directly from your data.
Should I always report the largest effect size possible?
No. Report the effect size metric most relevant to your research question and study design. For comparing two group means, use Cohen's d or Hedges' g. For correlations, use r. For binary outcomes, use Odds Ratio or Risk Ratio. Using multiple effect size metrics without justification (p-hacking equivalent for effect sizes) can inflate findings. Pre-specify the metric a priori.
What is the difference between Cohen's d and Hedges' g?
Cohen's d uses the sample standard deviation; Hedges' g applies a small-sample correction, making it slightly smaller than d for small samples. Both are standardized, comparable across studies. For large samples (n > 50), they are nearly identical. Use Hedges' g for meta-analysis of small-sample studies; either is acceptable for most purposes. Report which you used.
Sources
- Cohen, J. (1988). Statistical Power Analysis for the Behavioral Sciences (2nd ed.). Lawrence Erlbaum Associates. ISBN: 0-8058-0283-5
- Cumming, G. (2012). Understanding the New Statistics: Effect Sizes, Confidence Intervals, and Meta-Analysis. Routledge. ISBN: 0-415-87968-8
- Lakens, D. (2013). Calculating and Reporting Effect Sizes to Facilitate Cumulative Science: A Practical Primer for t-Tests and ANOVAs. Frontiers in Psychology, 4, 863. DOI: 10.3389/fpsyg.2013.00863 ↗
How to cite this page
ScholarGate. (2026, June 3). Effect Size: Quantifying the Magnitude of Research Findings. ScholarGate. https://scholargate.app/en/research-statistics/effect-size
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Confidence IntervalResearch Statistics↔ compare
- P-Value and Statistical SignificanceResearch Statistics↔ compare
- Statistical Power and Sample SizeResearch Statistics↔ compare
- Type I and Type II ErrorsResearch Statistics↔ compare