Statistical Power and Sample Size
Statistical Power Analysis and Sample Size Determination for Research Studies · Also known as: power analysis, sample size calculation, 1 minus beta, sensitivity
Statistical power is the probability of detecting a true effect if it exists (1 − β). Power analysis determines the sample size required to detect a hypothesized effect size with specified Type I error (α) and Type II error (β) rates. Introduced by Jacob Cohen (1988), power analysis is foundational to research design: underpowered studies produce inflated effect size estimates and are unlikely to replicate. The standard benchmark is 80% power (β = 0.20), though critical studies may require 90% power.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Conduct power analysis before starting any empirical study to determine sample size. Required for grant proposals and clinical trial protocols. Recommended even for small pilot studies (though with lower power targets, e.g., 50–60%). Use post-hoc power analysis with caution—it often serves as an excuse for low power rather than a discovery tool. If your study is underpowered, acknowledge this limitation and interpret non-significant findings carefully (failing to detect ≠ proof of no effect).
Strengths & limitations
- Prevents wasteful underpowered studies that are unlikely to detect true effects and are prone to false positives.
- Enables responsible study planning: researchers know a priori whether they have resources for adequate sample size.
- Facilitates reproducibility: adequately powered studies are more likely to replicate.
- Supports transparent science: pre-specified power calculations enhance credibility and reduce p-hacking temptation.
- Allows cost-benefit analysis: researchers can weigh the cost of larger samples against the benefit of higher detection probability.
- Requires specification of expected effect size before the study; in novel areas, this is difficult and often guesswork.
- Assumes the effect size specified a priori is correct; if the true effect is smaller, the study remains underpowered.
- Complex designs (multilevel, multivariate, mixed models) have no simple closed-form power formulas; simulation-based approaches are needed.
- Assumes no missing data, violations of assumptions, or confounding; real studies often face these challenges.
- Does not guarantee the study will yield correct conclusions; power is only one of many factors affecting research quality.
Frequently asked
What is the standard statistical power level?
80% power (β = 0.20) is conventional, balancing the desire to detect true effects with practical constraints. This means accepting a 20% risk of a Type II error (false negative). For critical applications (e.g., regulatory approval of life-saving treatments), 90% power is often used. Low-risk exploratory studies may use 50–60% power. Always justify your choice.
How do I choose the expected effect size for power analysis?
Use: (1) Pilot data or prior literature on similar interventions. (2) Minimal clinically important difference (MCID), the smallest effect that is practically relevant. (3) Cohen's benchmarks (d = 0.2/0.5/0.8) as a last resort. Never use p-hacking results or inflated effect sizes from underpowered studies. If in doubt, use a medium effect size (d = 0.5) as a conservative target.
Can I conduct a study with 60% power if resources are limited?
Yes, but acknowledge the limitation. Low power increases the risk of Type II errors and inflates effect size estimates (survivor bias). If you conduct an underpowered study, (a) pre-register it, (b) report the low power, and (c) interpret non-significant results carefully ('failing to find an effect' ≠ 'proving no effect'). Combine with replication to increase credibility.
Is post-hoc power analysis useful?
No. Post-hoc power (power calculated after seeing the data) is mathematically redundant with the p-value and is often misleading. A non-significant result with 'low post-hoc power' tells you the study was underpowered—nothing more. Plan power a priori. If you conduct an underpowered study, report it honestly and discuss limitations.
How does sample size affect effect size estimates?
In underpowered studies, only the largest effect sizes are detected (survivor bias), inflating the mean effect size estimate. This is called the 'decline effect' or 'replication failure.' Larger samples and higher power produce more accurate effect size estimates. This is why meta-analyses often find smaller average effects than early underpowered studies.
Sources
- Cohen, J. (1988). Statistical Power Analysis for the Behavioral Sciences (2nd ed.). Lawrence Erlbaum Associates. ISBN: 0-8058-0283-5
- Faul, F., Erdfelder, E., Lang, A.-G., & Buchner, A. (2007). G*Power 3: A Flexible Statistical Power Analysis Program for the Social, Behavioral, and Biomedical Sciences. Behavior Research Methods, 39(2), 175–191. DOI: 10.3758/BF03193146 ↗
- Button, K. S., Ioannidis, J. P. A., Mokrysz, C., Nosek, B. A., Flint, J., Robinson, E. S. J., & Munafò, M. R. (2013). Power failure: why small sample size undermines the reliability of neuroscience. Nature Reviews Neuroscience, 14(5), 365–376. DOI: 10.1038/nrn3475 ↗
How to cite this page
ScholarGate. (2026, June 3). Statistical Power Analysis and Sample Size Determination for Research Studies. ScholarGate. https://scholargate.app/en/research-statistics/statistical-power
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Effect SizeResearch Statistics↔ compare
- Null Hypothesis TestingResearch Statistics↔ compare
- P-Value and Statistical SignificanceResearch Statistics↔ compare
- Type I and Type II ErrorsResearch Statistics↔ compare