Psychometric Meta-Analysis
Also known as: Hunter-Schmidt Meta-Analysis, Validity Generalization, Artifact-Corrected Meta-Analysis, VG
Psychometric meta-analysis is the Hunter-Schmidt approach to cumulating research findings while correcting for the statistical artifacts that distort individual studies. Frank Schmidt and John Hunter developed it to solve the problem of validity generalization: across many studies the observed validity of a selection test varied widely, leading people to conclude that validity was situationally specific, when in fact most of the variation was an illusion produced by small samples, unreliable measures, and restricted ranges. Their 1977 Journal of Applied Psychology paper showed that once these artifacts are removed, the apparent variability shrinks and a stable true validity emerges that generalizes across settings. The full method, codified in their book Methods of Meta-Analysis, pools effect sizes, subtracts the variance due to sampling error, and corrects the mean and remaining variance for measurement unreliability and range restriction. It estimates not only the average true effect but how much it really varies and whether it generalizes.
Key highlights
- Separates real effect-size variation from sampling error, revealing that effects often generalize where they appeared situationally specific.
- Corrects observed effects for measurement unreliability and range restriction, recovering unbiased estimates of true-score relationships.
- Provides a credibility interval that directly tests whether an effect holds across settings, a clear criterion for generalization.
- Anchored decades of cumulative evidence in personnel selection, transforming how the field judges the validity of selection methods.
Intuition
This section is available to Pro members. Upgrade to Pro
How it works
This section is available to Pro members. Upgrade to Pro
When to use it
Use psychometric meta-analysis when you want to cumulate quantitative findings — especially validity correlations — across many studies and you have, or can estimate, information about the artifacts (reliability and range restriction) that distort those findings. It is the method of choice in personnel selection and other areas where measures are imperfect and samples are selected, because it estimates true-score effects and tests whether they generalize. It is less appropriate when artifact information is entirely unavailable (limiting the corrections to sampling error), when the studies are too few or too heterogeneous in construct to pool meaningfully, or when the research question concerns moderators that random-effects meta-regression handles more directly. As with any meta-analysis, it cannot fix a biased or incomplete literature, so attention to publication bias and study quality remains essential.
Strengths & limitations
- Separates real effect-size variation from sampling error, revealing that effects often generalize where they appeared situationally specific.
- Corrects observed effects for measurement unreliability and range restriction, recovering unbiased estimates of true-score relationships.
- Provides a credibility interval that directly tests whether an effect holds across settings, a clear criterion for generalization.
- Anchored decades of cumulative evidence in personnel selection, transforming how the field judges the validity of selection methods.
- Artifact corrections require reliability and range-restriction information that primary studies often fail to report, forcing reliance on assumed artifact distributions.
- Corrections can be sensitive to the assumed artifact values, and aggressive correcting can overstate true-score effects if those values are wrong.
- Like all meta-analyses, it inherits publication bias and selective reporting from the underlying literature.
- The bare-bones, sampling-error-only version can understate true heterogeneity when artifact variation is not modeled.
Common pitfalls
This section is available to Pro members. Upgrade to Pro
Applications
This section is available to Pro members. Upgrade to Pro
Frequently asked
How does psychometric meta-analysis differ from ordinary (Hedges-Olkin) meta-analysis?
Both pool effect sizes across studies, but the Hunter-Schmidt psychometric approach additionally corrects for statistical artifacts — sampling error, measurement unreliability, and range restriction — to estimate the true-score effect, whereas the Hedges-Olkin tradition focuses on the observed effect and uses fixed- or random-effects models with formal variance estimators. Psychometric meta-analysis also emphasizes the credibility interval for judging generalization rather than significance testing of the mean. In practice the schools have converged, but the defining feature of the Hunter-Schmidt method remains its systematic artifact correction aimed at recovering relationships at the construct level.
What is validity generalization and how is it decided?
Validity generalization is the conclusion that a predictor's validity holds across settings rather than being situationally specific. It is decided by correcting the observed mean validity and its variance for sampling error and artifacts, then constructing a credibility interval around the corrected true validity. If the lower bound of that interval — for example the 80 percent credibility value — is clearly above zero, validity generalizes, because even the least favorable true effects are positive. Schmidt and Hunter used exactly this logic to show that cognitive ability test validities generalize across jobs, overturning the earlier belief in situational specificity.
Doesn't correcting for artifacts just inflate the effect size?
Corrections raise the estimate, but their purpose is to remove known downward biases, not to inflate results arbitrarily. Measurement unreliability and range restriction provably attenuate observed correlations, so dividing by the square roots of the reliabilities and the range-restriction factor recovers the relationship that would be observed with perfect measurement in an unrestricted population. The risk is that wrong artifact values yield wrong corrections, which is why Hunter and Schmidt stress using well-justified reliability and range-restriction information, reporting both observed and corrected results, and being transparent about the assumed artifact distributions.
Sources
- 1.Hunter, J. E., & Schmidt, F. L. (2004). Methods of Meta-Analysis: Correcting Error and Bias in Research Findings (2nd ed.). Sage Publications.ISBN 9781412904797
- 2.Schmidt, F. L., & Hunter, J. E. (1977). Development of a general solution to the problem of validity generalization. Journal of Applied Psychology, 62(5), 529-540.
You have read it. What now?
Cite this page
ScholarGate. (2026, June 23). Psychometric Meta-Analysis. ScholarGate. https://scholargate.app/organizational-behavior/psychometric-meta-analysis