Skip to contentScholarGate
LibraryBookshelfDeskReview StudioAssistant
Sign in
On this page
IntuitionHow it worksWhen to use itStrengths & limitationsCommon pitfallsApplicationsFrequently asked🔒 Read the full methodSourcesRelated methods
Cite this pageSpotted an issue on this page? Report or suggest a fix →
Home›Epidemiology›Screening Test Evaluation — Evaluating Screening Tests and Programs
Process / pipelineClinical / epidemiology

Screening Test Evaluation — Evaluating Screening Tests and Programs

Evaluation of Screening Tests and Screening Programs · Also known as: screening study, screening performance evaluation, screening accuracy assessment, STE

Screening test evaluation is a systematic epidemiological approach for assessing whether a test or program can accurately and cost-effectively identify individuals with a condition before symptoms appear. It quantifies diagnostic performance metrics — sensitivity, specificity, predictive values, and the ROC curve — and evaluates whether a screening program meets established public health criteria for adoption and harm-benefit balance.

ScholarGate
  1. Process / pipeline
  2. v1
  3. 2 Sources
  4. PUBLISHED
Cite this page →
Tools & resources
Download slides
Learn & explore

Read the full method

Members only

Sign in with a free account to read this section.

Sign in

Method map

The neighbourhood of related methods — select a node to explore.

Screening Test Evaluation
Case-control studyCohort StudyCross-sectional epidemio…Diagnostic Accuracy Stud…Dose-Response AnalysisKaplan-Meier AnalysisAdaptive Diagnostic Accu…Adaptive Dose-Response A…Bayesian Diagnostic Accu…Bayesian Screening Test…

+12 more

When to use it

Use screening test evaluation when you need to determine whether a test or program should be adopted, continued, or modified for detecting a condition in an asymptomatic or pre-symptomatic population. It is appropriate when an established reference standard exists and when the research question goes beyond simple diagnostic accuracy to include population-level impact. Do not use it as a substitute for a diagnostic accuracy study when the goal is purely to measure test performance in symptomatic patients — in that context a standard diagnostic accuracy study design (STARD guidelines) is more appropriate. Screening evaluation is also not suitable when no effective early intervention exists, since accurate detection without treatment benefit does not justify a screening program.

Strengths & limitations

Strengths
  • Provides a comprehensive statistical framework covering all major performance dimensions: sensitivity, specificity, predictive values, and ROC-AUC.
  • Explicitly incorporates population-level harm-benefit analysis, not just test accuracy, making findings actionable for public health policy.
  • The Wilson-Jungner criteria provide a principled, internationally recognized checklist for deciding whether screening is appropriate.
  • Can accommodate continuous biomarkers, imaging scores, or questionnaire instruments by using ROC analysis to select the optimal threshold.
  • Supports cost-effectiveness modeling as a natural extension of program-level assessment.
Limitations
  • Requires a reliable reference standard, which may itself be costly, invasive, or imperfect — imperfect reference standards bias sensitivity and specificity estimates.
  • Predictive values (PPV, NPV) are population-specific and change with disease prevalence; values from one population may not transfer to another.
  • Lead-time bias can make screening appear to extend survival even when it does not; length-time bias inflates the apparent benefit of detecting slowly progressing cases.
  • Overdiagnosis — detecting conditions that would never have caused symptoms — is difficult to quantify and is often underestimated in evaluations.
  • Full program assessment requires long follow-up and large samples, making it resource-intensive.

Frequently asked

What is the difference between a screening test evaluation and a diagnostic accuracy study?

A diagnostic accuracy study assesses test performance in patients who already have symptoms or clinical suspicion of a condition. Screening test evaluation is applied to asymptomatic or general populations and adds program-level considerations — prevalence effects on predictive values, lead-time and length-time bias, overdiagnosis, coverage, and cost-effectiveness. Both use the same core metrics (sensitivity, specificity, AUC), but the context, study design, and interpretation differ substantially.

Why do positive predictive values matter more than sensitivity in low-prevalence populations?

In a low-prevalence population, even a test with 95% specificity will generate many false positives because the majority of the population is truly negative. PPV tells patients and clinicians the probability that a positive test result reflects true disease — which is what matters clinically. A test with 90% sensitivity and 90% specificity applied to a population with 1% prevalence yields a PPV of only about 8%; nine out of ten positives would be false alarms.

What is lead-time bias and why does it matter?

Lead time is the interval between screen-detected diagnosis and the time the disease would have been diagnosed clinically due to symptoms. If survival is measured from diagnosis, screened patients appear to survive longer even if they die at the same time as unscreened patients — their apparent longer survival is merely an artifact of earlier detection, not a true benefit. Evaluations must use disease-specific mortality in the screened versus control population, not survival from diagnosis, to avoid this bias.

How do I choose the best cut-off on a ROC curve?

There is no universally optimal threshold — the choice depends on the relative clinical and social costs of false positives and false negatives. Common approaches include the Youden index (maximizes sensitivity + specificity − 1), cost-benefit analysis weighting the two error types, or selecting a fixed sensitivity (e.g., 90%) first and then reading off the corresponding specificity. For serious conditions where missing a case is catastrophic, a high-sensitivity threshold is usually preferred even at the cost of more false positives.

What are the Wilson-Jungner criteria?

The Wilson-Jungner criteria are ten conditions proposed in the 1968 WHO report that should be met before a population screening program is introduced. Key criteria include: the condition should be an important health problem; there should be an accepted treatment for detected disease; facilities for diagnosis and treatment must be available; there should be a recognizable latent or early symptomatic stage; a suitable test or examination must exist; the test must be acceptable to the population; the natural history of the condition should be understood; and the cost should be balanced against the overall expenditure on medical care. These criteria remain foundational to screening policy worldwide.

Sources

  1. Wilson, J. M. G., & Jungner, G. (1968). Principles and Practice of Screening for Disease. World Health Organization. Public Health Papers No. 34. link ↗
  2. Pepe, M. S. (2003). The Statistical Evaluation of Medical Tests for Classification and Prediction. Oxford University Press. ISBN: 978-0198565826

How to cite this page

ScholarGate. (2026, June 3). Evaluation of Screening Tests and Screening Programs. ScholarGate. https://scholargate.app/en/epidemiology/screening-test-evaluation

Related methods

Case-control studyCohort StudyCross-sectional epidemiological studyDiagnostic Accuracy Study DesignDose-Response AnalysisKaplan-Meier Analysis

Which method?

Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.

  • Case-control studyEpidemiology↔ compare
  • Cohort StudyEpidemiology↔ compare
  • Cross-sectional epidemiological studyEpidemiology↔ compare
  • Diagnostic Accuracy Study DesignClinical Research↔ compare
  • Dose-Response AnalysisEpidemiology↔ compare
  • Kaplan-Meier AnalysisEpidemiology↔ compare
Compare side by side →

Referenced by

Adaptive Diagnostic Accuracy StudyAdaptive Dose-Response AnalysisBayesian Diagnostic Accuracy StudyBayesian Screening Test EvaluationCross-sectional epidemiological studyMatched Screening Test EvaluationMeta-analytic Diagnostic Accuracy StudyMulticenter Diagnostic Accuracy StudyMulticenter Screening Test EvaluationPragmatic diagnostic accuracy studyPragmatic phase IV studyPragmatic Screening Test EvaluationProspective Diagnostic Accuracy StudyProspective Screening Test EvaluationRetrospective diagnostic accuracy studyRisk-adjusted diagnostic accuracy studyRisk-adjusted screening test evaluation

Similar methods

Prospective Screening Test EvaluationRisk-adjusted screening test evaluationMulticenter Screening Test EvaluationDiagnostic Accuracy Study DesignSensitivity and SpecificityPragmatic Screening Test EvaluationBayesian Screening Test EvaluationMatched Screening Test Evaluation

Related reference concepts

Screening and Diagnostic Test EvaluationScreening Test Characteristics and PerformanceScreening Methodology and PrinciplesScreening Principles and Test EvaluationSecondary Prevention and ScreeningWilson-Jungner Criteria for Screening Programs

Spotted an issue on this page? Report or suggest a fix →

ScholarGate — Screening Test Evaluation (Evaluation of Screening Tests and Screening Programs). Retrieved 2026-07-20 from https://scholargate.app/en/epidemiology/screening-test-evaluation · Dataset: https://doi.org/10.5281/zenodo.20539026
Quick facts
Originator
Wilson & Jungner (WHO criteria, 1968); foundational work by Pepe, Altman, and others in statistical test evaluation
Year
1968 (Wilson-Jungner principles); statistical framework developed 1970s–2000s
Type
Observational diagnostic / epidemiological evaluation design
DataType
Binary or continuous test results, reference standard (gold standard) outcomes, population-level screening data
Subfamily
Clinical / epidemiology
Related methods
Case-control studyCohort StudyCross-sectional epidemiological studyDiagnostic Accuracy Study DesignDose-Response AnalysisKaplan-Meier Analysis
ScholarGate

A content-first reference library for research methods — what each one is, how it works, and where it comes from.

Open data (CC-BY)

Explore

  • Library
  • Search the library…
  • Browse by field
  • Fields
  • Journey
  • Compare
  • Which method?

Reference

  • Subjects
  • Atlas
  • Glossary
  • Methodology
  • Philosophy

Your tools

  • Bookshelf
  • Desk
  • Chat

Company

  • About
  • Pricing
  • Contact
  • Suggest a method

Entries are compiled from published sources for reference. Verifying the accuracy and suitability of any information for your own use remains your responsibility.

© 2026 ScholarGate · A research-method reference library
  • Privacy
  • Cookies
  • Terms
  • Delete account