Skip to contentScholarGate
LibraryBookshelfDeskReview StudioAssistant
Sign in
On this page
IntuitionHow it worksWhen to use itStrengths & limitationsCommon pitfallsApplicationsFrequently asked🔒 Read the full methodSourcesRelated methods
Cite this pageSpotted an issue on this page? Report or suggest a fix →
Home›Research Statistics›Sensitivity and Specificity
Process / pipelinediagnostic-testing

Sensitivity and Specificity

Sensitivity and Specificity in Diagnostic Testing and Binary Classification · Also known as: diagnostic accuracy, true positive rate, true negative rate, receiver operating characteristic

Sensitivity and specificity are fundamental metrics of diagnostic test accuracy. Sensitivity is the probability that a test correctly identifies a person with the disease (true positive rate: TP / (TP + FN)). Specificity is the probability that a test correctly identifies a person without the disease (true negative rate: TN / (TN + FP)). Every test involves a trade-off: increasing sensitivity (catching all sick people) often reduces specificity (more false alarms). Choice of test threshold depends on the clinical context: screening for serious diseases favors sensitivity; confirming a diagnosis favors specificity.

ScholarGate
  1. Process / pipeline
  2. v1
  3. 3 Sources
  4. PUBLISHED
Cite this page →
Tools & resources
Download slides
Learn & explore

Read the full method

Members only

Sign in with a free account to read this section.

Sign in

Method map

The neighbourhood of related methods — select a node to explore.

Sensitivity and Specificity
Effect SizeNull Hypothesis TestingP-Value and Statistical…Type I and Type II ErrorsROC analysis

When to use it

Use sensitivity and specificity to evaluate diagnostic tests in research and clinical settings. Sensitivity is critical for screening tests that aim to identify all potential cases (e.g., mammography for breast cancer, colonoscopy for colorectal cancer), where missing disease is costly. Specificity is critical for confirmatory tests that follow positive screening (e.g., biopsy), where false positives are costly. Use Receiver Operating Characteristic (ROC) curves to visualize the trade-off between sensitivity and specificity across different test thresholds, and to compare multiple tests. Optimal threshold depends on clinical context: if false negatives are catastrophic (missing cancer), prioritize sensitivity; if false positives are costly (unnecessary biopsies), prioritize specificity.

Strengths & limitations

Strengths
  • Provides clear, interpretable metrics of test accuracy that are independent of disease prevalence.
  • Sensitivity and specificity are stable properties of a test across different populations (unlike PPV/NPV, which depend on prevalence).
  • ROC curves enable visual comparison of multiple tests and selection of an optimal threshold.
  • Allows informed clinical decision-making by quantifying the probability of correct diagnosis.
  • Facilitates communication with patients about test limitations: e.g., 'This test catches 95% of cases but has a 10% false positive rate.'
Limitations
  • Sensitivity and specificity are often reported without information about prevalence and PPV/NPV, which are more relevant to clinical practice.
  • The choice of 'gold standard' affects estimates; if the gold standard is imperfect, sensitivity and specificity estimates are biased.
  • Sensitivity and specificity depend on the population tested; they may differ in populations with different disease severity, comorbidities, or demographics.
  • Trade-off between sensitivity and specificity requires choosing a threshold; different thresholds yield different metrics, and there is no universally optimal threshold.
  • Small sample size (few diseased individuals or few controls) yields unstable estimates with wide confidence intervals.

Frequently asked

What is the difference between sensitivity and positive predictive value (PPV)?

Sensitivity: probability a test is positive given disease is present [TP/(TP+FN)]. PPV: probability disease is present given test is positive [TP/(TP+FP)]. Sensitivity is a property of the test; PPV depends on disease prevalence. Example: a test with 90% sensitivity and 85% specificity has very different PPVs in a population where disease prevalence is 50% vs. 1%. Always report both sensitivity/specificity and PPV/NPV to fully characterize a test.

Which is more important: sensitivity or specificity?

Depends on context. Screening tests prioritize sensitivity: you want to catch all potential cases, even if some are false positives (these are then confirmed by a specific test). Diagnostic confirmation prioritizes specificity: you want high confidence disease is present before starting treatment. A good diagnostic pathway combines a sensitive screening test and a specific confirmatory test.

How does disease prevalence affect sensitivity and specificity?

Sensitivity and specificity are properties of the test itself and do NOT depend on prevalence. However, PPV and NPV depend heavily on prevalence. In a rare disease, even a specific test may have low PPV (most positives are false). In a common disease, even a less-specific test may have high PPV. This is Bayes' theorem. Always consider prevalence when interpreting test results clinically.

What is a ROC curve and how do I use it?

A Receiver Operating Characteristic (ROC) curve plots sensitivity (y-axis) vs. 1 - specificity (x-axis) as the test threshold varies. Each point on the curve represents a different threshold. The Area Under the Curve (AUC) quantifies overall accuracy: AUC = 1 (perfect test), 0.5 (random guessing). ROC curves allow: (1) Visual comparison of tests (test with curve furthest to upper-left is best). (2) Selection of optimal threshold based on clinical goals. (3) Calculation of AUC as a summary statistic.

If a test has 95% sensitivity and 90% specificity, is it a good test?

Not necessarily—it depends on prevalence, consequences of errors, and the clinical context. Also ask: (1) In the population of interest, what is prevalence? (2) What are PPV and NPV? (3) Are false positives or false negatives more costly? (4) How does this test compare to alternatives? A test with these metrics might be good for screening a common disease but inadequate for confirming a rare disease.

Sources

  1. Altman, D. G., & Bland, J. M. (1994). Diagnostic tests 1: Sensitivity and specificity. BMJ, 308(6943), 1552. link ↗
  2. Fawcett, T. (2006). An introduction to ROC analysis. Pattern Recognition Letters, 27(8), 861–874. DOI: 10.1016/j.patrec.2005.10.010 ↗
  3. Metz, C. E. (1978). Basic principles of ROC analysis. Seminars in Nuclear Medicine, 8(4), 283–298. DOI: 10.1016/S0001-2998(78)80014-2 ↗

How to cite this page

ScholarGate. (2026, June 3). Sensitivity and Specificity in Diagnostic Testing and Binary Classification. ScholarGate. https://scholargate.app/en/research-statistics/sensitivity-specificity

Related methods

Effect SizeNull Hypothesis TestingP-Value and Statistical SignificanceType I and Type II Errors

Which method?

Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.

  • Effect SizeResearch Statistics↔ compare
  • Null Hypothesis TestingResearch Statistics↔ compare
  • P-Value and Statistical SignificanceResearch Statistics↔ compare
  • Type I and Type II ErrorsResearch Statistics↔ compare
Compare side by side →

Referenced by

ROC analysis

Similar methods

Diagnostic Accuracy Study DesignSpecificityRecall (Sensitivity)Bayesian Screening Test EvaluationROC analysisScreening Test EvaluationPrecisionProspective Screening Test Evaluation

Related reference concepts

SensitivityScreening Test Characteristics and PerformanceSpecificityScreening and Diagnostic Test EvaluationReceiver Operating Characteristic CurvePositive Predictive Value

Spotted an issue on this page? Report or suggest a fix →

ScholarGate — Sensitivity and Specificity (Sensitivity and Specificity in Diagnostic Testing and Binary Classification). Retrieved 2026-07-21 from https://scholargate.app/en/research-statistics/sensitivity-specificity · Dataset: https://doi.org/10.5281/zenodo.20539026
Quick facts
Originator
Multiple sources in medical diagnosis and signal detection
Subfamily
diagnostic-testing
Year
1978
Type
Concept
Related methods
Effect SizeNull Hypothesis TestingP-Value and Statistical SignificanceType I and Type II Errors
ScholarGate

A content-first reference library for research methods — what each one is, how it works, and where it comes from.

Open data (CC-BY)

Explore

  • Library
  • Search the library…
  • Browse by field
  • Fields
  • Journey
  • Compare
  • Which method?

Reference

  • Subjects
  • Atlas
  • Glossary
  • Methodology
  • Philosophy

Your tools

  • Bookshelf
  • Desk
  • Chat

Company

  • About
  • Pricing
  • Contact
  • Suggest a method

Entries are compiled from published sources for reference. Verifying the accuracy and suitability of any information for your own use remains your responsibility.

© 2026 ScholarGate · A research-method reference library
  • Privacy
  • Cookies
  • Terms
  • Delete account