Skip to contentScholarGate
LibraryBookshelfDeskReview StudioAssistant
Sign in
On this page
IntuitionHow it worksWhen to use itStrengths & limitationsCommon pitfallsApplicationsFrequently asked🔒 Read the full methodSourcesRelated methods
Cite this pageSpotted an issue on this page? Report or suggest a fix →
Home›Clinical Research›Diagnostic Accuracy Study Design
Process / pipelinediagnostic evaluation

Diagnostic Accuracy Study Design

Diagnostic Test Accuracy Study (Sensitivity, Specificity, Likelihood Ratios) · Also known as: diagnostic accuracy study, test accuracy, STARD, diagnostic evaluation

A diagnostic accuracy study evaluates how well a new diagnostic test (or biomarker, imaging modality, clinical assessment) detects the presence or absence of disease compared to a reference standard (gold standard). Standardized since 2003 by the STARD (Standards for Reporting of Diagnostic Accuracy Studies) initiative, diagnostic accuracy studies are fundamental to clinical medicine, determining whether and how new tests can improve patient diagnosis and treatment.

ScholarGate
  1. Process / pipeline
  2. v1
  3. 3 Sources
  4. PUBLISHED
Cite this page →
Tools & resources
Download slides
Learn & explore

Read the full method

Members only

Sign in with a free account to read this section.

Sign in

Method map

The neighbourhood of related methods — select a node to explore.

Diagnostic Accuracy Study Design
Cohort Study DesignRandomized Controlled Tr…Adaptive Diagnostic Accu…Adaptive screening test…Bayesian Diagnostic Accu…Bayesian Screening Test…Case seriesCase-control studyCross-sectional epidemio…Matched Diagnostic Accur…

+17 more

When to use it

Use diagnostic accuracy studies when: (1) developing or evaluating a new diagnostic test, imaging modality, or biomarker, (2) comparing two diagnostic tests to determine which is superior, (3) determining whether a test can accurately detect disease in clinical populations, (4) assessing whether test performs similarly in subgroups (by age, sex, disease stage, comorbidity), (5) informing clinical practice guidelines on test adoption, (6) evaluating screening tests in asymptomatic populations, (7) assessing diagnostic tests in clinical decision-making (does it guide treatment choices?).

Strengths & limitations

Strengths
  • Direct and transparent: clearly answers 'Does this test accurately detect disease?'
  • Standardized methodology: STARD guidelines ensure consistent reporting and prevent bias.
  • Practical relevance: results directly inform clinical decision-making and test adoption.
  • Multiple metrics: sensitivity, specificity, predictive values, LRs, and ROC curves provide nuanced picture of test performance.
  • Subgroup assessment: can evaluate whether test works equally well across diverse populations (important for equity).
Limitations
  • Reference standard definition: 'gold standard' may not be perfect or universally agreed. If reference standard is flawed, whole study is flawed.
  • Spectrum bias: if study population does not represent clinical spectrum of disease (e.g., only severe cases, missing mild cases), accuracy estimates do not generalize.
  • Threshold selection: for continuous tests (e.g., biomarker levels), choice of cutoff determines sensitivity/specificity trade-off. Different thresholds yield different accuracy. Study must test range of thresholds or justify single cutoff.
  • Imperfect blinding: if test operator knows disease status, or reference standard assessor knows index test result, bias inflates accuracy.
  • Publication bias: negative or inconclusive accuracy studies may not be published, inflating published accuracy estimates.

Frequently asked

What is the difference between Sensitivity and Positive Predictive Value (PPV)?

Sensitivity = probability of positive test given disease is present. It asks: 'If I have the disease, will the test catch it?' High sensitivity means few false negatives (test rarely misses disease). PPV = probability of disease given positive test. It asks: 'If my test is positive, do I have disease?' PPV depends on disease prevalence: the same test has higher PPV in high-prevalence populations and lower PPV in low-prevalence populations. Sensitivity is fixed for a given test and threshold. For clinical practice, PPV often matters more (you want to know if positive result means disease), but it's population-dependent. Report both, and use Likelihood Ratios which are more stable across populations.

What is the ROC curve, and what does AUC mean?

ROC (Receiver Operating Characteristic) curve plots Sensitivity (true positive rate) vs 1−Specificity (false positive rate) as you vary the diagnostic threshold. For continuous tests (e.g., blood pressure, biomarker levels), each possible cutoff yields different sensitivity/specificity. The ROC curve shows this trade-off. AUC (Area Under Curve) is the area beneath the ROC curve. AUC = 0.5 means random guessing (diagonal line). AUC = 1.0 is perfect discrimination. Guidelines: AUC 0.5–0.7 = fair, 0.7–0.8 = acceptable, 0.8–0.9 = excellent, >0.9 = outstanding. AUC is threshold-independent and accounts for the full range of sensitivity/specificity trade-offs, useful for comparing tests.

What is a Likelihood Ratio, and why should I use it?

Likelihood Ratio (LR) is the ratio of sensitivity to specificity (or 1−specificity). LR+ = Sensitivity / (1−Specificity) tells how much a positive test increases disease probability. LR− = (1−Sensitivity) / Specificity tells how much a negative test decreases disease probability. For example, LR+ = 10 means a positive test makes disease 10 times more likely. LR− = 0.1 means a negative test makes disease 10 times less likely. Why LRs? They are threshold- and prevalence-independent: same LR applies in any population, unlike PPV/NPV which change with prevalence. Using Bayes' theorem, you convert pre-test probability (prevalence) to post-test probability via LR, directly answering 'Should I order this test?' and 'What does the result mean for my patient?'

What is the reference standard, and why must it be perfect?

The reference standard is the accepted 'gold standard' for diagnosing the disease: biopsy for cancer, angiography for CAD, PCR for COVID-19. The index test is compared against it. If the reference standard is imperfect (sometimes gives wrong answers), the entire study is compromised: you cannot know whether index test disagreements mean index test is wrong or reference standard is wrong. Perfect reference standard is ideal but often impossible. If reference standard has known imperfection, document it and discuss impact on results. Sensitivity and specificity estimated are technically relative to the imperfect reference, not true disease. Be transparent about reference standard limitations.

Sources

  1. Bossuyt, P. M., Reitsma, J. B., Bruns, D. E., Gatsonis, C. A., Glasziou, P. P., Irwig, L. M., ... & de Vet, H. C. (2003). Towards complete and accurate reporting of studies of diagnostic accuracy: the STARD initiative. Annals of Internal Medicine, 138(1), 40–44. DOI: 10.7326/0003-4819-138-1-200301070-00010 ↗
  2. Cohen, J. F., Korevaar, D. A., Altman, D. G., Bruns, D. E., Gatsonis, C. A., Hooft, L., ... & Bossuyt, P. M. (2016). STARD 2015 guidelines for reporting diagnostic accuracy studies: explanation and elaboration. BMJ Open, 6(11), e012799. DOI: 10.1136/bmjopen-2016-012799 ↗
  3. Jaeschke, R., Guyatt, G. H., & Sackett, D. L. (1994). Users' guides to the medical literature. III. How to use an article about a diagnostic test. A. Are the results of the study valid? JAMA, 271(5), 389–391. DOI: 10.1001/jama.1994.03510290071040 ↗

How to cite this page

ScholarGate. (2026, June 4). Diagnostic Test Accuracy Study (Sensitivity, Specificity, Likelihood Ratios). ScholarGate. https://scholargate.app/en/clinical-research/diagnostic-accuracy-study

Related methods

Cohort Study DesignRandomized Controlled Trial

Which method?

Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.

  • Cohort Study DesignClinical Research↔ compare
  • Randomized Controlled TrialExperimental design↔ compare
Compare side by side →

Referenced by

Adaptive Diagnostic Accuracy StudyAdaptive screening test evaluationBayesian Diagnostic Accuracy StudyBayesian Screening Test EvaluationCase seriesCase-control studyCross-sectional epidemiological studyMatched Diagnostic Accuracy StudyMatched Screening Test EvaluationMeta-analytic Diagnostic Accuracy StudyMulticenter Diagnostic Accuracy StudyMulticenter Screening Test EvaluationPhase I Clinical TrialPhase II clinical trialPragmatic diagnostic accuracy studyPragmatic Screening Test EvaluationProspective Case SeriesProspective Diagnostic Accuracy StudyProspective Screening Test EvaluationRetrospective diagnostic accuracy studyRetrospective phase II clinical trialRisk-adjusted case seriesRisk-adjusted diagnostic accuracy studyRisk-adjusted screening test evaluationScreening Test Evaluation

Similar methods

Sensitivity and SpecificityProspective Diagnostic Accuracy StudyProspective Screening Test EvaluationRetrospective diagnostic accuracy studyMatched Diagnostic Accuracy StudyPragmatic diagnostic accuracy studyMulticenter Diagnostic Accuracy StudyScreening Test Evaluation

Related reference concepts

Screening and Diagnostic Test EvaluationScreening Test Characteristics and PerformanceSensitivityReceiver Operating Characteristic CurveAnalytical Validation and Test AccuracyPositive Predictive Value

Spotted an issue on this page? Report or suggest a fix →

ScholarGate — Diagnostic Accuracy Study Design (Diagnostic Test Accuracy Study (Sensitivity, Specificity, Likelihood Ratios)). Retrieved 2026-07-20 from https://scholargate.app/en/clinical-research/diagnostic-accuracy-study · Dataset: https://doi.org/10.5281/zenodo.20539026
Quick facts
Originator
Bossuyt, Reitsma, and STARD group (2003); clinical epidemiology pioneers
Subfamily
diagnostic evaluation
Year
2003-2015
Type
Research Design
Related methods
Cohort Study DesignRandomized Controlled Trial
ScholarGate

A content-first reference library for research methods — what each one is, how it works, and where it comes from.

Open data (CC-BY)

Explore

  • Library
  • Search the library…
  • Browse by field
  • Fields
  • Journey
  • Compare
  • Which method?

Reference

  • Subjects
  • Atlas
  • Glossary
  • Methodology
  • Philosophy

Your tools

  • Bookshelf
  • Desk
  • Chat

Company

  • About
  • Pricing
  • Contact
  • Suggest a method

Entries are compiled from published sources for reference. Verifying the accuracy and suitability of any information for your own use remains your responsibility.

© 2026 ScholarGate · A research-method reference library
  • Privacy
  • Cookies
  • Terms
  • Delete account