Skip to contentScholarGate
LibraryBookshelfDeskReview StudioAssistant
Sign in
On this page
IntuitionHow it worksWhen to use itStrengths & limitationsCommon pitfallsApplicationsFrequently asked🔒 Read the full methodSourcesRelated methods
Cite this pageSpotted an issue on this page? Report or suggest a fix →
Home›Epidemiology›Retrospective Diagnostic Accuracy Study
Process / pipelineClinical / epidemiology

Retrospective Diagnostic Accuracy Study

Also known as: retrospective DAS, retrospective test accuracy study, retrospective index test evaluation, historical diagnostic accuracy study

A retrospective diagnostic accuracy study evaluates how well a diagnostic test (the index test) correctly identifies a target condition by applying it to previously collected data or archived specimens alongside a reference standard. Because both index test results and reference standard results are drawn from existing records or stored material rather than generated prospectively, this design is faster and less costly than a prospective counterpart — but carries specific methodological risks that must be controlled to produce valid estimates of sensitivity, specificity, and related measures.

ScholarGate
  1. Process / pipeline
  2. v1
  3. 2 Sources
  4. PUBLISHED
Cite this page →
Tools & resources
Download slides
Learn & explore

Read the full method

Members only

Sign in with a free account to read this section.

Sign in

Method map

The neighbourhood of related methods — select a node to explore.

Retrospective diagnostic accuracy study
Case-control studyDiagnostic Accuracy Stud…Prospective Diagnostic A…Retrospective Cohort Stu…Screening Test Evaluation

When to use it

A retrospective diagnostic accuracy study is appropriate when archived specimens, stored imaging, or existing medical records are available and include both the index test result (or material from which it can be freshly measured) and an independently ascertained reference standard result. It is well-suited to situations where prospective data collection would be ethically difficult, too slow, or prohibitively expensive, and where a preliminary accuracy estimate is needed to justify a larger prospective study. Do not use this design when the archived sample is highly selected (e.g., only patients with confirmed disease and known-negative controls) without representing the realistic clinical spectrum, as sensitivity and specificity estimates will be biased; in that circumstance, a prospective design or a nested case-control within a well-defined cohort is preferred. Also avoid when the reference standard was not applied independently of the index test result in the original records, creating verification bias.

Strengths & limitations

Strengths
  • Substantially faster and less expensive than prospective diagnostic accuracy studies because data collection has already occurred.
  • Can leverage large existing biobanks, registries, or electronic health records to study rare conditions with adequate sample sizes.
  • Avoids the ethical challenges of withholding a test from a prospective control group when the test is already in clinical use.
  • Allows evaluation of tests on historical specimens that are no longer obtainable (e.g., outbreak-era samples).
  • When high-quality records are available, can yield estimates comparable to prospective studies for well-defined patient populations.
Limitations
  • Spectrum bias is common: archived samples often over-represent severe or confirmed cases and under-represent mild or equivocal presentations, inflating sensitivity and specificity compared to real clinical practice.
  • Verification bias (partial verification) arises when not all participants received the reference standard — often because clinicians only sent high-suspicion patients for confirmatory testing.
  • The conditions under which the index test was originally performed may not be standardized, reducing reproducibility of extracted results.
  • Missing data on either index test results or reference standard outcomes can introduce selection bias if records are incomplete.
  • Blinding cannot always be guaranteed: clinicians who ordered the index test may have known the reference standard result when documenting test interpretation.

Frequently asked

What is the main difference between a retrospective and a prospective diagnostic accuracy study?

In a prospective design, participants are enrolled before the index test is applied, both the index test and reference standard are planned in advance, and the sequence of testing is controlled. In a retrospective design, both test results and reference standard outcomes come from data that already exist. The prospective design offers better control over spectrum of disease, standardization of test conditions, and verification completeness, but is slower and more expensive. The retrospective design is faster and cheaper but requires careful attention to spectrum bias and verification bias.

How do I handle participants for whom the reference standard result is missing?

Missing reference standard results must be reported transparently, not silently excluded. Describe how many participants are missing, why the reference standard was not applied, and whether the missingness is likely to be related to the index test result or disease status. Where feasible, perform a sensitivity analysis (e.g., assuming all missing cases are true positives or all true negatives) to bound the possible impact on accuracy estimates. A high proportion of missing reference standard results is a serious threat to validity.

Which quality assessment tool should I use when appraising retrospective diagnostic accuracy studies?

QUADAS-2 is the standard tool. It assesses four domains — patient selection, index test, reference standard, and flow and timing — and separately rates risk of bias and applicability concerns for each. For retrospective studies, particular scrutiny is warranted in the patient selection domain (was the sample representative of the intended clinical population?) and the flow and timing domain (did all patients receive the reference standard, and was it applied independently?).

Can I use a case-control design within a retrospective diagnostic accuracy framework?

A case-control selection — deliberately sampling confirmed disease cases and known-negative controls from a registry — is sometimes used for feasibility but inflates estimates of sensitivity and specificity relative to a consecutive or random sample. If you use this approach, report disease prevalence in the source population and apply predictive value calculations with caution; the artificially extreme spectrum means the estimates may not translate to routine clinical practice.

What sample size do I need?

Sample size should be calculated based on the expected sensitivity and specificity, the desired precision of confidence intervals, and the disease prevalence in the accessible archive. A common approach targets a 95% confidence interval half-width of no more than 0.05–0.10 for the primary accuracy measure. For rare diseases, a minimum of 30–50 cases is often cited as a lower bound to avoid very wide intervals, though larger samples are strongly preferred. Report the sample size justification explicitly in the methods.

Sources

  1. Bossuyt, P. M., Reitsma, J. B., Bruns, D. E., et al. (2015). STARD 2015: An Updated List of Essential Items for Reporting Diagnostic Accuracy Studies. BMJ, 351, h5527. DOI: 10.1136/bmj.h5527 ↗
  2. Whiting, P. F., Rutjes, A. W., Westwood, M. E., et al. (2011). QUADAS-2: A Revised Tool for the Quality Assessment of Diagnostic Accuracy Studies. Annals of Internal Medicine, 155(8), 529–536. DOI: 10.7326/0003-4819-155-8-201110180-00009 ↗

How to cite this page

ScholarGate. (2026, June 3). Retrospective Diagnostic Accuracy Study. ScholarGate. https://scholargate.app/en/epidemiology/retrospective-diagnostic-accuracy-study

Related methods

Case-control studyDiagnostic Accuracy Study DesignProspective Diagnostic Accuracy StudyRetrospective Cohort StudyScreening Test Evaluation

Which method?

Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.

  • Case-control studyEpidemiology↔ compare
  • Diagnostic Accuracy Study DesignClinical Research↔ compare
  • Prospective Diagnostic Accuracy StudyEpidemiology↔ compare
  • Retrospective Cohort StudyEpidemiology↔ compare
  • Screening Test EvaluationEpidemiology↔ compare
Compare side by side →

Referenced by

Prospective Diagnostic Accuracy Study

Similar methods

Prospective Diagnostic Accuracy StudyProspective Screening Test EvaluationDiagnostic Accuracy Study DesignMatched Diagnostic Accuracy StudyPragmatic diagnostic accuracy studyMulticenter Diagnostic Accuracy StudyMulticenter Screening Test EvaluationMatched Screening Test Evaluation

Related reference concepts

Screening and Diagnostic Test EvaluationSensitivityScreening Test Characteristics and PerformanceAnalytical Validation and Test AccuracyCase-Control StudyPositive Predictive Value

Spotted an issue on this page? Report or suggest a fix →

ScholarGate — Retrospective diagnostic accuracy study (Retrospective Diagnostic Accuracy Study). Retrieved 2026-07-21 from https://scholargate.app/en/epidemiology/retrospective-diagnostic-accuracy-study · Dataset: https://doi.org/10.5281/zenodo.20539026
Quick facts
Originator
Formalized through the STARD initiative led by Patrick Bossuyt and colleagues
Year
Formalized 2000s; STARD 2003, revised 2015
Type
Observational, retrospective study design
DataType
Previously collected biological specimens, medical records, imaging data, or stored samples with known reference standard results
Subfamily
Clinical / epidemiology
Related methods
Case-control studyDiagnostic Accuracy Study DesignProspective Diagnostic Accuracy StudyRetrospective Cohort StudyScreening Test Evaluation
ScholarGate

A content-first reference library for research methods — what each one is, how it works, and where it comes from.

Open data (CC-BY)

Explore

  • Library
  • Search the library…
  • Browse by field
  • Fields
  • Journey
  • Compare
  • Which method?

Reference

  • Subjects
  • Atlas
  • Glossary
  • Methodology
  • Philosophy

Your tools

  • Bookshelf
  • Desk
  • Chat

Company

  • About
  • Pricing
  • Contact
  • Suggest a method

Entries are compiled from published sources for reference. Verifying the accuracy and suitability of any information for your own use remains your responsibility.

© 2026 ScholarGate · A research-method reference library
  • Privacy
  • Cookies
  • Terms
  • Delete account