Skip to contentScholarGate
LibraryBookshelfDeskReview StudioAssistant
Sign in
On this page
IntuitionHow it worksWhen to use itStrengths & limitationsCommon pitfallsApplicationsFrequently asked🔒 Read the full methodSourcesRelated methods
Cite this pageSpotted an issue on this page? Report or suggest a fix →
Home›Epidemiology›Multicenter Diagnostic Accuracy Study
Process / pipelineClinical / epidemiology

Multicenter Diagnostic Accuracy Study

Also known as: multisite diagnostic accuracy study, multicenter DTA study, multicenter index test evaluation, STARD multicenter study

A multicenter diagnostic accuracy study evaluates how well an index test (e.g., a biomarker, imaging modality, or clinical prediction rule) identifies a target condition when conducted across two or more independent clinical sites. By recruiting patients from diverse settings, it produces estimates of sensitivity, specificity, and likelihood ratios that are more externally valid than those obtained from a single center, and it enables explicit assessment of how test performance varies across sites, patient populations, and operator skill levels.

ScholarGate
  1. Process / pipeline
  2. v1
  3. 2 Sources
  4. PUBLISHED
Cite this page →
Tools & resources
Download slides
Learn & explore

Read the full method

Members only

Sign in with a free account to read this section.

Sign in

Method map

The neighbourhood of related methods — select a node to explore.

Multicenter Diagnostic Accuracy Study
Diagnostic Accuracy Stud…Meta-analytic Diagnostic…Multicenter cohort studyMulticenter Randomized C…Screening Test Evaluation

When to use it

Use a multicenter diagnostic accuracy study when you need externally valid estimates of sensitivity and specificity for a test intended for use across diverse clinical settings, patient populations, or operator skill levels. It is particularly appropriate for tests approaching regulatory approval or clinical guideline adoption, where single-center data would be insufficient to convince guideline developers or health technology assessors. Do not use this design when the test is still in early-stage development (a single-center feasibility study is more efficient at that stage), when assembling a common reference standard across sites is logistically impossible, or when patient numbers at each site would be too small to detect meaningful heterogeneity — in that case, a pooled individual patient data meta-analysis of single-center studies may be preferable.

Strengths & limitations

Strengths
  • Produces more generalisable accuracy estimates by enrolling patients from diverse settings and populations.
  • Enables formal assessment of between-site heterogeneity, revealing whether a test is robust or context-dependent.
  • Larger total sample sizes are achievable, improving precision and statistical power to detect subgroup differences.
  • Findings are more credible to guideline developers, regulators, and clinicians than single-center data.
  • Allows head-to-head comparison of operator or laboratory performance across sites.
Limitations
  • Logistically demanding: requires coordinated protocols, centralised data management, and harmonised reference-standard procedures across all sites.
  • Between-site variation in patient mix and test application can create heterogeneity that complicates pooled estimates and their interpretation.
  • Higher cost and longer recruitment timelines than single-center studies, making them unsuitable for early-phase test development.
  • Verification bias can arise if the reference standard is not applied uniformly across sites or patient subgroups.

Frequently asked

How many sites do I need for a multicenter diagnostic accuracy study?

There is no universally fixed minimum, but most methodologists recommend at least five to ten sites to enable meaningful assessment of between-site heterogeneity. Fewer sites reduce the ability to detect systematic differences in test performance across settings. The required number also depends on total sample size targets: if each site can only contribute a small number of patients, more sites are needed to achieve adequate precision.

What statistical model should I use to pool results across sites?

The bivariate random-effects model (Reitsma et al., 2005) and the hierarchical summary ROC (HSROC) model (Rutter and Gatsonis, 2001) are the two standard approaches. Both account for the natural negative correlation between sensitivity and specificity and for between-study (between-site) heterogeneity. Software implementations are available in SAS (PROC NLMIXED), R (mada, meta4diag), and Stata.

How do I handle sites that use slightly different versions of the index test?

Document all index-test versions and quantify technical differences before analysis. If differences are minor (e.g., same assay platform, different reagent lot), sensitivity analyses excluding discrepant sites can assess impact. If differences are substantial (e.g., qualitatively different platforms), the study is effectively comparing two tests and should be analysed and reported as such, with test version as a key subgroup variable.

Is registration of a diagnostic accuracy study mandatory?

Registration is not universally mandated by law, but STARD 2015 strongly recommends it and a growing number of journals require it as a condition of submission. Registration in PROSPERO (for systematic reviews) or ClinicalTrials.gov (for prospective studies) improves transparency, reduces publication bias, and is increasingly expected by regulators and guideline developers.

How is a multicenter diagnostic accuracy study different from a meta-analysis of single-center studies?

In a multicenter study, all sites follow a common prospective protocol with centralised data management, harmonised reference standards, and pre-specified analysis plans. This eliminates many sources of between-study heterogeneity that plague retrospective meta-analyses. A meta-analysis of single-center studies pools independently conducted studies and must contend with heterogeneous patient selection, reference standards, and reporting quality. When feasible, a prospective multicenter study is methodologically superior.

Sources

  1. Bossuyt, P. M., Reitsma, J. B., Bruns, D. E., Gatsonis, C. A., Glasziou, P. P., Irwig, L., Lijmer, J. G., Moher, D., Rennie, D., de Vet, H. C. W., Kressel, H. Y., Rifai, N., Golub, R. M., Altman, D. G., Hooft, L., Korevaar, D. A., & Cohen, J. F. (2015). STARD 2015: An Updated List of Essential Items for Reporting Diagnostic Accuracy Studies. BMJ, 351, h5527. DOI: 10.1136/bmj.h5527 ↗
  2. Rutjes, A. W. S., Reitsma, J. B., Coomarasamy, A., Khan, K. S., & Bossuyt, P. M. M. (2006). Evaluation of diagnostic tests when there is no gold standard: A review of methods. Health Technology Assessment, 10(50), iii–iv, 1–121. DOI: 10.3310/hta11500 ↗

How to cite this page

ScholarGate. (2026, June 3). Multicenter Diagnostic Accuracy Study. ScholarGate. https://scholargate.app/en/epidemiology/multicenter-diagnostic-accuracy-study

Related methods

Diagnostic Accuracy Study DesignMeta-analytic Diagnostic Accuracy StudyMulticenter cohort studyMulticenter Randomized Clinical TrialScreening Test Evaluation

Which method?

Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.

  • Diagnostic Accuracy Study DesignClinical Research↔ compare
  • Meta-analytic Diagnostic Accuracy StudyEpidemiology↔ compare
  • Multicenter cohort studyEpidemiology↔ compare
  • Multicenter Randomized Clinical TrialEpidemiology↔ compare
  • Screening Test EvaluationEpidemiology↔ compare
Compare side by side →

Similar methods

Multicenter Screening Test EvaluationProspective Diagnostic Accuracy StudyAdaptive Diagnostic Accuracy StudyProspective Screening Test EvaluationPragmatic diagnostic accuracy studyDiagnostic Accuracy Study DesignRetrospective diagnostic accuracy studyMeta-analytic Diagnostic Accuracy Study

Related reference concepts

Screening and Diagnostic Test EvaluationAnalytical Validation and Test AccuracyScreening Test Characteristics and PerformanceReceiver Operating Characteristic CurveSensitivityCritical Appraisal Tools and Checklists

Spotted an issue on this page? Report or suggest a fix →

ScholarGate — Multicenter Diagnostic Accuracy Study (Multicenter Diagnostic Accuracy Study). Retrieved 2026-07-21 from https://scholargate.app/en/epidemiology/multicenter-diagnostic-accuracy-study · Dataset: https://doi.org/10.5281/zenodo.20539026
Quick facts
Originator
STARD Group (Bossuyt, Reitsma et al.)
Year
2003 (STARD statement first published; updated 2015)
Type
Observational diagnostic study design
DataType
Patient-level test results, reference standard outcomes, site-level covariates
Subfamily
Clinical / epidemiology
Related methods
Diagnostic Accuracy Study DesignMeta-analytic Diagnostic Accuracy StudyMulticenter cohort studyMulticenter Randomized Clinical TrialScreening Test Evaluation
ScholarGate

A content-first reference library for research methods — what each one is, how it works, and where it comes from.

Open data (CC-BY)

Explore

  • Library
  • Search the library…
  • Browse by field
  • Fields
  • Journey
  • Compare
  • Which method?

Reference

  • Subjects
  • Atlas
  • Glossary
  • Methodology
  • Philosophy

Your tools

  • Bookshelf
  • Desk
  • Chat

Company

  • About
  • Pricing
  • Contact
  • Suggest a method

Entries are compiled from published sources for reference. Verifying the accuracy and suitability of any information for your own use remains your responsibility.

© 2026 ScholarGate · A research-method reference library
  • Privacy
  • Cookies
  • Terms
  • Delete account