Skip to contentScholarGate
LibraryBookshelfDeskReview StudioAssistant
Sign in
On this page
IntuitionHow it worksWhen to use itStrengths & limitationsCommon pitfallsApplicationsFrequently asked🔒 Read the full methodSourcesRelated methods
Cite this pageSpotted an issue on this page? Report or suggest a fix →
Home›Econometrics›Diebold-Mariano Test of Equal Predictive Accuracy
Hypothesis testForecast evaluation

Diebold-Mariano Test of Equal Predictive Accuracy

Also known as: DM Test, Test of Equal Forecast Accuracy, Diebold-Mariano Forecast Comparison Test, Tahmin Doğruluğu Eşitliği Testi

The Diebold-Mariano (DM) test, introduced by Diebold and Mariano in 1995, is a widely used non-parametric procedure for formally comparing the predictive accuracy of two competing forecasting models. It evaluates whether the difference in forecast errors between two models is statistically significant, without requiring nested models or specific distributional assumptions about the forecasts, making it broadly applicable across economics, finance, and time-series analysis.

ScholarGate
  1. Hypothesis test
  2. v1
  3. 1 Sources
  4. PUBLISHED
Cite this page →
Tools & resources
Download slides
Learn & explore

Read the full method

Members only

Sign in with a free account to read this section.

Sign in

Method map

The neighbourhood of related methods — select a node to explore.

Diebold-Mariano Test
Giacomini-White TestModel Confidence SetPesaran-Timmermann TestTime-Series Cross-Valida…

When to use it

Use the Diebold-Mariano test when comparing the out-of-sample predictive accuracy of two non-nested or nested forecasting models applied to the same target variable and evaluation period. It requires a sequence of genuine out-of-sample forecast errors, an appropriate loss function, and a sample large enough for the normal approximation to hold. The test is not appropriate for in-sample fit comparisons. For multi-horizon overlapping forecasts, Harvey, Leybourne, and Newbold (1997) proposed a small-sample correction. When comparing more than two models simultaneously, the Model Confidence Set is a natural alternative.

Strengths & limitations

Strengths
  • Distribution-free: requires no parametric assumptions about forecast errors or data-generating processes
  • Flexible loss function: accommodates squared error, absolute error, and asymmetric losses
  • Handles autocorrelation: long-run variance estimator corrects for serially correlated loss differentials from multi-step-ahead forecasts
  • Widely applicable: valid for nested and non-nested models alike
Limitations
  • Requires a sufficient number of out-of-sample observations for the normal approximation to be reliable
  • Standard form can be undersized in small samples; Harvey-Leybourne-Newbold correction is recommended for finite samples
  • Sensitive to the choice of loss function — different loss functions may yield conflicting conclusions
  • Does not account for parameter estimation uncertainty in the forecasting models, unlike the Giacomini-White test

Frequently asked

Can the DM test be used for nested model comparisons?

The original Diebold-Mariano test was not designed for nested models, where one model nests the other as a special case. In that setting, the loss differential degenerates under the null, making the standard normal approximation invalid. Clark and McCracken (2001, 2005) developed modified tests specifically for nested model comparisons, and those should be used instead when nesting is present.

How many out-of-sample observations are needed?

There is no universal minimum, but as a practical guideline at least 30 to 50 out-of-sample observations are typically needed for the asymptotic normal distribution to provide reliable inference. With fewer observations, the Harvey-Leybourne-Newbold small-sample correction — which replaces the standard normal critical values with t-distribution critical values — substantially improves size control and is strongly recommended.

Does the choice of loss function affect the result?

Yes, and this is intentional. The DM test evaluates predictive accuracy relative to a specific loss function specified by the researcher. Squared-error loss penalizes large errors more heavily, while absolute-error loss is more robust to outliers. If the two loss functions lead to conflicting conclusions about which model is better, this reveals that the ranking of models is loss-function-dependent and the researcher should choose the criterion that best reflects the actual forecasting objective.

Sources

  1. Diebold, F. X., & Mariano, R. S. (1995). Comparing predictive accuracy. Journal of Business & Economic Statistics, 13(3), 253–263. DOI: 10.1080/07350015.1995.10524599 ↗

How to cite this page

ScholarGate. (2026, June 2). Diebold-Mariano Test of Equal Predictive Accuracy. ScholarGate. https://scholargate.app/en/econometrics/diebold-mariano-test

Related methods

Giacomini-White TestModel Confidence SetPesaran-Timmermann Test

Which method?

Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.

  • Giacomini-White TestEconometrics↔ compare
  • Model Confidence SetEconometrics↔ compare
  • Pesaran-Timmermann TestEconometrics↔ compare
Compare side by side →

Referenced by

Giacomini-White TestModel Confidence SetPesaran-Timmermann TestTime-Series Cross-Validation

Similar methods

Giacomini-White TestPesaran-Timmermann TestModel Confidence SetTime-Series Cross-ValidationARCH-LM TestGranger CausalityTime series Bayesian model averagingLjung-Box Test

Related reference concepts

Econometric ModelingPredictive Information CriteriaBayesian Model Comparison and SelectionEconometricsCross-ValidationLikelihood-Ratio Tests

Spotted an issue on this page? Report or suggest a fix →

ScholarGate — Diebold-Mariano Test (Diebold-Mariano Test of Equal Predictive Accuracy). Retrieved 2026-07-21 from https://scholargate.app/en/econometrics/diebold-mariano-test · Dataset: https://doi.org/10.5281/zenodo.20539026
Quick facts
Originator
Francis Diebold & Roberto Mariano
Year
1995
Type
Non-parametric forecast comparison test
Subfamily
Forecast evaluation
Null Hypothesis
Equal predictive accuracy between two competing forecasts
Loss Function
Flexible; supports squared error, absolute error, and user-defined loss
Related methods
Giacomini-White TestModel Confidence SetPesaran-Timmermann Test
ScholarGate

A content-first reference library for research methods — what each one is, how it works, and where it comes from.

Open data (CC-BY)

Explore

  • Library
  • Search the library…
  • Browse by field
  • Fields
  • Journey
  • Compare
  • Which method?

Reference

  • Subjects
  • Atlas
  • Glossary
  • Methodology
  • Philosophy

Your tools

  • Bookshelf
  • Desk
  • Chat

Company

  • About
  • Pricing
  • Contact
  • Suggest a method

Entries are compiled from published sources for reference. Verifying the accuracy and suitability of any information for your own use remains your responsibility.

© 2026 ScholarGate · A research-method reference library
  • Privacy
  • Cookies
  • Terms
  • Delete account