Hypothesis testStatisticsTest

Equivalence Test (TOST — Two One-Sided Tests)

Also known as: TOST, two one-sided tests, bioequivalence test, Eşdeğerlik Testi (TOST — Two One-Sided Tests)

OriginatorDonald J. SchuirmannYear1987Sources2Related methods8

The equivalence test using the Two One-Sided Tests (TOST) procedure is a parametric hypothesis test designed to demonstrate that the difference between two group means falls within a pre-specified equivalence region ±Δ. Introduced by Schuirmann (1987) in the context of pharmaceutical bioequivalence, TOST reverses the logic of classical null-hypothesis testing: instead of trying to detect a difference, it provides positive evidence of similarity.

Key highlights

  • Provides positive statistical evidence of equivalence rather than merely a failure to reject a null difference.
  • Directly linked to the 90% confidence interval, giving a transparent and interpretable summary.
  • Well-established regulatory precedent in pharmaceutical and regulatory sciences (FDA, EMA guidelines).
  • Adaptable to non-normal data via a Wilcoxon-based TOST variant.

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Use TOST whenever the scientific claim requires demonstrating similarity rather than difference: bioequivalence of drug formulations, validation that a new measurement instrument agrees closely with an established one, or confirmation that an intervention does not meaningfully alter a baseline measure. The key assumptions are approximate normality of the outcome in each group, at least roughly equal variances (or use the Welch form), independent observations, and — critically — that the equivalence margin Δ is defined on substantive grounds before data collection, not chosen post hoc to fit the results. A minimum of about 20 observations per group is recommended; TOST is often underpowered with small samples because demonstrating absence of a meaningful difference requires more data than demonstrating its presence.

Strengths & limitations

Strengths
  • Provides positive statistical evidence of equivalence rather than merely a failure to reject a null difference.
  • Directly linked to the 90% confidence interval, giving a transparent and interpretable summary.
  • Well-established regulatory precedent in pharmaceutical and regulatory sciences (FDA, EMA guidelines).
  • Adaptable to non-normal data via a Wilcoxon-based TOST variant.
Limitations
  • The equivalence margin Δ must be specified before data collection; post-hoc choice of Δ invalidates the inference.
  • Typically requires larger samples than a standard difference test to achieve adequate power.
  • Does not replace measures of practical importance: a statistically confirmed equivalence within a very wide Δ may still be practically uninformative.
  • The pooled-variance form assumes equal variances; unequal variances require the Welch adjustment.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

How is TOST different from simply not rejecting a standard t-test?

A non-significant standard t-test only shows that you lack strong evidence of a difference — it does not confirm equivalence. TOST actively tests whether the difference is small enough to be practically negligible by checking that a confidence interval for the mean difference falls entirely within the pre-specified equivalence bounds. The two conclusions are logically different.

How should I choose the equivalence margin Δ?

Δ must be chosen on substantive grounds before data collection, not by looking at the data. In pharmacology it is typically set at ±20% of the reference mean. In other fields, you should define the smallest difference that would be practically meaningful in your context. Lakens (2017) recommends expressing Δ in standardised units (e.g., Cohen's d = 0.5) when no natural metric criterion exists.

Why is the 90% confidence interval used rather than the 95% interval?

Because TOST runs two one-sided tests each at α = 0.05, the combined procedure corresponds mathematically to checking whether a 90% confidence interval for the mean difference falls inside [−Δ, +Δ]. Using a 95% interval would be too conservative and is the wrong correspondence. This is a common source of confusion when interpreting TOST results.

Can I use TOST with non-normal data?

Yes. For non-normal distributions, a Wilcoxon-based TOST variant is available that uses the same two-one-sided logic but replaces the t statistic with a non-parametric rank-based counterpart. The equivalence margin must then be stated on the original measurement scale or in terms of shift.

Sources

  1. 1.
    Schuirmann, D.J. (1987). A Comparison of the Two One-Sided Tests Procedure and the Power Approach for Assessing the Equivalence of Average Bioavailability. Journal of Pharmacokinetics and Biopharmaceutics, 15(6), 657–680.
  2. 2.
    Lakens, D. (2017). Equivalence Tests: A Practical Primer for t Tests, Correlations, and Meta-Analyses. Social Psychological and Personality Science, 8(4), 355–362.

You have read it. What now?

Cite this page

ScholarGate. (2026, June 1). Equivalence Test (TOST). ScholarGate. https://scholargate.app/statistics/equivalence-test-tost