Skip to contentScholarGate
LibraryBookshelfDeskReview StudioAssistant
Sign in
On this page
IntuitionHow it worksWhen to use itStrengths & limitationsCommon pitfallsApplicationsFrequently asked🔒 Read the full methodSourcesRelated methods
Cite this pageSpotted an issue on this page? Report or suggest a fix →
Home›Statistics›Equivalence Test (TOST — Two One-Sided Tests)
Hypothesis test

Equivalence Test (TOST — Two One-Sided Tests)

Two One-Sided Tests Procedure for Equivalence · Also known as: TOST, two one-sided tests, bioequivalence test, Eşdeğerlik Testi (TOST — Two One-Sided Tests)

The equivalence test using the Two One-Sided Tests (TOST) procedure is a parametric hypothesis test designed to demonstrate that the difference between two group means falls within a pre-specified equivalence region ±Δ. Introduced by Schuirmann (1987) in the context of pharmaceutical bioequivalence, TOST reverses the logic of classical null-hypothesis testing: instead of trying to detect a difference, it provides positive evidence of similarity.

ScholarGate
  1. Hypothesis test
  2. v1
  3. 2 Sources
  4. PUBLISHED
Cite this page →
Tools & resources
Download slides
Learn & explore

Read the full method

Members only

Sign in with a free account to read this section.

Sign in

Method map

The neighbourhood of related methods — select a node to explore.

Equivalence Test (TOST)
Independent t-testMann-Whitney U testOne-way ANOVAPaired t-testWelch t-testBioequivalence AnalysisBland-Altman AnalysisEquivalence / Non-Inferi…

When to use it

Use TOST whenever the scientific claim requires demonstrating similarity rather than difference: bioequivalence of drug formulations, validation that a new measurement instrument agrees closely with an established one, or confirmation that an intervention does not meaningfully alter a baseline measure. The key assumptions are approximate normality of the outcome in each group, at least roughly equal variances (or use the Welch form), independent observations, and — critically — that the equivalence margin Δ is defined on substantive grounds before data collection, not chosen post hoc to fit the results. A minimum of about 20 observations per group is recommended; TOST is often underpowered with small samples because demonstrating absence of a meaningful difference requires more data than demonstrating its presence.

Strengths & limitations

Strengths
  • Provides positive statistical evidence of equivalence rather than merely a failure to reject a null difference.
  • Directly linked to the 90% confidence interval, giving a transparent and interpretable summary.
  • Well-established regulatory precedent in pharmaceutical and regulatory sciences (FDA, EMA guidelines).
  • Adaptable to non-normal data via a Wilcoxon-based TOST variant.
Limitations
  • The equivalence margin Δ must be specified before data collection; post-hoc choice of Δ invalidates the inference.
  • Typically requires larger samples than a standard difference test to achieve adequate power.
  • Does not replace measures of practical importance: a statistically confirmed equivalence within a very wide Δ may still be practically uninformative.
  • The pooled-variance form assumes equal variances; unequal variances require the Welch adjustment.

Frequently asked

How is TOST different from simply not rejecting a standard t-test?

A non-significant standard t-test only shows that you lack strong evidence of a difference — it does not confirm equivalence. TOST actively tests whether the difference is small enough to be practically negligible by checking that a confidence interval for the mean difference falls entirely within the pre-specified equivalence bounds. The two conclusions are logically different.

How should I choose the equivalence margin Δ?

Δ must be chosen on substantive grounds before data collection, not by looking at the data. In pharmacology it is typically set at ±20% of the reference mean. In other fields, you should define the smallest difference that would be practically meaningful in your context. Lakens (2017) recommends expressing Δ in standardised units (e.g., Cohen's d = 0.5) when no natural metric criterion exists.

Why is the 90% confidence interval used rather than the 95% interval?

Because TOST runs two one-sided tests each at α = 0.05, the combined procedure corresponds mathematically to checking whether a 90% confidence interval for the mean difference falls inside [−Δ, +Δ]. Using a 95% interval would be too conservative and is the wrong correspondence. This is a common source of confusion when interpreting TOST results.

Can I use TOST with non-normal data?

Yes. For non-normal distributions, a Wilcoxon-based TOST variant is available that uses the same two-one-sided logic but replaces the t statistic with a non-parametric rank-based counterpart. The equivalence margin must then be stated on the original measurement scale or in terms of shift.

Sources

  1. Schuirmann, D.J. (1987). A Comparison of the Two One-Sided Tests Procedure and the Power Approach for Assessing the Equivalence of Average Bioavailability. Journal of Pharmacokinetics and Biopharmaceutics, 15(6), 657–680. DOI: 10.1007/BF01068419 ↗
  2. Lakens, D. (2017). Equivalence Tests: A Practical Primer for t Tests, Correlations, and Meta-Analyses. Social Psychological and Personality Science, 8(4), 355–362. DOI: 10.1177/1948550617697177 ↗

How to cite this page

ScholarGate. (2026, June 1). Two One-Sided Tests Procedure for Equivalence. ScholarGate. https://scholargate.app/en/statistics/equivalence-test-tost

Related methods

Independent t-testMann-Whitney U testOne-way ANOVAPaired t-testWelch t-test

Which method?

Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.

  • Independent t-testStatistics↔ compare
  • Mann-Whitney U testStatistics↔ compare
  • One-way ANOVAStatistics↔ compare
  • Paired t-testStatistics↔ compare
  • Welch t-testStatistics↔ compare
Compare side by side →

Referenced by

Bioequivalence AnalysisBland-Altman AnalysisEquivalence / Non-Inferiority Trial

Similar methods

Equivalence / Non-Inferiority TrialBioequivalence AnalysisIndependent t-testIndependent samples t-testPaired t-testOne-sample t-testRobust one-sample t-testPower Analysis for t-test

Related reference concepts

Bioequivalence Studies and AssessmentHypothesis Testing FrameworkStatistical Power and Sample SizeStatistical Estimation and InferenceBioavailability and BioequivalenceStatistical Hypothesis Testing

Spotted an issue on this page? Report or suggest a fix →

ScholarGate — Equivalence Test (TOST) (Two One-Sided Tests Procedure for Equivalence). Retrieved 2026-07-21 from https://scholargate.app/en/statistics/equivalence-test-tost · Dataset: https://doi.org/10.5281/zenodo.20539026
Quick facts
Originator
Donald J. Schuirmann
Year
1987
Family
Hypothesis test
Type
Parametric equivalence test
Groups
2
Outcome
continuous
Parametric
Yes
Distribution
Student t (two one-sided)
EquivalenceBounds
user-specified ±Δ
MinSample
20
Related methods
Independent t-testMann-Whitney U testOne-way ANOVAPaired t-testWelch t-test
ScholarGate

A content-first reference library for research methods — what each one is, how it works, and where it comes from.

Open data (CC-BY)

Explore

  • Library
  • Search the library…
  • Browse by field
  • Fields
  • Journey
  • Compare
  • Which method?

Reference

  • Subjects
  • Atlas
  • Glossary
  • Methodology
  • Philosophy

Your tools

  • Bookshelf
  • Desk
  • Chat

Company

  • About
  • Pricing
  • Contact
  • Suggest a method

Entries are compiled from published sources for reference. Verifying the accuracy and suitability of any information for your own use remains your responsibility.

© 2026 ScholarGate · A research-method reference library
  • Privacy
  • Cookies
  • Terms
  • Delete account