Equivalence Test (TOST — Two One-Sided Tests)
Two One-Sided Tests Procedure for Equivalence · Also known as: TOST, two one-sided tests, bioequivalence test, Eşdeğerlik Testi (TOST — Two One-Sided Tests)
The equivalence test using the Two One-Sided Tests (TOST) procedure is a parametric hypothesis test designed to demonstrate that the difference between two group means falls within a pre-specified equivalence region ±Δ. Introduced by Schuirmann (1987) in the context of pharmaceutical bioequivalence, TOST reverses the logic of classical null-hypothesis testing: instead of trying to detect a difference, it provides positive evidence of similarity.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Use TOST whenever the scientific claim requires demonstrating similarity rather than difference: bioequivalence of drug formulations, validation that a new measurement instrument agrees closely with an established one, or confirmation that an intervention does not meaningfully alter a baseline measure. The key assumptions are approximate normality of the outcome in each group, at least roughly equal variances (or use the Welch form), independent observations, and — critically — that the equivalence margin Δ is defined on substantive grounds before data collection, not chosen post hoc to fit the results. A minimum of about 20 observations per group is recommended; TOST is often underpowered with small samples because demonstrating absence of a meaningful difference requires more data than demonstrating its presence.
Strengths & limitations
- Provides positive statistical evidence of equivalence rather than merely a failure to reject a null difference.
- Directly linked to the 90% confidence interval, giving a transparent and interpretable summary.
- Well-established regulatory precedent in pharmaceutical and regulatory sciences (FDA, EMA guidelines).
- Adaptable to non-normal data via a Wilcoxon-based TOST variant.
- The equivalence margin Δ must be specified before data collection; post-hoc choice of Δ invalidates the inference.
- Typically requires larger samples than a standard difference test to achieve adequate power.
- Does not replace measures of practical importance: a statistically confirmed equivalence within a very wide Δ may still be practically uninformative.
- The pooled-variance form assumes equal variances; unequal variances require the Welch adjustment.
Frequently asked
How is TOST different from simply not rejecting a standard t-test?
A non-significant standard t-test only shows that you lack strong evidence of a difference — it does not confirm equivalence. TOST actively tests whether the difference is small enough to be practically negligible by checking that a confidence interval for the mean difference falls entirely within the pre-specified equivalence bounds. The two conclusions are logically different.
How should I choose the equivalence margin Δ?
Δ must be chosen on substantive grounds before data collection, not by looking at the data. In pharmacology it is typically set at ±20% of the reference mean. In other fields, you should define the smallest difference that would be practically meaningful in your context. Lakens (2017) recommends expressing Δ in standardised units (e.g., Cohen's d = 0.5) when no natural metric criterion exists.
Why is the 90% confidence interval used rather than the 95% interval?
Because TOST runs two one-sided tests each at α = 0.05, the combined procedure corresponds mathematically to checking whether a 90% confidence interval for the mean difference falls inside [−Δ, +Δ]. Using a 95% interval would be too conservative and is the wrong correspondence. This is a common source of confusion when interpreting TOST results.
Can I use TOST with non-normal data?
Yes. For non-normal distributions, a Wilcoxon-based TOST variant is available that uses the same two-one-sided logic but replaces the t statistic with a non-parametric rank-based counterpart. The equivalence margin must then be stated on the original measurement scale or in terms of shift.
Sources
- Schuirmann, D.J. (1987). A Comparison of the Two One-Sided Tests Procedure and the Power Approach for Assessing the Equivalence of Average Bioavailability. Journal of Pharmacokinetics and Biopharmaceutics, 15(6), 657–680. DOI: 10.1007/BF01068419 ↗
- Lakens, D. (2017). Equivalence Tests: A Practical Primer for t Tests, Correlations, and Meta-Analyses. Social Psychological and Personality Science, 8(4), 355–362. DOI: 10.1177/1948550617697177 ↗
How to cite this page
ScholarGate. (2026, June 1). Two One-Sided Tests Procedure for Equivalence. ScholarGate. https://scholargate.app/en/statistics/equivalence-test-tost
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Independent t-testStatistics↔ compare
- Mann-Whitney U testStatistics↔ compare
- One-way ANOVAStatistics↔ compare
- Paired t-testStatistics↔ compare
- Welch t-testStatistics↔ compare