Skip to contentScholarGate
LibraryBookshelfDeskReview StudioAssistant
Sign in
On this page
IntuitionHow it worksWhen to use itStrengths & limitationsCommon pitfallsApplicationsFrequently asked🔒 Read the full methodSourcesRelated methods
Cite this pageSpotted an issue on this page? Report or suggest a fix →
Home›Experimental design›Crossover A/B Test — Within-Subject A/B Testing Design
Process / pipelineExperimental design

Crossover A/B Test — Within-Subject A/B Testing Design

Crossover A/B Testing Design · Also known as: within-subject A/B test, crossover split test, repeated-measures A/B test, AB crossover experiment

A crossover A/B test is an experimental design in which the same participants or units are exposed to both treatment A and treatment B in sequence, with each serving as their own control. By eliminating between-subject variability, the design achieves higher statistical power than a standard parallel A/B test at the same sample size, but it requires careful handling of carryover effects and time-period confounds.

ScholarGate
  1. Process / pipeline
  2. v1
  3. 2 Sources
  4. PUBLISHED
Cite this page →
Tools & resources
Download slides
Learn & explore

Read the full method

Members only

Sign in with a free account to read this section.

Sign in

Method map

The neighbourhood of related methods — select a node to explore.

Crossover A/B Test
AB DesignAdaptive A/B testBlocked A/B TestCrossover Randomized Con…Factorial A/B TestMulti-arm experiment

When to use it

Use a crossover A/B test when participants are scarce or recruitment is expensive, when individual differences are large relative to the treatment effect (making within-subject comparison powerful), and when carryover effects are either negligible or can be managed with a washout. It is well-suited to stable phenomena — chronic condition management, pricing sensitivity, UI preference — where the user's baseline is unlikely to shift dramatically between periods. Avoid it when the first treatment permanently changes the participant (learning effects, habit formation), when the outcome of interest cannot be measured twice (irreversible events, one-time purchases), when the washout needed would be impractically long, or when the total experiment duration becomes so long that attrition or external temporal trends become serious threats to validity.

Strengths & limitations

Strengths
  • Each participant serves as their own control, eliminating between-person variance and substantially increasing statistical power compared to a parallel design of the same size.
  • Requires fewer participants to achieve equivalent power — critical in low-traffic or hard-to-recruit populations.
  • Produces within-subject estimates of treatment effects, which are often more precise and interpretable than between-group estimates.
  • Period and sequence effects can be explicitly estimated and adjusted for in a properly analyzed crossover design.
  • Participants experience both conditions, which can improve ecological validity when the phenomenon depends on individual preference or tolerance.
Limitations
  • Carryover effects — the residual influence of the first treatment on the second period — are a fundamental threat; they cannot always be fully eliminated even with a washout period.
  • The total experiment duration is longer than a parallel A/B test because participants must complete two exposure periods, increasing the risk of attrition and external trend confounds.
  • The statistical model is more complex than a simple two-sample comparison, requiring adjustment for period, sequence, and potential interaction effects.
  • Cannot be used when the outcome is a one-time or irreversible event, or when treatment effects are expected to persist indefinitely.
  • Period effects — changes in the environment or participant behavior between Period 1 and Period 2 — can bias results if they are not evenly distributed across sequences.

Frequently asked

How long should the washout period be?

The washout should last at least as long as it takes the treatment effect to dissipate to a negligible level. In pharmacology this is often five half-lives of the drug. In behavioral or online experiments it may be one to two weeks, or zero if the intervention leaves no measurable residual influence. There is no universal rule; base the decision on prior evidence or a pilot study.

What happens if carryover is significant?

If the formal carryover test (treatment-by-period interaction) is significant, Period 2 data are contaminated and cannot be used for an unbiased treatment comparison. The recommended approach is to use only Period 1 data, which effectively reduces the crossover to a parallel design and sacrifices the efficiency gain. This is why detecting carryover risk before designing the study is so important.

Is a crossover A/B test the same as a within-subject design in psychology?

Conceptually yes — both compare the same participants under multiple conditions. The crossover label is most common in clinical and online experimentation contexts, while 'within-subject design' or 'repeated-measures design' is the standard terminology in psychology. The key difference in practice is that the crossover design explicitly randomizes the order of conditions and models the sequence and period effects, whereas some within-subject designs expose all participants to conditions in the same fixed order.

Can I run a crossover A/B test with more than two conditions?

Yes. Higher-order crossover designs (e.g., three-period, three-treatment) exist and are used in practice. Williams (1949) provided balanced sequences for any number of treatments. However, complexity and participant burden grow with each additional condition, and managing carryover between more than two treatments is considerably harder. For online experiments, two-condition crossovers are by far the most common.

How do I determine the required sample size?

Sample size for a crossover design uses the within-subject standard deviation (not the total standard deviation), which is typically substantially smaller than the between-subject figure. Standard formulas for paired comparisons apply, adjusted for the desired power and significance level. Because crossover designs are more efficient, the required N is often 30–60% of what a parallel A/B test would need for the same effect size.

Sources

  1. Jones, B., & Kenward, M. G. (2014). Design and Analysis of Cross-Over Trials (3rd ed.). Chapman and Hall/CRC. ISBN: 9781439861424
  2. Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing. Cambridge University Press. ISBN: 9781108724227

How to cite this page

ScholarGate. (2026, June 3). Crossover A/B Testing Design. ScholarGate. https://scholargate.app/en/experimental-design/crossover-ab-test

Related methods

AB DesignAdaptive A/B testBlocked A/B TestCrossover Randomized Controlled TrialFactorial A/B TestMulti-arm experiment

Which method?

Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.

  • AB DesignExperimental design↔ compare
  • Adaptive A/B testExperimental design↔ compare
  • Blocked A/B TestExperimental design↔ compare
  • Crossover Randomized Controlled TrialExperimental design↔ compare
  • Factorial A/B TestExperimental design↔ compare
  • Multi-arm experimentExperimental design↔ compare
Compare side by side →

Similar methods

Crossover DesignCrossover Randomized Controlled TrialCrossover Control Group Experimental DesignCrossover Pretest-Posttest Experimental DesignCrossover Field ExperimentCrossover Laboratory ExperimentCrossover Factorial ExperimentCrossover multi-arm experiment

Related reference concepts

Bioequivalence Studies and AssessmentPermutation TestsRandomization and BlockingUser and Online EvaluationCross-Validation and ResamplingStatistical Power and Sample Size

Spotted an issue on this page? Report or suggest a fix →

ScholarGate — Crossover A/B Test (Crossover A/B Testing Design). Retrieved 2026-07-21 from https://scholargate.app/en/experimental-design/crossover-ab-test · Dataset: https://doi.org/10.5281/zenodo.20539026
Quick facts
Originator
Crossover design: E. J. Williams (1949); A/B testing framework: Ronald Fisher (experimental roots); modern online application widely attributed to Google and Microsoft experimentation teams
Year
1949 (crossover design); 2000s (online A/B application)
Type
Within-subject controlled experiment
DataType
Continuous, binary, or count outcome data measured on the same units across periods
Subfamily
Experimental design
Related methods
AB DesignAdaptive A/B testBlocked A/B TestCrossover Randomized Controlled TrialFactorial A/B TestMulti-arm experiment
ScholarGate

A content-first reference library for research methods — what each one is, how it works, and where it comes from.

Open data (CC-BY)

Explore

  • Library
  • Search the library…
  • Browse by field
  • Fields
  • Journey
  • Compare
  • Which method?

Reference

  • Subjects
  • Atlas
  • Glossary
  • Methodology
  • Philosophy

Your tools

  • Bookshelf
  • Desk
  • Chat

Company

  • About
  • Pricing
  • Contact
  • Suggest a method

Entries are compiled from published sources for reference. Verifying the accuracy and suitability of any information for your own use remains your responsibility.

© 2026 ScholarGate · A research-method reference library
  • Privacy
  • Cookies
  • Terms
  • Delete account