Cluster Randomized Multiple Baseline Design
Also known as: CR-MBD, cluster-randomized MBD, group-randomized multiple baseline, multilevel multiple baseline design
The cluster randomized multiple baseline design combines cluster-level random assignment with the logic of the multiple baseline design. Intact groups — such as classrooms, schools, or clinics — are randomly assigned to receive an intervention at staggered time points. This preserves the within-unit repeated-measure logic of the multiple baseline while adding the causal warrant of random assignment at the cluster level.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Use this design when: (1) the intervention is delivered at the group or organizational level and individual randomization is not feasible; (2) you need to demonstrate a replicable, causally interpretable effect across multiple clusters; (3) time-series data can be collected frequently enough to observe baseline stability and post-intervention change. It is especially suited to school-based, community, or health-system interventions. Do NOT use it when clusters cannot be identified as meaningfully intact units, when fewer than three tiers are achievable, when stable baselines are implausible (rapid naturally occurring change), or when the intervention cannot be ethically withheld from some clusters during their extended baseline phase.
Strengths & limitations
- Combines the replication logic of the multiple baseline with the causal warrant of random assignment, strengthening internal validity over non-randomized single-subject designs.
- Respects natural clustering — avoids contamination that would arise from randomizing individuals within the same group.
- Requires far fewer clusters than a parallel-group cluster RCT to demonstrate an effect, making it feasible in low-resource or novel intervention contexts.
- Provides time-series data that reveal the trajectory of change, not just a pre-post snapshot.
- The staggered baseline serves as a within-study replication: if the effect appears three or more times only when intervention begins, alternative explanations are weakened.
- Requires a minimum of three tiers; with only two, the replication evidence is insufficient to rule out coincidental timing.
- Clusters in later tiers face longer baselines, raising ethical concerns if the intervention is expected to be beneficial.
- Intraclass correlation within clusters inflates standard errors; analysis must account for the multilevel structure or conclusions will be anti-conservative.
- The design assumes no intervention diffusion across clusters — if control clusters are exposed to the treatment (contamination), baseline stability is compromised.
- Visual analysis, while primary, is susceptible to analyst bias; supplementary quantitative tests are advisable.
Frequently asked
How is this different from a standard cluster RCT?
A standard cluster RCT assigns clusters to treatment or control and compares them at a single post-test. A cluster randomized multiple baseline design assigns clusters to staggered intervention start points and follows each cluster with dense repeated measures over time. This produces a within-study replication of the effect at each intervention onset, which strengthens causal inference and reveals the trajectory of change — information a single post-test cannot provide.
How many clusters do I need?
A minimum of three clusters (one per tier) is required to provide the replication needed for causal inference. In practice, having two or more clusters per tier — giving six or more clusters total — substantially increases power and generalizability. Power analysis should account for the intraclass correlation within clusters.
What if my baseline is not stable before the intervention begins?
An unstable or trending baseline is a serious design threat. If the outcome is already changing before intervention, any post-intervention change cannot be unambiguously attributed to the treatment. Investigators should delay introducing the intervention in a given tier until the baseline trend is clearly stable and flat, even if this extends the study. If stability is consistently unattainable, a different design may be more appropriate.
Can I use standard regression to analyze the data?
Standard OLS regression ignores the clustered data structure (observations within clusters are not independent) and the repeated-measures structure (time points within a cluster are autocorrelated). Multilevel or hierarchical models that account for cluster-level random effects, or randomization tests designed for single-subject data, are required for valid inference.
Is it ethical to withhold treatment from later tiers for a long baseline?
This is the primary ethical tension in the design. Researchers should pre-specify the maximum baseline length and build in provisions for early intervention if a participant or cluster reaches a harm threshold. When the intervention is not yet proven effective — as is often the case in a pilot study — withholding it during baseline is generally defensible. For established effective treatments, the design may not be ethically justifiable.
Sources
- Murray, D. M. (1998). Design and Analysis of Group-Randomized Trials. Oxford University Press. ISBN: 978-0195120424
- Kowalski, J. T., & Shadish, W. R. (2015). Incorporating cluster randomization into multiple baseline designs. Journal of School Psychology, 53(6), 435-449. link ↗
How to cite this page
ScholarGate. (2026, June 3). Cluster Randomized Multiple Baseline Design. ScholarGate. https://scholargate.app/en/experimental-design/cluster-randomized-multiple-baseline-design
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- ABAB designExperimental design↔ compare
- Cluster Randomized Controlled TrialExperimental design↔ compare
- Multiple Baseline DesignExperimental design↔ compare
- Randomized Controlled TrialExperimental design↔ compare
- Single-Subject Experimental DesignExperimental design↔ compare