Process / pipelineExperimental designExperimental designPipeline

Cluster Randomized Solomon Four-Group Design

Also known as: CR-S4GD, cluster-randomized four-group design, group-randomized Solomon design, Solomon four-group cluster trial

OriginatorRichard L. Solomon (four-group logic, 1949); cluster randomization methods developed by Murray and colleagues in the 1990sYear1949 (Solomon design); cluster extension formalized in 1990sSources2Related methods6

The cluster randomized Solomon four-group design combines cluster randomization — assigning intact groups such as schools, clinics, or communities to conditions — with the Solomon four-group structure that isolates the effect of pretesting. Four clusters (or sets of clusters) are created: two receive the treatment and two serve as controls, with only one treatment cluster and one control cluster receiving a pretest, while the others go straight to the posttest. This structure simultaneously controls for pretest sensitization and the logistical constraint that individual randomization is infeasible.

Key highlights

  • Simultaneously controls for the testing effect (pretest sensitization) and the inability to randomize individuals, addressing two threats to internal validity at once.
  • Produces an estimate of the testing effect itself, providing methodological transparency and additional scientific information.
  • Cluster randomization protects against contamination when the treatment is delivered at the group level.
  • The 2×2 factorial structure permits estimation of treatment × pretest interaction, flagging whether the treatment is only effective for pretested participants.
  • Well-suited to real-world settings such as schools, clinics, and community programs where disrupting existing groupings is not feasible.

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Use this design when individual randomization is logistically or ethically impossible (e.g., whole-class instruction, community-level interventions, clinic protocols), AND you suspect that administering a pretest will itself alter participants' behavior or outcomes — a threat called pretest sensitization or testing effect. It is especially valuable in education, public health, and organizational research where clusters are natural groupings. Do not use it when you have fewer than four clusters per arm, as the design requires adequate cluster-level statistical power; with very few clusters, variance estimation is unreliable. It is also unnecessary if pretesting is not reactive (e.g., objective biomarker measures that participants cannot prepare for) — a simpler two-arm cluster RCT suffices in that case.

Strengths & limitations

Strengths
  • Simultaneously controls for the testing effect (pretest sensitization) and the inability to randomize individuals, addressing two threats to internal validity at once.
  • Produces an estimate of the testing effect itself, providing methodological transparency and additional scientific information.
  • Cluster randomization protects against contamination when the treatment is delivered at the group level.
  • The 2×2 factorial structure permits estimation of treatment × pretest interaction, flagging whether the treatment is only effective for pretested participants.
  • Well-suited to real-world settings such as schools, clinics, and community programs where disrupting existing groupings is not feasible.
Limitations
  • Requires substantially more clusters than a standard two-arm cluster RCT; each of the four arms needs adequate cluster-level power, making the design resource-intensive.
  • The intraclass correlation (ICC) inflates variance within arms, requiring larger within-cluster sample sizes or more clusters to achieve equivalent power compared to individual randomization.
  • Analysis is more complex than a simple pretest-posttest design; incorrect use of individual-level standard errors without accounting for clustering produces anti-conservative inference.
  • Practical coordination of four separate arms — some pretested, some not — in field settings increases administrative burden and the risk of protocol deviation.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

How is this design different from a standard cluster RCT?

A standard cluster RCT typically has two arms — treatment and control — and usually includes a pretest for both. The cluster randomized Solomon four-group design adds two additional arms that skip the pretest, allowing the researcher to estimate and statistically remove the sensitizing effect of pretesting itself. This matters when the pre-assessment could teach participants, raise awareness, or otherwise change behavior independently of the treatment.

How many clusters do I need per arm?

A minimum of four to six clusters per arm is generally recommended as a practical floor, but formal power analysis should drive the decision. You need to know (or estimate) the ICC, mean cluster size, and expected effect size. The design effect (DEFF = 1 + (m-1) × ICC) inflates the required sample size compared to an individual-level design. With ICCs typical in educational settings (0.05–0.20) and class sizes of 25–30, design effects of 2–5 are common, roughly doubling to quintupling the number of clusters needed.

Can I use this design when I have very few available clusters?

It is not advisable. With fewer than four clusters per arm, cluster-level variance estimates are highly unstable, confidence intervals are very wide, and the study will almost certainly be underpowered. If you have very few clusters, consider a within-cluster crossover design or switch to a quasi-experimental approach instead.

What statistical model should I use for the analysis?

The preferred approach is a mixed-effects linear (or generalized linear) model with treatment condition and pretest exposure as fixed factors, and cluster as a random intercept. Alternatively, you can compute cluster-level posttest means and run a 2×2 ANOVA on those means, which automatically respects the unit of randomization. Generalized estimating equations (GEE) are another option. Individual-level OLS without a cluster random effect is incorrect and will overstate significance.

Is the cluster randomized Solomon four-group design used in clinical trials?

It is less common in clinical trials than in educational and public health research, partly because individual randomization is more often feasible in clinical settings and biomarker outcomes are typically not reactive to baseline measurement. However, it has appeared in cluster-randomized trials of behavioral and psychosocial interventions in healthcare, where awareness raised by a baseline questionnaire could plausibly change patient behavior before the intervention begins.

Sources

  1. 1.
    Solomon, R. L. (1949). An extension of control group design. Psychological Bulletin, 46(2), 137–150.
  2. 2.
    Murray, D. M. (1998). Design and Analysis of Group-Randomized Trials. Oxford University Press.
    ISBN 978-0195100877

You have read it. What now?

Cite this page

ScholarGate. (2026, June 3). Cluster Randomized Solomon Four-Group Design. ScholarGate. https://scholargate.app/experimental-design/cluster-randomized-solomon-four-group-design

Cluster Randomized Solomon Four-Group Design | ScholarGate