Interrupted Time Series in Education Research
Interrupted Time Series Analysis in Education Research · Also known as: ITS in education, educational ITS, segmented regression in education, policy interrupted time series
Interrupted time series (ITS) analysis is a quasi-experimental design that estimates the causal effect of an education policy or intervention by examining whether an outcome trend changes abruptly at the point of implementation. Applied to education, it is used to evaluate reforms, curriculum changes, testing policies, and school interventions using routinely collected longitudinal data without a randomised control group.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Use ITS in education research when a policy, programme, or reform is introduced at a clearly defined time point and you have a sufficient number of observations both before and after — typically at least 8 pre- and 8 post-intervention time points, though 20+ per segment is preferable. The method fits administrative panel data such as annual state test results, monthly attendance records, or semester-level enrollment figures. Do not use ITS when the intervention date is ambiguous or gradual, when you have fewer than about 8 observations per segment (power is too low), when many other policy changes occurred simultaneously (confounding threats), or when the outcome series shows severe non-stationarity that cannot be modelled.
Strengths & limitations
- Provides a credible causal estimate from routinely collected administrative data without requiring a randomised control group.
- Separates two distinct causal pathways: an immediate level change and a longer-term trend change, giving richer information than a simple pre-post comparison.
- Particularly powerful for evaluating large-scale, population-level education reforms where randomisation is infeasible.
- Can be extended with a control series (comparative ITS) to further strengthen causal claims by ruling out concurrent confounds.
- Transparent and interpretable: results map directly onto the observed time-series graph, making communication to policy audiences straightforward.
- Requires a sufficient number of pre-intervention time points to estimate the baseline trend reliably; sparse data yields wide confidence intervals and low power.
- Cannot rule out concurrent events (history threat): if another policy changed at the same time, their effects are confounded.
- Assumes the intervention date is sharp and known precisely; phased or gradual rollouts complicate estimation.
- Autocorrelation in the series must be modelled explicitly, and misspecification of the error structure can bias standard errors and inference.
Frequently asked
How many time points do I need before and after the intervention?
A common rule of thumb is at least 8 pre- and 8 post-intervention observations, but 20 or more per segment is preferable for stable slope estimates and adequate power. With fewer than 8 points in either segment, the trend estimates are unreliable and confidence intervals are very wide.
What is the difference between a level change and a slope change?
A level change (β₂) is an immediate jump or drop in the outcome at the moment the intervention begins — like scores suddenly rising after a new programme starts. A slope change (β₃) is a gradual acceleration or deceleration of the trend after the intervention — like scores growing faster or slower than they were before. Both can occur together, and ITS estimates them separately.
Why does autocorrelation matter, and how do I handle it?
Repeated observations on the same population over time are usually positively correlated with recent past values, which means standard OLS standard errors are too small and p-values are misleadingly low. Use autocorrelation-corrected methods such as Newey-West heteroscedasticity-and-autocorrelation-consistent (HAC) standard errors, ARIMA error models, or generalised least squares to obtain valid inference.
Can I add a control group to strengthen the design?
Yes. Comparative ITS adds a parallel outcome series from a similar population that did not receive the intervention. If both the level and slope changes are near zero in the control series while they are significant in the treated series, this substantially strengthens the causal claim by ruling out concurrent confounds such as national trends or seasonal effects.
Is ITS accepted as rigorous evidence by education research standards?
Yes. The What Works Clearinghouse (WWC) and the ESSA evidence tiers classify well-designed ITS studies as Tier 2 (moderate) evidence — stronger than simple pre-post or correlational designs, though below randomised controlled trials. The key quality indicators are adequate segment length, autocorrelation correction, and absence of concurrent policy changes.
Sources
- Shadish, W. R., Cook, T. D., & Campbell, D. T. (2002). Experimental and Quasi-Experimental Designs for Generalized Causal Inference. Houghton Mifflin. ISBN: 978-0395615560
- Wagner, A. K., Soumerai, S. B., Zhang, F., & Ross-Degnan, D. (2002). Segmented regression analysis of interrupted time series studies in medication use research. Journal of Clinical Pharmacy and Therapeutics, 27(4), 299-309. DOI: 10.1046/j.1365-2710.2002.00430.x ↗
How to cite this page
ScholarGate. (2026, June 3). Interrupted Time Series Analysis in Education Research. ScholarGate. https://scholargate.app/en/causal-inference/interrupted-time-series-in-education-research
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Difference-in-DifferencesEconometrics↔ compare
- Panel Fixed EffectsEconometrics↔ compare
- Propensity Score MatchingResearch Statistics↔ compare
- Synthetic Control MethodCausal inference↔ compare