Process / pipelineHuman GeographyCluster detection and disease surveillancePipeline

Scan Statistic Cluster Detection

Also known as: Kulldorff Scan Statistic, Spatial Scan Statistic, SaTScan Cluster Detection

OriginatorMartin KulldorffYear1997Sources1Related methods4

The spatial scan statistic, introduced by Martin Kulldorff in 1997, is a method for detecting and testing the significance of spatial clusters of events such as disease cases. It moves windows of many sizes and positions across the study region, treating each window as a candidate cluster, and scores it by a likelihood ratio comparing the rate of events inside the window to the rate outside. The window with the highest score is the most likely cluster, and its significance is assessed by Monte Carlo simulation, giving a principled answer to the recurring question of whether an apparent hotspot is real or chance.

Key highlights

  • Provides a formal significance test for clusters while correctly adjusting for multiple testing across all windows.
  • Locates clusters without prior knowledge of their size or position, scanning many scales simultaneously.
  • Supports several probability models (Poisson, Bernoulli, ordinal, space-time) for different data types.
  • Backed by free, widely used software (SaTScan) and an extensive validation literature in epidemiology.

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Use the spatial scan statistic when you have point or area-level counts of events together with a population or expected baseline, and you need to know whether and where statistically significant clusters occur — the classic settings being disease surveillance, outbreak detection, and crime or accident hotspot analysis. It is well suited to prospective surveillance (repeatedly scanning incoming data for emerging clusters) and to space-time extensions. It is less appropriate when you want a smooth continuous risk surface rather than discrete clusters, when the expected counts cannot be credibly specified, or when the events are not plausibly generated by an inhomogeneous Poisson-like process.

Strengths & limitations

Strengths
  • Provides a formal significance test for clusters while correctly adjusting for multiple testing across all windows.
  • Locates clusters without prior knowledge of their size or position, scanning many scales simultaneously.
  • Supports several probability models (Poisson, Bernoulli, ordinal, space-time) for different data types.
  • Backed by free, widely used software (SaTScan) and an extensive validation literature in epidemiology.
Limitations
  • Circular windows are biased toward compact clusters and can miss or distort elongated or irregular ones.
  • Monte Carlo inference is computationally intensive, especially for large datasets and space-time scans.
  • Results depend on the chosen maximum window size and the specified expected (baseline) counts.
  • Detects the location of clusters but does not by itself explain their causes or rule out confounding.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

How does the scan statistic correct for multiple comparisons?

Although it evaluates a huge number of overlapping windows, it bases inference on the single maximum likelihood ratio rather than on each window separately. The null distribution of that maximum is obtained by Monte Carlo simulation — repeatedly reshuffling the cases under the no-cluster hypothesis and recomputing the maximum each time. Comparing the observed maximum to this simulated distribution yields a p-value that already accounts for having scanned everywhere, so no further multiple-testing adjustment is needed.

What shapes of cluster can the scan statistic find?

The classic version uses circular windows of varying radius, which is most powerful for roughly compact clusters but can poorly capture elongated or irregular ones. Extensions use elliptical windows and 'flexibly shaped' scans built from sets of adjacent regions to detect non-circular clusters. The trade-off is that more flexible window shapes increase computation and, if unconstrained, can over-fit, so the choice depends on the plausible geometry of the phenomenon.

What is the difference between retrospective and prospective scanning?

Retrospective scanning analyses a fixed historical dataset once to find where clusters occurred. Prospective scanning is applied repeatedly to accumulating data — for example daily surveillance — to detect clusters as they emerge, and it adjusts for the repeated analyses over time so that the running surveillance maintains correct overall significance. The space-time scan statistic underlies most prospective public-health early-warning systems.

Sources

  1. 1.
    Kulldorff, M. (1997). A spatial scan statistic. Communications in Statistics – Theory and Methods, 26(6), 1481–1496.

You have read it. What now?

Cite this page

ScholarGate. (2026, June 22). Scan Statistic Cluster Detection. ScholarGate. https://scholargate.app/human-geography/scan-statistic-cluster-detection