Process / pipelineSpatial EpidemiologySpatial cluster detection / scan statisticsPipeline

Spatial Scan Statistic

Also known as: Kulldorff Scan Statistic, SaTScan Cluster Detection, Circular Scan Statistic, Spatial Likelihood-Ratio Scan

OriginatorMartin Kulldorff (with Neville Nagarwalla)Year1997Sources2Related methods7

The spatial scan statistic is a likelihood-ratio method for detecting localized clusters of disease without pre-specifying where they are. Introduced by Martin Kulldorff and Neville Nagarwalla (1995) and generalized by Kulldorff (1997), it slides a circular window of varying size and position across the study region, and for each candidate window compares the observed-to-expected case ratio inside the window against outside it using a likelihood ratio under a Poisson or Bernoulli model. The window that maximizes the likelihood ratio is the most likely cluster, and its statistical significance is obtained by Monte Carlo simulation under the null of no clustering, which correctly accounts for the enormous multiplicity of windows examined. Implemented in the widely used SaTScan software, the method has become the standard tool for screening surveillance data for spatial and space-time disease clusters.

Key highlights

  • Detects clusters of unknown location and size without pre-specifying where to look, avoiding the bias of testing a hand-picked area.
  • Monte Carlo inference adjusts automatically for the massive multiplicity of overlapping windows, controlling false positives.
  • Provides interpretable output per cluster: location, size, observed vs expected counts, relative risk, and a p-value.
  • Flexible across data types and models (Poisson, Bernoulli, space-time, multinomial) and supported by mature, free software (SaTScan).

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Use the spatial scan statistic when you have geolocated disease data and want to test whether localized clusters exist and pinpoint where, without specifying candidate locations in advance. It is well suited to routine surveillance, outbreak investigation, and screening of cancer or chronic-disease registries, and it extends naturally to space-time scanning for emerging clusters. The method assumes you can compute expected counts from a population at risk (or use a control group) and that the cluster shape is reasonably approximated by the chosen window geometry, typically circular, though elliptical and irregular variants exist. It is less appropriate when clusters are expected to be highly irregular or linear and only circular windows are used, when the population denominator is poorly measured, or when the scientific question is smooth risk estimation across the map rather than detection of discrete clusters, for which Bayesian disease mapping is a better fit.

Strengths & limitations

Strengths
  • Detects clusters of unknown location and size without pre-specifying where to look, avoiding the bias of testing a hand-picked area.
  • Monte Carlo inference adjusts automatically for the massive multiplicity of overlapping windows, controlling false positives.
  • Provides interpretable output per cluster: location, size, observed vs expected counts, relative risk, and a p-value.
  • Flexible across data types and models (Poisson, Bernoulli, space-time, multinomial) and supported by mature, free software (SaTScan).
Limitations
  • The classic circular window has low power for irregularly shaped or elongated clusters and can over- or under-estimate their extent.
  • Results are sensitive to the maximum-window-size setting, which trades off detecting small versus large clusters.
  • It detects and tests clusters but does not produce a smooth, full-map estimate of relative risk.
  • Accuracy depends on correct expected counts; poor population denominators or unadjusted confounders produce spurious or missed clusters.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

How does the scan statistic avoid the multiple-testing problem from examining so many windows?

It uses Monte Carlo hypothesis testing rather than analytic p-values. After computing the maximum likelihood ratio over all candidate windows in the real data, it generates many datasets under the null of no clustering (randomly allocating cases in proportion to population) and recomputes the same maximized statistic for each. The observed statistic's rank among the simulated maxima yields the p-value. Because each simulated dataset's statistic is itself a maximum over all windows, the null distribution already embodies the full multiplicity of the search, so the test is correctly sized despite scanning an enormous number of overlapping circles.

What does the maximum cluster size setting do and how should I choose it?

The maximum cluster size caps how large a scanning window can grow, typically expressed as a fraction of the population at risk (50% is the conventional default). A smaller cap focuses power on tight, local clusters and can split a large diffuse cluster, while a larger cap can detect broad excesses but may report implausibly large 'clusters' that cover much of the map. Because results depend on this choice, good practice is to run the scan at several maximum sizes and report sensitivity, choosing a value motivated by the plausible spatial scale of the disease process under study rather than accepting the default uncritically.

When should I use a spatial scan statistic instead of disease mapping?

Use the scan statistic when the goal is detection and hypothesis testing: deciding whether discrete clusters exist and locating them with a p-value, as in surveillance and outbreak investigation. Use Bayesian disease mapping when the goal is estimation: producing a smooth, full-map picture of relative risk with stabilized small-area rates and uncertainty. The two are complementary, scanning answers 'is there a significant hot spot and where,' while mapping answers 'what is the estimated risk surface everywhere.' Many studies run both, using the scan to flag clusters and a smoothed map to characterize the underlying risk pattern.

Sources

  1. 1.
    Kulldorff, M. (1997). A spatial scan statistic. Communications in Statistics - Theory and Methods, 26(6), 1481-1496.
  2. 2.
    Kulldorff, M., & Nagarwalla, N. (1995). Spatial disease clusters: detection and inference. Statistics in Medicine, 14(8), 799-810.

You have read it. What now?

Cite this page

ScholarGate. (2026, June 23). Spatial Scan Statistic. ScholarGate. https://scholargate.app/spatial-epidemiology/spatial-scan-statistic