Besag-York-Mollie Model
Also known as: BYM Model, Convolution Prior Model, CAR Convolution Model, BYM2 Reparameterization
The Besag-York-Mollie (BYM) model is the workhorse hierarchical Bayesian model for small-area disease mapping. Proposed by Julian Besag, Jeremy York, and Annie Mollie (1991), it models area-level disease counts with a Poisson likelihood whose log relative risk is the sum of two random effects: a spatially structured component, given an intrinsic conditional autoregressive (ICAR) prior that borrows strength from neighboring areas, and an unstructured component capturing area-specific heterogeneity that is not spatially patterned. This convolution of structured and unstructured effects lets the model smooth noisy small-area rates toward local and global means while distinguishing genuine spatial trend from independent overdispersion. Because the original parameterization makes the two variance components hard to interpret and depends on the graph, Riebler, Sorbye, Simpson, and Rue (2016) introduced the scaled BYM2 reparameterization, which mixes a scaled spatial effect and an unstructured effect through a single interpretable mixing parameter and a total-variance parameter, improving prior specification and identifiability.
Key highlights
- Stabilizes unreliable small-area rates by borrowing strength from neighbors, yielding smoother, more trustworthy risk maps.
- Separates spatially structured variation from unstructured heterogeneity, aiding interpretation of whether risk is location-linked.
- Sits in a flexible hierarchical Bayesian framework that propagates uncertainty and accommodates covariates and extensions (space-time, multivariate).
- The BYM2 reparameterization gives interpretable, near-orthogonal hyperparameters with priors that transfer across different maps.
Intuition
This section is available to Pro members. Upgrade to Pro
How it works
This section is available to Pro members. Upgrade to Pro
When to use it
Use the BYM model when you have counts of a disease or event aggregated to small areas, want to estimate and map the underlying relative risk, and need to stabilize noisy rates by borrowing strength across space. It is the default choice when areas have a meaningful adjacency structure, when raw small-area rates are unstable because of sparse counts, and when you want to distinguish spatially structured variation (suggesting location-linked exposures) from unstructured heterogeneity. Prefer the BYM2 reparameterization when you care about interpreting how much variation is spatial, want priors that transfer across regions, or face identifiability concerns. The model is less appropriate when the scientific aim is testing for discrete clusters with a p-value (use a scan statistic instead), when areas have no sensible neighborhood graph, or when the count process is far from Poisson in ways the random effects cannot absorb. It assumes the chosen adjacency structure reasonably reflects spatial dependence.
Strengths & limitations
- Stabilizes unreliable small-area rates by borrowing strength from neighbors, yielding smoother, more trustworthy risk maps.
- Separates spatially structured variation from unstructured heterogeneity, aiding interpretation of whether risk is location-linked.
- Sits in a flexible hierarchical Bayesian framework that propagates uncertainty and accommodates covariates and extensions (space-time, multivariate).
- The BYM2 reparameterization gives interpretable, near-orthogonal hyperparameters with priors that transfer across different maps.
- In its original form the two variance components are weakly identified and not on comparable scales, complicating interpretation and prior choice.
- Results depend on the adjacency definition, and the binary first-order neighbor graph is a crude model of true spatial dependence.
- Smoothing can oversmooth genuine sharp boundaries or attenuate true localized excesses (ecological smoothing bias).
- Fitting requires Bayesian computation (MCMC or INLA) and care with the improper ICAR prior, constraints, and convergence.
Common pitfalls
This section is available to Pro members. Upgrade to Pro
Applications
This section is available to Pro members. Upgrade to Pro
Frequently asked
Why does the BYM model need both a structured and an unstructured random effect?
Because geographic variation in disease risk has two distinct flavors. Some of it is spatially smooth, where nearby areas resemble each other because they share location-linked exposures, and the structured ICAR effect captures this by shrinking each area toward its neighbors. Some of it is idiosyncratic, where a single area deviates for reasons unrelated to location, such as data artifacts or purely local factors, and the unstructured effect captures this independent heterogeneity. With only the structured effect, the model would force every anomaly to spread across neighbors; with only the unstructured effect, it would ignore real spatial smoothing. Including both lets the data decide, through the two variance components, how much of the variation is spatial.
What problem does the BYM2 reparameterization solve?
The original BYM model has two variance parameters that are only weakly identified, because at each area the data mainly inform the sum of the two random effects, not each separately, and the two are measured on incomparable scales that depend on the adjacency graph. This makes it hard to choose priors or to interpret how much variation is spatial. Riebler and colleagues' BYM2 rewrites the model with a single total-variance parameter and one mixing parameter between zero and one giving the spatial fraction, and it scales the ICAR component so the prior means the same thing on any map. The hyperparameters become near-orthogonal and interpretable, and penalized-complexity priors can be placed on them sensibly.
How is the BYM model usually fitted in practice?
Historically it was fitted by MCMC in WinBUGS or OpenBUGS, but the modern standard is integrated nested Laplace approximation via the R-INLA package, which fits the BYM and BYM2 models quickly and accurately for latent Gaussian models like this one. Stan (and brms) and Nimble are also used, particularly for custom extensions. In all cases the analyst supplies the observed and expected counts and an adjacency graph, specifies the ICAR plus heterogeneity (or BYM2) structure, places priors on the hyperparameters (penalized-complexity priors are recommended for BYM2), and then summarizes the posterior relative risks with maps, credible intervals, and exceedance probabilities rather than point estimates alone.
Sources
- 1.Besag, J., York, J., & Mollie, A. (1991). Bayesian image restoration, with two applications in spatial statistics. Annals of the Institute of Statistical Mathematics, 43(1), 1-20.
- 2.Riebler, A., Sorbye, S. H., Simpson, D., & Rue, H. (2016). An intuitive Bayesian spatial model for disease mapping that accounts for scaling. Statistical Methods in Medical Research, 25(4), 1145-1165.
You have read it. What now?
Cite this page
ScholarGate. (2026, June 23). Besag-York-Mollie Model. ScholarGate. https://scholargate.app/spatial-epidemiology/besag-york-mollie-model