Multivariate Spatio-Temporal Shared Component Models
- Multivariate spatio-temporal shared component models are Bayesian hierarchical frameworks that decompose risk into shared latent factors and outcome-specific variations across space and time.
- They extend traditional disease mapping by integrating explicit spatial and temporal smoothing with flexible scaling to address incomplete or sparse data.
- Applications in cancer mapping and incidence–mortality studies highlight their practical role in disentangling common etiologic drivers and improving public health forecasts.
Searching arXiv for the cited papers to ground the article in the current literature. Multivariate spatio-temporal shared component models are Bayesian hierarchical models for jointly analyzing multiple related outcomes over space and time by decomposing risk into latent components that are shared across outcomes, together with outcome-specific terms. In disease mapping, the central use case is that several diseases, or incidence and mortality for the same disease, may reflect common latent risk factors while still exhibiting distinct spatial, temporal, or spatio-temporal deviations. The resulting framework combines multivariate borrowing of information, explicit spatial smoothing, temporal smoothing, and, in some formulations, outcome-specific scaling of shared latent structure. Recent work has extended the classical shared spatial component formulation to include shared temporal effects, shared spatio-temporal interactions, flexible time-varying modulation, and forecasting under incomplete registry coverage (Mahaki et al., 2017, Retegui et al., 2024, Retegui et al., 29 Jul 2025).
1. Conceptual foundations and scope
The shared component model is a joint disease-mapping construction in which several outcomes are linked through one or more latent components representing common underlying causes. Rather than specifying only a generic multivariate covariance structure, the model represents disease risk as a sum of latent components, each associated with a specific subset of diseases. This makes the latent terms interpretable as surrogates for latent risk factors when exposure data are unavailable. The epidemiologic appeal is therefore not merely multivariate smoothing, but the separation of shared versus disease-specific variation (Mahaki et al., 2017).
In the spatio-temporal setting, the extension from purely spatial shared component models adds temporal trends and, in some formulations, space-time interaction terms. This allows assessment of whether latent risk-factor patterns persist over time, whether they evolve temporally, and whether particular area-year combinations deviate from additive spatial and temporal structure. The Iran cancer model is explicit on this point: each component is shared by different subsets of diseases, spatial and temporal trends are considered for each component, and the relative weight of these trends for each component for each relevant disease can be estimated (Mahaki et al., 2017).
A second major use case is bivariate modeling of incidence and mortality for the same cancer. In that setting, shared components encode common disease burden and common area-level risk structure, while outcome-specific terms absorb features that differ between incidence and mortality because of diagnosis, treatment, survival, or coding practices. This logic underlies both the rare-cancer models with flexible shared interactions and the England registry-completion models that use mortality to improve inference on incomplete or delayed incidence series (Retegui et al., 2024, Retegui et al., 29 Jul 2025).
A common misconception is that any multivariate spatio-temporal model with dependence across outcomes is automatically a classical shared component model. The literature distinguishes between explicit shared-component decompositions and broader latent Gaussian or latent-process frameworks. The multivariate spatio-temporal mixed effects model (MSTM) contains a common latent state vector shared across outcomes, but it does not explicitly separate one named shared component plus outcome-specific structured residual fields in the classical disease-mapping sense (Bradley et al., 2014). Similarly, tinyVAST can encode latent factors, simultaneous effects, lagged effects, and recursive dependencies, but it is introduced as a graph-based multivariate spatio-temporal latent Gaussian GLMM framework rather than as a canonical shared-component package (Thorson et al., 2024).
2. Hierarchical formulations
A standard multivariate spatio-temporal shared component formulation begins with a Poisson likelihood for observed counts and a log-risk decomposition into intercepts and latent structure. In the Iran cancer model, for province , year , and cancer ,
With four latent shared components, the additive spatio-temporal specification is
where is the spatial effect of shared component , is the temporal effect of shared component , and and 0 are disease-specific spatial and temporal scaling parameters. Model variants add disease-specific heterogeneity 1, a common space-time interaction 2, or both (Mahaki et al., 2017).
In the bivariate incidence–mortality literature, the formulation is typically written directly in terms of log-rates. The baseline multivariate rare-cancer model uses a shared spatial effect, outcome-specific temporal effects, and outcome-specific spatio-temporal interactions:
3
4
Its shared-interaction extension replaces 5 and 6 with one common latent interaction field 7, first with a single scaling parameter 8, and then with time-varying scaling:
9
0
The novelty is that the same latent spatio-temporal interaction is used for both outcomes, but its contribution is modulated over time through 1 (Retegui et al., 2024).
The England registry-completion models use the same bivariate Poisson structure for incidence and mortality, with population-at-risk offsets rather than expected counts. Three variants are considered: a model with only a shared spatial component, a model with shared spatial and shared temporal components, and a model with a shared spatial component plus a flexible shared spatio-temporal interaction. The most distinctive formulation is
2
3
Here the omission of an incidence-specific temporal effect is deliberate: the shared spatio-temporal interaction is intended to capture incidence’s evolving temporal behavior jointly with mortality when incidence is incomplete (Retegui et al., 29 Jul 2025).
3. Shared structure, scaling, and identifiability
The defining feature of these models is that dependence between outcomes is induced hierarchically through common latent components in the linear predictors. In the Iran formulation, a disease may share one or more latent components with other diseases, but the strength of association is disease-specific through 4 and 5. A component contributes only to diseases for which it is considered relevant; otherwise its weight is effectively zero. This allows a priori epidemiologic mappings such as smoking shared by esophagus, stomach, bladder, and lung cancers, or overweight/obesity shared by esophagus, colorectal, prostate, and breast cancers (Mahaki et al., 2017).
In bivariate incidence–mortality models, sharedness is typically represented through reciprocal loading parameterizations. A shared spatial field 6 enters incidence as 7 and mortality as 8. When a shared spatio-temporal interaction is included, the same pattern is repeated with 9 or 0. This reciprocal structure is a standard identifiability device in two-outcome shared component models, because it prevents duplication of a common latent field by two unconstrained loadings (Retegui et al., 2024, Retegui et al., 29 Jul 2025).
The flexible-shared-interaction literature makes the identifiability argument explicit. For two outcomes with reciprocal loadings, Held’s condition
1
is automatically satisfied. For the flexible interaction model, there are 2 scaling parameters in total, 3 for incidence and 4 for mortality, so
5
This makes the time-varying reciprocal construction identifiable by design (Retegui et al., 2024).
The same literature also distinguishes shared from outcome-specific residual structure. In the Iran model, disease-specific heterogeneity is represented by 6, and the paper assigns a zero-mean multivariate normal distribution with covariance matrix 7 to allow for correlation between the relevant diseases in each space-time unit (Mahaki et al., 2017). In the rare-cancer and registry-completion models, outcome-specific departures can enter through unstructured spatial effects such as 8, 9, or 0, or through outcome-specific temporal effects 1 and, in some models, outcome-specific spatio-temporal interactions (Retegui et al., 2024, Retegui et al., 29 Jul 2025).
This suggests a useful taxonomy. Classical multivariate spatio-temporal shared component models are explicit additive decompositions of shared and specific latent terms. Broader latent-process models such as the MSTM instead represent dependence through a common dynamic low-rank state 2 entering all outcomes,
3
with 4 common across all 5 processes. The resulting structure is closely analogous to shared latent components, but it is less explicit about shared versus outcome-specific structured decomposition (Bradley et al., 2014).
4. Spatial, temporal, and interaction priors
The latent spatial, temporal, and spatio-temporal structures are typically represented through Gaussian Markov random field priors. For shared spatial effects in the bivariate incidence–mortality literature, the neighborhood matrix is “defined by Besag,” with adjacency based on shared borders. The prior is
6
which is an intrinsic CAR/ICAR-type prior expressed through the graph Laplacian precision matrix 7 (Retegui et al., 29 Jul 2025). The rare-cancer paper describes the same construction as an intrinsic CAR (iCAR/Besag) prior (Retegui et al., 2024).
Temporal effects are commonly modeled as first-order random walks. In the registry-completion setting,
8
where 9 is “determined by the temporal structure matrix of a first order random walk” (Retegui et al., 29 Jul 2025). The Iran cancer model uses the same RW1 logic for each shared temporal component 0, describing it as the one-dimensional analog of a CAR prior (Mahaki et al., 2017).
Spatio-temporal interaction priors often follow Knorr-Held’s four classical types. The interaction vector has prior
1
The four structures are:
- Type I: 2
- Type II: 3
- Type III: 4
- Type IV: 5
The registry-completion paper reports that Type II consistently performed best for all models, meaning the interaction is temporally structured but spatially unstructured at the interaction level (Retegui et al., 29 Jul 2025). The rare-cancer study evaluates all four types in simulation and in the Great Britain application (Retegui et al., 2024).
Because these latent effects overlap with intercepts and with one another, centering constraints are necessary. The literature explicitly imposes sum-to-zero constraints on spatial and temporal effects, and additional constraints on spatio-temporal interactions depending on the interaction type. For example, Type I uses
6
while Type IV uses
7
These remove confounding with main spatial and temporal effects (Retegui et al., 2024, Retegui et al., 29 Jul 2025).
Hyperpriors are correspondingly simple. The Iran paper uses 8 for spatial and temporal precisions, 9 for residual covariance, and log-normal-type priors induced by
0
for positive scaling weights (Mahaki et al., 2017). The bivariate incidence–mortality papers use 1, 2, and 3, together with vague priors for latent standard deviations or precisions (Retegui et al., 2024, Retegui et al., 29 Jul 2025).
5. Estimation, computation, and adjacent scalable frameworks
Estimation strategies reflect the evolution of the field. The Iran cancer model was fitted by MCMC in WinBUGS with 50,000 simulated draws, the first 20,000 discarded as burn-in, and every 10th retained. Convergence was assessed using Gelman–Rubin diagnostics and graphical inspection of sample paths. Model comparison used DIC, with reported values 4, 5, 6, and 7 for Models A–D, respectively, so Model B was preferred (Mahaki et al., 2017).
More recent bivariate models are fitted with integrated nested Laplace approximation. The registry-completion paper uses R-INLA because the models are latent Gaussian with structured random effects and missing responses, and missing incidence observations are entered as NA so posterior marginals are returned for the corresponding linear predictors 8 (Retegui et al., 29 Jul 2025). The rare-cancer paper also uses INLA, but its flexible shared-interaction model is not directly available in R-INLA; the authors therefore implement it through rgeneric and the copy feature to reuse the same latent interaction field in both incidence and mortality equations (Retegui et al., 2024).
Predictive model comparison is central in the more recent literature. The England registry-completion study evaluates models using absolute relative bias, Dawid–Sebastiani score, and interval score, and bases model selection on out-of-sample predictive validation rather than DIC or WAIC (Retegui et al., 29 Jul 2025). The rare-cancer paper uses DIC, WAIC, and logarithmic score for fit, and MARB, MRRMSE, interval score, credible interval length, and coverage for predictive assessment (Retegui et al., 2024).
Two adjacent frameworks are methodologically important because they generalize shared-latent-structure ideas beyond the classical disease-mapping decomposition. The MSTM is a fully Bayesian multivariate spatio-temporal mixed effects model with Gibbs sampling, Kalman filtering and smoothing for the latent states, and reduced-rank Moran’s I basis functions. It is designed for areal data with multivariate-spatio-temporal dependencies, accommodates missingness and differing spatial support, and was demonstrated on a dataset with 7,530,037 observations and 3,680 spatial fields (Bradley et al., 2014). tinyVAST is an R package for latent Gaussian GLMMs that uses formula syntax, arrow notation, and arrow-and-lag notation to encode simultaneous, lagged, and recursive dependencies. It encompasses vector autoregressive, empirical orthogonal functions, spatial factor analysis, and ARIMA models, and it uses TMB, Laplace approximation, and sparse precision matrices (Thorson et al., 2024). Neither framework is introduced primarily as a classical shared-component model, but both can express shared latent structures.
6. Applications, empirical findings, and methodological significance
The Iran application illustrates the classical multivariate disease-mapping use case. Seven cancers were analyzed across 30 provinces and 5 years using four hypothesized latent components: smoking; overweight/obesity; inadequate fruit and vegetable consumption; and low physical activity. The preferred model by DIC included residual heterogeneity but not the extra space-time interaction term. The paper reports that esophagus and stomach had highest risk in the northern part of Iran; lung had higher risk in the northwest; prostate and breast had high risk in Isfahan, Yazd, Fars, Tehran, Semnan, Mazandaran, and Razavi Khorasan; and temporal effects of all four shared components were nearly constant across the 5 years (Mahaki et al., 2017).
The Great Britain rare-cancer study shows how shared spatio-temporal interactions can improve estimation when counts are small. The data consist of 142 administrative health care districts, 9 biennial periods from 2002–2019, and bivariate incidence–mortality outcomes for pancreatic cancer and leukaemia. For pancreatic cancer, the selected model was a shared spatial component plus a shared interaction with one scaling parameter and Type I interaction. For leukaemia, the selected model included a mortality-specific unstructured spatial effect, a shared interaction, 7 scaling parameters, and Type IV interaction. The paper reports that models with shared spatio-temporal interactions outperform those with only outcome-specific interactions, and that for leukaemia all posterior medians of the 9 were greater than 1, so the shared interaction affects incidence more strongly than mortality (Retegui et al., 2024).
The England lung-cancer registry-completion study shifts the shared-component framework from explanatory disease mapping toward prediction under incomplete observation. Male lung cancer incidence and mortality were modeled over 106 clinical commissioning groups from 2001–2019. Incidence availability in validation was deliberately staggered to mimic registry start dates: 70% available in 2001–2003, 75% in 2004–2006, 81% in 2007–2009, 88% in 2010–2012, 93% in 2013–2015, and 100% from 2016 onward. The core result is that the model with a shared spatial component and flexible shared spatio-temporal interaction performed best overall, especially for incomplete incidence series and short-term forecasting. In 2001–2003, MARB was 0, 1, and 2 for Models 1–3, respectively; for three-years-ahead national forecasting, MARB and credible interval length were 3 and 4 for Model 1, 5 and 6 for Model 2, and 7 and 8 for Model 3 (Retegui et al., 29 Jul 2025).
Taken together, these studies show two distinct but compatible interpretations of multivariate spatio-temporal shared component models. In the classical disease-mapping tradition, shared components are latent surrogates for common etiologic drivers, with disease-specific spatial and temporal weights. In the more recent bivariate incidence–mortality literature, the shared component acts as a predictive bridge between related outcomes, especially when one outcome is delayed, incomplete, or sparse. A plausible implication is that the field is moving from purely explanatory shared spatial mapping toward shared latent spatio-temporal constructions designed simultaneously for smoothing, harmonization, completion, and short-horizon forecasting (Mahaki et al., 2017, Retegui et al., 2024, Retegui et al., 29 Jul 2025).