---
title: 'Spat-RPM: Spatial Domain Random Partition Model'
url: https://www.emergentmind.com/topics/spatial-domain-random-partition-model-spat-rpm
type: topic
---

# Spat-RPM: Spatial Domain Random Partition Model

Spatial Domain Random Partition Model (Spat-RPM) denotes, in the recent literature, a class of models that place a random partition on a spatial domain, spatial lattice, areal map, voxel grid, or spatial index set, and then define cluster-specific stochastic structure conditional on that partition. This suggests that Spat-RPM functions less as a single canonical specification than as an umbrella label for several related constructions, including spatial product partition models, similarity-weighted nonparametric mixtures, temporally evolving areal partitions, Potts-Gibbs image partitions, Voronoi-partition Cox processes, and graph-based local feature-selection models [1504.04489] [2308.01704] [2312.12396] [2507.17600].

## 1. Lineage and model family

The earliest formulation in this collection is the spatial product partition model of Page and Quintana, which extends product partition models so that the partitioning of locations into spatially dependent clusters is explicitly modeled through distance-based cohesion functions [1504.04489]. Subsequent work broadened the same organizing idea to functional spatial data through a similarity-based generalized Dirichlet process, to large spatio-temporal areal datasets through regime-specific random partitions and change-points, to multivariate disease mapping through short-boundary cohesions and MDAGAR spatial effects, and to sequences of temporally correlated spatial partitions driven by spanning trees [2308.01704] [2312.12396] [2401.08303] [2501.04601].

A parallel line applies the same partition-first logic to different data modalities. Teo Shu Xian and Wade use a Potts-Gibbs random partition for scalar-on-image regression, so that spatially neighboring voxels tend to share a common regression coefficient [2206.11051]. Nolau, Gonçalves and Gamerman model nonstationary spatial point processes by random Voronoi partitions with conditionally independent Gaussian processes across regions [2507.17600]. In spatial multimodal omics, the partition is imposed on a tessellated domain and coupled to cluster-wise feature selection [2605.30658]. A conceptually distinct usage appears in Durieu and Wang, where a random partition of $\mathbb{N}^2$ determines the covariance structure of a discrete random field whose scaling limit is a fractional Brownian sheet [1709.00934].

| Paper | Partition mechanism | Observation model |
|---|---|---|
| Page and Quintana [1504.04489] | Spatial PPM with cohesion $h(S_h;s_{S_h})$ | Cluster-specific likelihoods, often Gaussian regression |
| Wakayama, Sugasawa and Kobayashi [2308.01704] | SGDP with pairwise similarity reweighting | GP clustering of functional data |
| Change-point areal model [2312.12396] | Regime-specific areal random partitions | Harmonic regression with Leroux-CAR effects |
| Pavani and Quintana [2401.08303] | PPM with short-boundary cohesion $C_{\rm HB}$ | Multivariate Gaussian model with AR-seasonal terms |
| Teo Shu Xian and Wade [2206.11051] | Potts-Gibbs partition on image lattice | Scalar-on-image regression |
| Nolau, Gonçalves and Gamerman [2507.17600] | Random Voronoi tessellation | Cox process for point patterns |

Taken together, these formulations show that the common element is not a single likelihood or prior family, but the explicit use of a latent spatial partition as the main structural object. The partition may be over sites, areal units, meshes, voxels, graph vertices, or Voronoi cells; the likelihood may be Gaussian, Poisson, Poisson-inverse-Gaussian, Gaussian-process, or point-process based.

## 2. Partition priors and spatial regularization

A central distinction among Spat-RPM formulations is how they encode spatial contiguity, locality, or compactness. In the spatial PPM of Page and Quintana, the prior is
$$
p(\pi)\propto \prod_{h=1}^{k(\pi)} h(S_h;s_{S_h}),
$$
where the cohesion $h(\cdot)$ is modified to depend on spatial locations [1504.04489]. Their four cohesion choices include a penalty by sum-of-distances to the centroid, a hard-threshold neighbor rule, a prior-predictive similarity, and a posterior-predictive “double-dipping” cohesion. In this formulation, the mass parameter $M$ retains the usual role that larger $M$ implies more clusters a priori [1504.04489].

For areal data, a different device is to penalize boundary length directly. In the change-point spatio-temporal areal model, the areal product-partition prior within regime $r$ is
$$
P(\rho_r\mid \kappa,\xi)
=
\mathcal K(\kappa,\xi,I)\;
\kappa^{K_r}
\prod_{j=1}^{K_r}
\Gamma(|C^r_j|)
\exp\Bigl\{-\xi\sum_{i\in C^r_j}\ell^j(i)\Bigr\},
$$
where $\xi\ge 0$ penalizes cluster boundary length and $\sum_{i\in C^r_j}\ell^j(i)$ is the perimeter of cluster $C^r_j$ [2312.12396]. Pavani and Quintana use the Hegarty-Blackwell short-boundary cohesion
$$
C_{\rm HB}(S_j)=\eta^{\ell(S_j)},
$$
with $0\le \eta\le 1$, so that small $\eta$ favors few large contiguous regions and $\eta\to 1$ makes many small or even disconnected pieces more likely [2401.08303].

A different strategy is to preserve the cluster-count behavior of a nonparametric prior while biasing assignments toward nearby or adjacent units. In the SGDP model, pairwise similarities $s_{ii'}$ modify only the existing-cluster weights, while the new-cluster probability remains exactly as in the underlying generalized Dirichlet process [2308.01704]. Proposition 1 states that, for fixed $(\alpha,\beta,\tau)$, the SGDP prior strictly favors existing clusters containing more neighbors of $i$ [2308.01704]. This cleanly separates adjacency preference from the mechanism controlling total cluster counts.

Potts-Gibbs image models use a Markov random field penalty on shared labels. The partition prior has the form
$$
\pi(\mathcal C)
=
\frac{1}{Z(\upsilon,\phi)}
\exp\Bigl\{\upsilon\sum_{j\sim k}\mathbb{I}(z_j=z_k)\Bigr\}
V_p(M)\prod_{m=1}^M W_m(\phi),
$$
so larger $\upsilon$ makes label configurations with large contiguous blocks more probable [2206.11051]. In temporally evolving areal models driven by spanning trees, each season-specific partition is obtained by pruning edges of a spanning tree with probability $\rho_s$, guaranteeing spatially constrained clustering by construction [2501.04601]. In Voronoi-partition Cox processes, the domain is partitioned by random generators $U_1,\dots,U_K$, and a repulsive prior on $U$ is used to avoid collapse and label-switching [2507.17600]. In spatial multimodal omics, the domain is tessellated into blocks, a spanning tree is drawn on the block graph, and cutting $k-1$ edges yields $k$ connected components [2605.30658].

This variety shows that “spatial regularization” in Spat-RPMs is not tied to a single mathematical device. It may be implemented through cohesion functions, perimeter penalties, similarity reweighting, Potts interactions, tree pruning, or random tessellations.

## 3. Observation models and latent structures

Conditioning on the partition, different Spat-RPMs specify markedly different local models. In the change-point areal model, for time $t$ in regime $r=r_t$ and areal unit $i$,
$$
Y_{it}
=
\bm x_t'\bm\beta_{ir}+u_{ir}+\epsilon_{it},
\qquad
\epsilon_{it}\sim N(0,\sigma^2_{\epsilon_r}),
$$
with $\bm x_t$ a $K$-dimensional harmonic regression design and $\bm u_r\sim N_I(\bm 0,\tau_r^2Q(\zeta_r,W)^{-1})$ under a Leroux-CAR precision [2312.12396]. Regimes such as weekday-day, weekday-night, weekend-day, and weekend-night are induced by a vector of change-points, and each regime has its own partition of the areal units [2312.12396].

In the SGDP formulation for functional spatial data, each location $i$ carries a function $y_i(x)$, and conditional on cluster assignment $z_i=j$,
$$
y_i(x)\sim GP(\theta_j(x),C_y),
\qquad
\theta_j(\cdot)\sim GP(m_\theta,C_\theta),
$$
so clustering is over mean functions rather than scalar responses [2308.01704]. The Tokyo application further indexes observations by period indicators for weekdays, holidays, and day-before-holiday, with separate SGDP priors and Gaussian-process atoms by period [2308.01704].

For multivariate disease data, Pavani and Quintana specify
$$
y_{itd}\sim \mathcal N\bigl(X_{itd}^\top\beta_d+Z_{itd}^\top\gamma_{c_i,d}^*+\phi_{id},\sigma_{c_i,d}^{2*}\bigr),
$$
where $Z_{itd}$ may include AR(3) and annual seasonal lags, $\gamma_{j,d}^*$ is cluster- and disease-specific, and the spatial random effects follow an MDAGAR prior that allows within-area and neighbor-area cross-disease effects [2401.08303]. The temporal dependence in this model is defined for areal clusters rather than for each area independently.

The spatio-temporal dengue model driven by spanning trees uses overdispersed counts. Conditional on offset $O_{it}$, mean $\lambda_{it}$, and latent heterogeneity $z_{is}$,
$$
y_{it}\sim \mathrm{Poisson}(O_{it}\lambda_{it}z_{is}),
\qquad
z_{is}\sim \mathrm{IG}(1,\psi_{is}),
$$
which yields a marginal Poisson-inverse-Gaussian distribution with mean $O_{it}\lambda_{it}$ and variance $O_{it}\lambda_{it}+(O_{it}\lambda_{it})^2\psi_{is}^{-1}$ [2501.04601]. The mean contains a cluster-season intercept and the dispersion has its own regression [2501.04601].

In scalar-on-image regression, the partition groups voxels with a common coefficient, reducing the image predictor from $p$ voxels to $M$ region-level coefficients [2206.11051]. In spatial multimodal omics, each cluster carries a sparse regression with a binary active-set indicator $\nu_j$, so the model performs local variable selection and local basis selection [2605.30658]. For point processes, the Voronoi-partition Cox process writes
$$
\lambda(s)=\sum_{k=1}^K \lambda_k^*F(\beta_k(s))\mathbf 1\{s\in R_k\},
$$
with conditionally independent Gaussian-process priors on the latent functions across regions [2507.17600].

Durieu and Wang’s random-field construction lies outside the inferential Bayesian clustering paradigm but remains partition-based at its core. A product partition of $\mathbb N^2$ is first built, then $\pm1$ spins are assigned to each cell by identical or alternating rules, and the resulting field has covariance determined by co-membership in the same partition block [1709.00934].

## 4. Posterior inference and computational strategies

Posterior computation in Spat-RPMs is generally MCMC-based, but the update mechanism depends strongly on the partition prior. The change-point areal model uses a Metropolis-within-Gibbs sampler that imputes missing $Y_{it}$, samples cluster means $\bm\beta^*_{jr}$, updates hyperparameters, draws spatial effects $\bm u_r$, updates variances, reallocates labels $s_i^r$, and samples change-points $\bar t_m$ [2312.12396]. Spatial contiguity enters directly through the cohesion factors in the allocation step and through the CAR precision $Q(\zeta,W)$ [2312.12396]. Because $Q(\zeta,W)$ is sparse, solving systems is $O(I)$ per regime via sparse Cholesky or conjugate-gradient, and the overall cost per sweep is $O\bigl(\sum_r[I\,K_r + I + T\,I]\bigr)$, linear in $I,T$ for fixed average $K_r$ [2312.12396].

The SGDP functional model uses Gibbs updates for cluster means and labels, inverse-gamma updates for scale parameters, and Metropolis-Hastings for length-scale parameters and for $(\alpha,\beta,\tau)$ on the transformed scale [2308.01704]. The spatial PPM framework of Page and Quintana uses a Gibbs-type MCMC in the spirit of Neal’s Algorithm 8, with optional split-merge Metropolis steps as in Jain and Neal to improve mixing [1504.04489]. Pavani and Quintana also use Neal’s Algorithm 8 for partition reassignment, with conjugate Gibbs updates for regression, temporal, variance, and MDAGAR parameters, and Metropolis-Hastings for each $\alpha_d$ [2401.08303].

Potts-Gibbs image models face an intractable partition function $Z(\upsilon,\phi)$ and therefore use a generalized Swendsen-Wang scheme with auxiliary bond variables [2206.11051]. The sampler constructs nested clusters and reassigns them to existing or auxiliary clusters with probabilities combining Gibbs-type weight ratios, marginal likelihoods, and a Potts correction factor [2206.11051]. This addresses the slow mixing of single-site Gibbs when $\upsilon$ is large.

Two recent models use trans-dimensional or exact samplers to address large partition spaces. The spatial multimodal omics model introduces an informed RJ-MCMC with Birth, Death, Change, and Hyper moves, exploiting conjugacy in $\theta_j$ so that each move only evaluates a marginal likelihood [2605.30658]. The nonstationary spatial point-process model develops an exact, discretization-free MCMC based on thinning augmentation, retrospective sampling, and Metropolis-Hastings updates for the random Voronoi generators [2507.17600]. Its Gaussian-process cost is reduced from a single $\mathcal O(n^3)$ calculation to $\sum_{k=1}^K n_k^3$ because each region is handled separately [2507.17600].

Temporally correlated spanning-tree partitions are updated by a partially collapsed Gibbs step over tree-cut indicators, followed by compatibility-preserving tree updates using Prim’s or Kruskal’s algorithm on random edge weights [2501.04601]. No reversible-jump steps are needed in that model [2501.04601].

## 5. Theoretical properties and interpretation

Several Spat-RPM papers establish explicit probabilistic properties of the partition prior or of the full posterior. In the SGDP model, the probability of creating a brand-new cluster at each step remains exactly as in the underlying GDP, and when $\alpha\beta>1$ the expected total number of clusters remains finite as $n\to\infty$, guarding against over-clustering [2308.01704]. This is a precise statement that adjacency bias and cluster-count control are separated in that construction.

The spatial multimodal omics model derives coupled hyperparameter conditions linking domain partition and local feature-selection priors, under which consistency theory and posterior contraction rates of both the domain partition and feature selection are established [2605.30658]. The paper states cluster-number consistency, partition contraction in mis-clustered area fraction, feature-selection consistency, and coefficient and predictive contraction at the common rate $O\!\bigl(K^{-1}\log^{r_0}n\bigr)$ under regularity assumptions [2605.30658]. This is one of the strongest asymptotic analyses among the models grouped under the Spat-RPM label.

Durieu and Wang provide a different kind of theory: functional central limit theorems for discrete random fields constructed from product partitions of $\mathbb N^2$ [1709.00934]. In the Karlin, Hammond-Sheffield, and mixed models, normalized rectangular partial sums converge to a fractional Brownian sheet with Hurst indices spanning $(0,1/2)$, $(1/2,1)$, and mixed regimes [1709.00934]. The covariance structure factorizes by coordinate because the underlying one-dimensional partitions are independent [1709.00934].

Across the Bayesian clustering papers, a recurring interpretive theme is that the partition prior determines which notion of “spatial smoothness” is being encoded. In short-boundary or perimeter-penalized models, the prior directly penalizes cut edges or boundary length [2312.12396] [2401.08303]. In Potts-Gibbs models, smoothness is local label agreement on a neighborhood graph [2206.11051]. In SGDP, the effect is not a boundary penalty but a reweighting toward existing clusters containing more similar or adjacent units [2308.01704]. This suggests that contiguous maps in different Spat-RPMs may represent distinct prior semantics even when they appear visually similar.

## 6. Applications and empirical behavior

Spat-RPMs have been developed largely in response to concrete applied problems. The change-point areal model is motivated by mobile phone usage in the Metropolitan area of Milan and is used to spatially cluster population patterns of mobile phone usage while allowing different hierarchical structures across time points [2312.12396]. The SGDP model is applied to hourly population flow in the central seven wards of Tokyo, discretized into $n=452$ square meshes of size $500\,\mathrm{m}^2$ over $T=30$ days [2308.01704]. In that application, posterior means of $\alpha_\ell\beta_\ell$ are all greater than $1$, posterior $\tau_\ell$ is far below $1$, weekday clusters recover interpretable “business-core,” “downtown,” and “residential” patterns, SGDP and GDP have almost identical cluster counts but SGDP yields markedly smoother and more contiguous cluster maps, and pre-holiday and holiday periods capture the “Friday night” surge in leisure districts [2308.01704].

In education data from Greater Santiago, the spatial PPM is compared against Spatial Stick-Breaking and standard GP regression. The simulation study reports adjusted Rand index, MSPE, and LPML/WAIC, and the real-data application finds that CPS gives notably better WAIC/LPML than SSB, Cohesion 1 fitted best, Cohesion 4 gave lowest MSPE, and the posterior mean number of clusters varied from approximately $9$ to approximately $35$ [1504.04489]. Predictive maps recovered the known east-west school-performance gradient [1504.04489].

For mosquito-borne disease data, two related strands appear. Pavani and Quintana model dengue and chikungunya in 145 microregions of Southeast Brazil from 2018 to 2022 and report that the proposal compares well to competing alternatives while clustering areas where temporal trends are similar and exploring temporal and spatial correlation between diseases [2401.08303]. Pavani, Loschi and Quintana analyze weekly dengue counts from 2018 to 2023 in the same region using temporally dependent sequences of spatial partitions; WAIC selects an AR(1) prior on $\{\pi_s\}$, the estimated number of clusters varies seasonally from $5$ to $18$, peaking in autumn, and regression results indicate that temperature raises both mean and dispersion, humidity raises mean but lowers dispersion, and HDI moderately raises mean [2501.04601].

The point-process formulation has been tested on synthetic discontinuity and hotspot examples and on two real datasets. In the discontinuity experiment, for any $K\ge 2$ the model substantially outperforms stationary and kernel methods, the smallest MSE occurs at the correct $K=2$, and oversizing to $K=6$ increases MSE by only approximately $10\%$, whereas stationary and kernel estimators show more than $50\%$ worse MSE [2507.17600]. In the hotspot experiment, fitted models with $K=10,15,20$ recover both background and hotspots almost perfectly [2507.17600]. In real data, the model is applied to *Beilschmiedia pendula* tree locations in Panama and to Mato Grosso fire-occurrence points, with posterior intensity maps agreeing with prior analyses and highlighting north-central hotspots consistent with known deforestation patterns [2507.17600].

In spatial multimodal omics, the model is applied to a breast-cancer Visium CytAssist dataset with $4\,169$ spots, $17\,957$ genes, and $35$ proteins, focusing on CD8A protein as response [2605.30658]. After screening and pre-selection, $234$ genes remain, WAIC tunes $(K,\alpha,\lambda)$, and the fitted model discovers $K=3$ clusters with distinct gene signatures associated with tumor-rich, stromal/tertiary-lymphoid, and mixed tumor-immune regions [2605.30658]. Those signatures correlate with tumor-marker proteins and scRNA-seq deconvoluted T/B-cell proportions [2605.30658].

The empirical record therefore supports a broad but coherent interpretation of Spat-RPMs: they are models in which uncertainty about the number, location, geometry, and temporal evolution of spatial clusters is represented explicitly, rather than being absorbed indirectly into a covariance function or a single global smoothness parameter.

Source: https://www.emergentmind.com/topics/spatial-domain-random-partition-model-spat-rpm