Papers
Topics
Authors
Recent
Search
2000 character limit reached

Separable Spatio-Temporal Hawkes Process

Updated 8 July 2026
  • Separable spatio-temporal Hawkes processes are self-exciting point processes that split triggering functions into independent temporal and spatial components.
  • This factorization enables efficient likelihood evaluation, simulation, and estimation, making them ideal for modeling events like earthquakes, crime, and social media bursts.
  • The separable assumption simplifies computations while providing a baseline for comparing more flexible nonseparable and neural extensions in modern research.

Searching arXiv for papers on separable spatio-temporal Hawkes processes and closely related models. {"query":"separable spatio-temporal Hawkes process arXiv Hawkes space-time separable kernel", "max_results": 10} {"query":"spatio-temporal Hawkes process separable kernel arXiv", "max_results": 10} A separable spatio-temporal Hawkes process is a self-exciting point process on space-time whose conditional intensity is written as a background rate plus a sum of triggering contributions from past events, with the triggering kernel factorized into a purely temporal term and a purely spatial term. In the formulations used across multivariate, Bayesian, nonparametric, and simulation-oriented treatments, this factorization typically takes the form g(Δt,Δs)=gT(Δt)gS(Δs)g(\Delta t,\Delta \mathbf{s})=g_T(\Delta t)g_S(\Delta \mathbf{s}), so that temporal decay and spatial spread are modeled independently and then multiplied (Yuan et al., 2018, Zhou et al., 2022, Bernabeu et al., 18 Nov 2025). The separable specification is therefore both a modeling assumption and a computational device: it is widely used because it preserves the Hawkes branching interpretation while making likelihood evaluation, simulation, and estimation more adaptable, but recent work also treats it as a restrictive assumption that can be relaxed by nonseparable, graphon-based, additive-GP, or neural constructions (Jun et al., 2022, Baars et al., 2024, Chukwuemeka et al., 27 Feb 2026, Liu et al., 30 Mar 2026).

1. Formal definition in space and time

On a spatial domain DR2D \subset \mathbb{R}^2 and time interval [0,T)[0,T), the standard spatio-temporal Hawkes conditional intensity is written as

λ(s,tHts)=μ(s,t)+i:ti<tg(s,t,si,ti),\lambda(\mathbf{s},t\mid\mathcal{H}^{\mathbf{s}}_t)=\mu(\mathbf{s},t)+\sum_{i:t_i<t} g(\mathbf{s},t,\mathbf{s}_i,t_i),

where μ(s,t)\mu(\mathbf{s},t) is the background rate and g(s,t,si,ti)g(\mathbf{s},t,\mathbf{s}_i,t_i) is the triggering function (Jun et al., 2022). In multivariate form, for LL interacting processes,

λl(t,s)=ul(t,s)+m=1Li:tm,i<tαm,lgm,l(ttm,i,ssm,i),\lambda_l^*(t,\mathbf{s})=u_l(t,\mathbf{s})+\sum_{m=1}^L\sum_{i:t_{m,i}<t}\alpha_{m,l}\,g_{m,l}(t-t_{m,i},\mathbf{s}-\mathbf{s}_{m,i}),

where αm,l\alpha_{m,l} is the expected number of offspring in process ll triggered by a parent in process DR2D \subset \mathbb{R}^20 (Zhou et al., 2022).

This background-plus-triggering decomposition is the defining Hawkes structure. The background term models exogenous arrivals, while the summation over past events models endogenous self-excitation or cross-excitation. Several papers emphasize that this additive decomposition is distinct from separability of the triggering kernel itself: all Hawkes models split exogenous and endogenous intensity additively, but only some impose a product factorization across time and space (Liu et al., 30 Mar 2026).

A closely related multivariate reconstruction formulation writes, for node or process DR2D \subset \mathbb{R}^21,

DR2D \subset \mathbb{R}^22

with DR2D \subset \mathbb{R}^23 representing the triggering strength from one process to another (Yuan et al., 2018). In this form, the separable spatio-temporal Hawkes process also serves as a model of latent interaction structure.

2. Separable triggering kernels

The separable assumption states that the triggering kernel can be written as a product of a temporal kernel and a spatial kernel: DR2D \subset \mathbb{R}^24 This factorization is explicit in the main spatio-temporal model for aggregated-data Bayesian inference, where

DR2D \subset \mathbb{R}^25

(Zhou et al., 2022). A standard separable kernel in terrorism modeling is likewise given as an exponential temporal term multiplied by a Gaussian spatial term (Jun et al., 2022).

The factorization can be parametric, semiparametric, or nonparametric. In a nonparametric multivariate network-reconstruction model, the separability assumption is retained while the temporal factor DR2D \subset \mathbb{R}^26 and the isotropic spatial factor DR2D \subset \mathbb{R}^27 are estimated from data rather than fixed a priori (Yuan et al., 2018). In a randomized-kernel model, the triggering intensity is also explicitly separable, but the Gaussian spatial kernel is approximated by random Fourier features so that the likelihood can be written using scalable matrix operations (Ilhan et al., 2020).

Representative paper Temporal factor Spatial factor
(Zhou et al., 2022) Exponential DR2D \subset \mathbb{R}^28 Gaussian
(Ilhan et al., 2020) Exponential DR2D \subset \mathbb{R}^29 Gaussian via RFF approximation
(Bernabeu et al., 18 Nov 2025) Exponential or power law Gaussian or exponential
(Yuan et al., 2018) Exponential benchmark or learned [0,T)[0,T)0 Gaussian benchmark or learned isotropic [0,T)[0,T)1

In simulation-oriented work, separability is used explicitly because “in many applications, it is common to decompose the triggering function into separable components” and because separable triggerings are “more computationally adaptable” to the presented methods (Bernabeu et al., 18 Nov 2025). That computational role is a recurring reason for the persistence of separable models even where more flexible kernels are available.

3. Branching structure, multivariate interaction, and stability

Separable spatio-temporal Hawkes processes inherit the branching-process interpretation of classical Hawkes models. Latent variables [0,T)[0,T)2 and [0,T)[0,T)3 are introduced so that [0,T)[0,T)4 denotes the probability that event [0,T)[0,T)5 is a background event and [0,T)[0,T)6 the probability that event [0,T)[0,T)7 is triggered by event [0,T)[0,T)8 (Bernabeu et al., 18 Nov 2025). In a related EM-type formulation, the latent indicators [0,T)[0,T)9 and λ(s,tHts)=μ(s,t)+i:ti<tg(s,t,si,ti),\lambda(\mathbf{s},t\mid\mathcal{H}^{\mathbf{s}}_t)=\mu(\mathbf{s},t)+\sum_{i:t_i<t} g(\mathbf{s},t,\mathbf{s}_i,t_i),0 encode whether event λ(s,tHts)=μ(s,t)+i:ti<tg(s,t,si,ti),\lambda(\mathbf{s},t\mid\mathcal{H}^{\mathbf{s}}_t)=\mu(\mathbf{s},t)+\sum_{i:t_i<t} g(\mathbf{s},t,\mathbf{s}_i,t_i),1 triggers event λ(s,tHts)=μ(s,t)+i:ti<tg(s,t,si,ti),\lambda(\mathbf{s},t\mid\mathcal{H}^{\mathbf{s}}_t)=\mu(\mathbf{s},t)+\sum_{i:t_i<t} g(\mathbf{s},t,\mathbf{s}_i,t_i),2 or whether event λ(s,tHts)=μ(s,t)+i:ti<tg(s,t,si,ti),\lambda(\mathbf{s},t\mid\mathcal{H}^{\mathbf{s}}_t)=\mu(\mathbf{s},t)+\sum_{i:t_i<t} g(\mathbf{s},t,\mathbf{s}_i,t_i),3 is a background event, and the expected complete-data likelihood then separates immigrant and offspring contributions (Yuan et al., 2018).

For multivariate models, separability operates at the kernel level while cross-process interaction is carried by a triggering matrix. In the aggregated-data Bayesian formulation,

λ(s,tHts)=μ(s,t)+i:ti<tg(s,t,si,ti),\lambda(\mathbf{s},t\mid\mathcal{H}^{\mathbf{s}}_t)=\mu(\mathbf{s},t)+\sum_{i:t_i<t} g(\mathbf{s},t,\mathbf{s}_i,t_i),4

with

λ(s,tHts)=μ(s,t)+i:ti<tg(s,t,si,ti),\lambda(\mathbf{s},t\mid\mathcal{H}^{\mathbf{s}}_t)=\mu(\mathbf{s},t)+\sum_{i:t_i<t} g(\mathbf{s},t,\mathbf{s}_i,t_i),5

so self-excitation corresponds to λ(s,tHts)=μ(s,t)+i:ti<tg(s,t,si,ti),\lambda(\mathbf{s},t\mid\mathcal{H}^{\mathbf{s}}_t)=\mu(\mathbf{s},t)+\sum_{i:t_i<t} g(\mathbf{s},t,\mathbf{s}_i,t_i),6 and cross-excitation to λ(s,tHts)=μ(s,t)+i:ti<tg(s,t,si,ti),\lambda(\mathbf{s},t\mid\mathcal{H}^{\mathbf{s}}_t)=\mu(\mathbf{s},t)+\sum_{i:t_i<t} g(\mathbf{s},t,\mathbf{s}_i,t_i),7 (Zhou et al., 2022). In network reconstruction, the matrix λ(s,tHts)=μ(s,t)+i:ti<tg(s,t,si,ti),\lambda(\mathbf{s},t\mid\mathcal{H}^{\mathbf{s}}_t)=\mu(\mathbf{s},t)+\sum_{i:t_i<t} g(\mathbf{s},t,\mathbf{s}_i,t_i),8 is interpreted directly as a weighted adjacency matrix, with λ(s,tHts)=μ(s,t)+i:ti<tg(s,t,si,ti),\lambda(\mathbf{s},t\mid\mathcal{H}^{\mathbf{s}}_t)=\mu(\mathbf{s},t)+\sum_{i:t_i<t} g(\mathbf{s},t,\mathbf{s}_i,t_i),9 implying a latent directed edge μ(s,t)\mu(\mathbf{s},t)0 (Yuan et al., 2018).

Normalization of the triggering kernel leads to a simple stability condition in several treatments. When the spatial and temporal triggering functions are normalized so that their product integrates to one over space and time, the stationarity condition simplifies to

μ(s,t)\mu(\mathbf{s},t)1

with μ(s,t)\mu(\mathbf{s},t)2 interpreted as the reproduction number or branching ratio (Bernabeu et al., 18 Nov 2025). A related nonparametric Bayesian framework states the non-explosion condition as

μ(s,t)\mu(\mathbf{s},t)3

again linking process stability to the integrated triggering strength (Liu et al., 30 Mar 2026). This suggests that separability does not alter the basic Hawkes subcriticality criterion; rather, it structures how that criterion is parameterized.

4. Likelihoods, estimation, and computational strategies

The canonical likelihood for a spatio-temporal Hawkes process is

μ(s,t)\mu(\mathbf{s},t)4

or, in equivalent notation,

μ(s,t)\mu(\mathbf{s},t)5

(Jun et al., 2022, Bernabeu et al., 18 Nov 2025). Under separability, the compensator can be decomposed into background and offspring terms, and the offspring integral itself can be expressed through separate temporal and spatial factors (Bernabeu et al., 18 Nov 2025).

Three broad inferential traditions recur in the literature. The first is direct maximum-likelihood estimation, often with numerical approximation of the compensator. The second is EM or stochastic declustering, which uses the branching representation to alternate between posterior parent probabilities and parameter updates (Yuan et al., 2018, Bernabeu et al., 18 Nov 2025). The third is Bayesian data augmentation, where latent exact event times, latent locations, and latent branching labels are imputed when observations are aggregated into time bins and spatial regions (Zhou et al., 2022).

Several computational variants specialize the separable model. A randomized-kernel approach retains a separable exponential–Gaussian triggering structure but replaces the Gaussian spatial kernel with a random Fourier feature approximation, then learns the parameters by mini-batch gradient descent on the negative log-likelihood while enforcing positivity by softplus and orthonormality by projection (Ilhan et al., 2020). A flexible parametric inference framework based on kernels with finite support, domain discretization, and approximate precomputations does not require separability in general, but its synthetic studies instantiate

μ(s,t)\mu(\mathbf{s},t)6

so the solver applies directly to separable space-time triggering kernels (Siviero et al., 2024). A Bayesian framework for aggregated data uses Gibbs sampling for conjugate parameters and Metropolis–Hastings for nonconjugate parameters under the exponential–Gaussian separable model (Zhou et al., 2022).

A unified comparison of three inference strategies for separable models—likelihood, EM, and inlabru—reports that all three methods are generally accurate; likelihood and inlabru are usually the most accurate in parameter recovery; EM is fastest; and inlabru often has the lowest MAE overall in the reported tables (Bernabeu et al., 18 Nov 2025). The same study also reports a clear grid-size tradeoff for likelihood: increasing grid resolution improves MAE and reduces variability, but runtime increases sharply.

5. Empirical roles and application domains

Separable spatio-temporal Hawkes processes are used across terrorism, earthquakes, crime, social media, traffic forecasting, and network reconstruction. In terrorism modeling, a standard separable exponential–Gaussian trigger serves as the baseline from which more flexible nonstationary and nonseparable models are developed (Jun et al., 2022). In a randomized-kernel framework for spatio-temporal Hawkes processes, synthetic and real-data experiments on earthquake and Chicago crime data show that the proposed RFF-GD method outperforms EM-based estimation and stochastic declustering in negative log-likelihood per event and AIC, while also revealing interpretable type-to-type excitation patterns (Ilhan et al., 2020).

In network reconstruction, the separable formulation

μ(s,t)\mu(\mathbf{s},t)7

is used to infer latent connectivity from event cascades with spatial locations (Yuan et al., 2018). The paper reports that incorporating spatial information improves recovery of reciprocity, edge structure, community structure, and kernel shape relative to temporal-only Hawkes models. On Gowalla friendship networks, the nonparametric spatiotemporal model achieves higher AUC than the temporal and Bayesian baselines, and on crime-event data the inferred parent-probability matrix supports stochastic declustering and motif analysis (Yuan et al., 2018).

For aggregated observations, separable models remain practically useful because exact timestamps and exact spatial locations are often unavailable. A Bayesian approach for temporally and spatially aggregated counts shows that μ(s,t)\mu(\mathbf{s},t)8 and μ(s,t)\mu(\mathbf{s},t)9 are robust to aggregation, that g(s,t,si,ti)g(\mathbf{s},t,\mathbf{s}_i,t_i)0 is mainly affected by temporal aggregation, and that g(s,t,si,ti)g(\mathbf{s},t,\mathbf{s}_i,t_i)1 is mainly affected by spatial aggregation (Zhou et al., 2022). In the Iraq application, the fitted multivariate separable model yields strong self-excitation in both IED attacks and airstrikes, weak cross-excitation overall, a non-negligible estimate g(s,t,si,ti)g(\mathbf{s},t,\mathbf{s}_i,t_i)2 from airstrikes to IEDs, and a near-zero estimate g(s,t,si,ti)g(\mathbf{s},t,\mathbf{s}_i,t_i)3 in the reverse direction (Zhou et al., 2022).

Simulation studies devoted specifically to separable spatio-temporal Hawkes models compare exponential or power-law temporal kernels with Gaussian or exponential spatial kernels under constant background (Bernabeu et al., 18 Nov 2025). Those experiments indicate that stronger spatial clustering and longer temporal memory can improve estimation accuracy, and they provide practical guidance about when separability is especially advantageous for simulation and inference.

6. Limits of separability and major generalizations

A common misconception is that any model exhibiting a product structure between space and time is therefore a separable spatio-temporal Hawkes process in the classical sense. Recent work explicitly rejects that identification. In a graphon-induced spatiotemporal Hawkes process, the linear excitation term has the multiplicative form

g(s,t,si,ti)g(\mathbf{s},t,\mathbf{s}_i,t_i)4

but the model is positioned as a graphon-based generalization of multivariate and network Hawkes processes rather than as a classical separable space-time kernel model (Baars et al., 2024). In a multivariate spatio-temporal neural Hawkes process, the latent memory decay contains the separable exponential factor

g(s,t,si,ti)g(\mathbf{s},t,\mathbf{s}_i,t_i)5

yet the paper states that the overall intensity is not constrained to be separable because there is no predefined triggering kernel g(s,t,si,ti)g(\mathbf{s},t,\mathbf{s}_i,t_i)6 (Chukwuemeka et al., 27 Feb 2026).

A second misconception is that separability is merely a harmless simplification. Multiple empirical studies present it as restrictive when space-time interaction is substantively important. In terrorism data, separability becomes limiting when cross-triggering peaks away from zero spatial lag or when spatial dispersion changes with elapsed time; the Nigeria analysis therefore favors a fully flexible model with nonseparable marginal and cross-triggering structure, while the Afghanistan analysis finds that nonstationary triggering is more consequential than nonseparability once population-dependent effects are included (Jun et al., 2022). In earthquake modeling, earlier separable ETAS formulations are criticized for forcing the same temporal decay profile at every spatial distance, motivating a non-separable triggering density with spatially varying productivity and anisotropy (Kwon et al., 2022).

Generalizations proceed along several directions. One route is additive Gaussian-process modeling of the background rate and triggering kernel, where the latent covariance can be separable or additive: g(s,t,si,ti)g(\mathbf{s},t,\mathbf{s}_i,t_i)7 This formulation is presented as improving interpretability by decoupling temporal, spatial, and interaction effects while often improving numerical stability relative to product-separable kernels (Liu et al., 30 Mar 2026). Another route is flexible finite-support parametric inference that can estimate any differentiable spatio-temporal kernel and uses separable kernels only as one family among several experimentally evaluated choices (Siviero et al., 2024).

Separable spatio-temporal Hawkes processes thus occupy a dual position in the literature. They remain a standard reference model for self-exciting event data because they preserve the branching interpretation, admit clear temporal and spatial parameterization, and are computationally adaptable to likelihood, EM, Bayesian, and simulation procedures (Bernabeu et al., 18 Nov 2025, Zhou et al., 2022). At the same time, they now function equally as a baseline against which nonseparable, multivariate, graph-structured, and latent-state models are defined and assessed (Jun et al., 2022, Baars et al., 2024, Chukwuemeka et al., 27 Feb 2026, Liu et al., 30 Mar 2026).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Separable Spatio-Temporal Hawkes Process.