Papers
Topics
Authors
Recent
Search
2000 character limit reached

Volatility-Modulated Volterra Processes

Updated 17 April 2026
  • Volatility-modulated Volterra processes are non-Markovian stochastic systems characterized by path-dependent dynamics and a stochastic volatility factor.
  • They employ convolution-based drift and diffusion functions with decaying kernels (exponential or polynomial) to manage memory effects.
  • A low-dimensional reduction via Riemannian NPC geometry and Hyper-Geometric Networks enables efficient approximation and robust simulation.

A volatility-modulated Volterra process is a class of non-Markovian stochastic process characterized by path-dependent dynamics where both the drift and diffusion coefficients are convolutions (or more general functionals) of the past trajectory, modulated by a stochastic volatility factor. This framework generalizes the classical semimartingale models, incorporates rough and persistent memory effects, enables a direct connection to stochastic volatility and modulation phenomena observed in financial time series, and is central to contemporary rough volatility modeling and high-dimensional stochastic analysis.

1. Mathematical Structure of Volatility-Modulated Volterra Processes

Consider a filtered probability space and discrete time t=0,,Tt=0,\dots,T (with continuous-time analogues provided via stochastic integrals). A dd-dimensional volatility-modulated Volterra process (Xt)t=0T(X_t)_{t=0}^T is defined recursively by

Xt+1=Xt+Drift(t,X0:t)+Diffusion(t,X0:t,S0:t)WtX_{t+1} = X_t + \text{Drift}(t, X_{0:t}) + \text{Diffusion}(t, X_{0:t}, S_{0:t}) \cdot W_t

where:

  • WtW_t are i.i.d. Nd(0,I)\mathcal{N}_d(0, I) (standard Gaussian noise).
  • StSym(d)S_t \in \mathrm{Sym}(d) is an independent, latent stochastic volatility factor with almost sure Frobenius bound StFR\|S_t\|_F \leq R.
  • The drift and diffusion are of Volterra type:

Drift(t,x0:t)=r=0tκ(t,r)μ(t,xr) Diffusion(t,x0:t,s0:t)=exp(12r=0tκ(t,r)[σ(t,xr)+sr])\begin{aligned} \text{Drift}(t, x_{0:t}) &= \sum_{r=0}^t \kappa(t, r) \cdot \mu(t, x_r) \ \text{Diffusion}(t, x_{0:t}, s_{0:t}) &= \exp \left( \frac{1}{2} \sum_{r=0}^t \kappa(t, r)[\sigma(t, x_r) + s_r] \right) \end{aligned}

where μ:R1+dRd\mu : \mathbb{R}^{1+d} \to \mathbb{R}^d, dd0 are dd1-Lipschitz, and dd2 is a nonnegative kernel with dd3, dd4 for dd5.

Decay conditions on dd6 ensure controlled memory:

  • Exponential decay: dd7 for some dd8, dd9,
  • Polynomial decay: (Xt)t=0T(X_t)_{t=0}^T0 for (Xt)t=0T(X_t)_{t=0}^T1, (Xt)t=0T(X_t)_{t=0}^T2.

These settings ensure that the impact of distant history on present dynamics diminishes sufficiently rapidly, which is necessary for stability and tractability of the models (Arabpour et al., 2024).

2. Conditional Law and Intrinsic Geometry

Given a realized path (Xt)t=0T(X_t)_{t=0}^T3, the conditional law of (Xt)t=0T(X_t)_{t=0}^T4, conditional on both (Xt)t=0T(X_t)_{t=0}^T5, is Gaussian: (Xt)t=0T(X_t)_{t=0}^T6 Typically, (Xt)t=0T(X_t)_{t=0}^T7 is unobserved, only known to satisfy a bound. Marginalizing (Xt)t=0T(X_t)_{t=0}^T8, one obtains a random probability measure (Xt)t=0T(X_t)_{t=0}^T9 on the space of Xt+1=Xt+Drift(t,X0:t)+Diffusion(t,X0:t,S0:t)WtX_{t+1} = X_t + \text{Drift}(t, X_{0:t}) + \text{Diffusion}(t, X_{0:t}, S_{0:t}) \cdot W_t0-dimensional Gaussian measures. The conditional law Xt+1=Xt+Drift(t,X0:t)+Diffusion(t,X0:t,S0:t)WtX_{t+1} = X_t + \text{Drift}(t, X_{0:t}) + \text{Diffusion}(t, X_{0:t}, S_{0:t}) \cdot W_t1 (with Xt+1=Xt+Drift(t,X0:t)+Diffusion(t,X0:t,S0:t)WtX_{t+1} = X_t + \text{Drift}(t, X_{0:t}) + \text{Diffusion}(t, X_{0:t}, S_{0:t}) \cdot W_t2 integrated out) is then the barycenter (intrinsic mean) of Xt+1=Xt+Drift(t,X0:t)+Diffusion(t,X0:t,S0:t)WtX_{t+1} = X_t + \text{Drift}(t, X_{0:t}) + \text{Diffusion}(t, X_{0:t}, S_{0:t}) \cdot W_t3 in the manifold of nondegenerate Gaussian laws endowed with a canonical nonpositive-curvature (NPC) Riemannian metric (Arabpour et al., 2024).

The explicit conditional map Xt+1=Xt+Drift(t,X0:t)+Diffusion(t,X0:t,S0:t)WtX_{t+1} = X_t + \text{Drift}(t, X_{0:t}) + \text{Diffusion}(t, X_{0:t}, S_{0:t}) \cdot W_t4 is infinite-dimensional and neither smooth nor tractable, leading to a severe curse of dimensionality.

3. Low-Dimensional Reduction via Riemannian NPC Geometry

To address dimensionality, the law Xt+1=Xt+Drift(t,X0:t)+Diffusion(t,X0:t,S0:t)WtX_{t+1} = X_t + \text{Drift}(t, X_{0:t}) + \text{Diffusion}(t, X_{0:t}, S_{0:t}) \cdot W_t5 is projected onto the finite-dimensional statistical manifold Xt+1=Xt+Drift(t,X0:t)+Diffusion(t,X0:t,S0:t)WtX_{t+1} = X_t + \text{Drift}(t, X_{0:t}) + \text{Diffusion}(t, X_{0:t}, S_{0:t}) \cdot W_t6 of Gaussian distributions, equipped with the Xt+1=Xt+Drift(t,X0:t)+Diffusion(t,X0:t,S0:t)WtX_{t+1} = X_t + \text{Drift}(t, X_{0:t}) + \text{Diffusion}(t, X_{0:t}, S_{0:t}) \cdot W_t7-NPC Riemannian structure: Xt+1=Xt+Drift(t,X0:t)+Diffusion(t,X0:t,S0:t)WtX_{t+1} = X_t + \text{Drift}(t, X_{0:t}) + \text{Diffusion}(t, X_{0:t}, S_{0:t}) \cdot W_t8 This space Xt+1=Xt+Drift(t,X0:t)+Diffusion(t,X0:t,S0:t)WtX_{t+1} = X_t + \text{Drift}(t, X_{0:t}) + \text{Diffusion}(t, X_{0:t}, S_{0:t}) \cdot W_t9 is CAT(0)—complete, simply connected, and non-positively curved—allowing for uniquely defined barycenters (Arabpour et al., 2024). The explicit geodesic distance is

WtW_t0

where WtW_t1 are eigenvalues of WtW_t2.

For each WtW_t3, define the map WtW_t4: WtW_t5 thus replacing the high-dimensional conditional law by its unique barycenter in WtW_t6. The mapping WtW_t7 is Lipschitz and only weakly depends on remote past due to decay of WtW_t8 (quantified in Theorem 3.15).

4. Hypernetwork Architecture for Universal Approximation

The low-dimensional projected law WtW_t9 is not, in general, analytically tractable. To learn the mapping Nd(0,I)\mathcal{N}_d(0, I)0, (Arabpour et al., 2024) introduces a sequential deep-learning model called the Hyper-Geometric Network (HGN):

  • For each fixed time Nd(0,I)\mathcal{N}_d(0, I)1, approximate the "static" Nd(0,I)\mathcal{N}_d(0, I)2 using a geometric deep net (GDN), i.e., a feedforward ReLU MLP acting on log/exp chart representations of Nd(0,I)\mathcal{N}_d(0, I)3.
  • A hypernetwork Nd(0,I)\mathcal{N}_d(0, I)4 iteratively evolves a latent state Nd(0,I)\mathcal{N}_d(0, I)5 and a linear read-out Nd(0,I)\mathcal{N}_d(0, I)6 produces the parameters Nd(0,I)\mathcal{N}_d(0, I)7 for the per-time GDN.
  • This construction yields a mixture-of-experts interpretation in which each Nd(0,I)\mathcal{N}_d(0, I)8 has a specialist expert determined by the hypernetwork trajectory.

Universal approximation theorems (see Theorem 4.14) demonstrate that any causal map Nd(0,I)\mathcal{N}_d(0, I)9 of finite memory and regularity can be uniformly approximated by an HGN with parameter complexity polynomial in both the time horizon StSym(d)S_t \in \mathrm{Sym}(d)0 and error level StSym(d)S_t \in \mathrm{Sym}(d)1. The architecture can encode non-stationary and time-inhomogeneous behavior efficiently (Arabpour et al., 2024).

5. Implementation and Empirical Insights

Algorithmic implementation consists of:

  • Synthetic data generation for StSym(d)S_t \in \mathrm{Sym}(d)2 (Algorithm 6.1).
  • Computing the barycentric Gaussian projection StSym(d)S_t \in \mathrm{Sym}(d)3 (Algorithm 6.2).
  • Training pipeline (Algorithm 6.3): Each per-time expert is fit independently via intrinsic MSE on StSym(d)S_t \in \mathrm{Sym}(d)4; the hypernetwork StSym(d)S_t \in \mathrm{Sym}(d)5 is then fit by least squares across StSym(d)S_t \in \mathrm{Sym}(d)6.
  • The loss combines expert-wise geodesic MSE on StSym(d)S_t \in \mathrm{Sym}(d)7 and a hyper loss for temporal parameter consistency.

Ablation studies confirm the model tracks the performance of per-time GDNs for a wide range of process parameters, volatility-of-volatility, memory, and problem dimension. Training is parallel in time, avoiding backpropagation through the whole trajectory. Empirical error decays and memory truncation effects are consistent with theoretical predictions for both Lipschitz continuity and memory decay (Arabpour et al., 2024).

6. Theoretical Implications and Research Directions

This approach provides a rigorous dimension reduction for conditional evolution laws of general volatility-modulated Volterra processes, encapsulates the non-smooth structure (memory, non-stationarity, roughness), and enables effective statistical or deep-learning-based prediction of law-valued outputs. The NPC-projected law framework bridges infinite-dimensional stochastic analysis with differentiable geometry and deep learning. It enables practical, robust model estimation and simulation for processes central in rough volatility, path-dependent SPDEs, high-frequency finance, and other fields requiring precise control of high-dimensional pathwise uncertainty (Arabpour et al., 2024).

This methodology, emphasizing the conditional law's geometry and its learnability on statistical manifolds, is a key development for non-Markovian, memory-driven stochastic systems. It brings together recent advances in stochastic process theory, statistical geometry, and deep learning, and opens new avenues for efficient simulation, calibration, and control in rough volatility modeling and beyond.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Volatility-Modulated Volterra Processes.