Papers
Topics
Authors
Recent
Search
2000 character limit reached

RBF-Lifted Signature Kernel

Updated 31 December 2025
  • RBF-lifted signature kernel is a universal, positive definite kernel that lifts continuous paths into a reproducing kernel Hilbert space using Gaussian RBF and signature transforms.
  • It employs random Fourier feature approximations to efficiently scale the computation of signature kernels while ensuring uniform error bounds under sub-Gaussian assumptions.
  • Variants like Diagonal Projection and Tensor Random Projection provide effective trade-offs, making the approach practical for large-scale time-series and sequence analysis.

The RBF-lifted signature kernel is a universal and characteristic positive definite kernel on the space of continuous paths, combining the expressivity of tensor algebra signatures with the nonlinear similarity afforded by the Gaussian radial basis function (RBF). It measures path similarity by lifting Euclidean increments into a reproducing kernel Hilbert space (RKHS), applying the signature transformation, and computing the Hilbert-Schmidt inner product in the resulting tensor algebra. Recent research demonstrates both explicit constructions and scalable random-feature approximations for this kernel, enabling efficient application to large-scale sequence and time-series analysis tasks (Toth et al., 2023, Piatti et al., 29 Dec 2025).

1. Mathematical Definition and Construction

Let x,y:[0,T]→Rdx, y : [0,T] \to \mathbb{R}^d be continuous paths of finite pp-variation. The full signature of xx over [0,T][0,T] is given by

Sig⁡(x)0,T=(1, S1(x), S2(x), … )∈T((Rd))\operatorname{Sig}(x)_{0,T} = \left(1,\,S^1(x),\,S^2(x),\,\dots\right) \in T((\mathbb{R}^d))

where

Sk(x)=∫0<t1<⋯<tk<Tdxt1⊗⋯⊗dxtk∈(Rd)⊗k.S^k(x) = \int_{0 < t_1 < \cdots < t_k < T} dx_{t_1} \otimes \cdots \otimes dx_{t_k} \in (\mathbb{R}^d)^{\otimes k}.

For the RBF-lifted signature kernel,

  • The static RBF kernel is κσ(u,v)=exp⁡(−12σ2∥u−v∥2)\kappa_\sigma(u, v) = \exp\Big(-\frac{1}{2\sigma^2}\|u - v\|^2\Big).
  • Its RKHS, Hσ\mathcal{H}_\sigma, is L2(Rd,dμσ)L^2(\mathbb{R}^d, d\mu_\sigma), with dμσd\mu_\sigma being the spectral measure associated to pp0 via Bochner’s theorem.
  • The feature map is pp1.

To construct the kernel,

  • Lift the path pp2 pointwise into pp3 via pp4,
  • Compute the signature in pp5,
  • The RBF-lifted signature kernel is

pp6

This construction is universal and characteristic on path space and possesses invariance and stability properties inherited from both the signature and the RBF kernel (Piatti et al., 29 Dec 2025).

2. Random Fourier Signature Feature Approximations

Direct computation of pp7 is infeasible for long paths due to exponential growth in tensor dimensions. Random Fourier feature-based acceleration replaces the static embedding by randomized mappings to approximate the inner products.

  • Draw pp8 i.i.d. samples pp9.
  • The RFF map is: xx0
  • In signature computations, replace kernel embeddings xx1 with xx2 at each step and truncate at signature level xx3.

The resultant random feature signature map is: xx4 whose inner product yields an unbiased estimator for the truncated RBF-lifted signature kernel: xx5 The expected value of this estimator is exactly the truncated signature kernel, xx6 (Toth et al., 2023).

3. Uniform Approximation Guarantees

Concentration inequalities bound the uniform error of the RFF-accelerated kernel over compact path spaces.

For compact convex xx7 with paths of bounded xx8-variation xx9, under sub-Gaussian moment bounds on the RFF distribution, there exist constants [0,T][0,T]0 such that for each signature level [0,T][0,T]1 and error [0,T][0,T]2: [0,T][0,T]3 with precise bounds detailed in Theorem 3.1 and equation (3.13) (Toth et al., 2023). To guarantee uniform error [0,T][0,T]4 for all [0,T][0,T]5 with probability [0,T][0,T]6, one requires [0,T][0,T]7 RFF draws.

4. Computational Complexity and Scalable Variants

The exact signature kernel (e.g., as in Király–Oberhauser 2019) incurs [0,T][0,T]8 time for [0,T][0,T]9 paths of length Sig⁡(x)0,T=(1, S1(x), S2(x), … )∈T((Rd))\operatorname{Sig}(x)_{0,T} = \left(1,\,S^1(x),\,S^2(x),\,\dots\right) \in T((\mathbb{R}^d))0 and truncation Sig⁡(x)0,T=(1, S1(x), S2(x), … )∈T((Rd))\operatorname{Sig}(x)_{0,T} = \left(1,\,S^1(x),\,S^2(x),\,\dots\right) \in T((\mathbb{R}^d))1. RBF-lifted signature kernels with RFF approximation scale as: Sig⁡(x)0,T=(1, S1(x), S2(x), … )∈T((Rd))\operatorname{Sig}(x)_{0,T} = \left(1,\,S^1(x),\,S^2(x),\,\dots\right) \in T((\mathbb{R}^d))2 eliminating the quadratic dependence in both dataset size and sequence length (Toth et al., 2023).

To further improve scalability, two variants are introduced:

Variant Feature Dimension Time Complexity
Diagonal Projection (DP) Sig⁡(x)0,T=(1, S1(x), S2(x), … )∈T((Rd))\operatorname{Sig}(x)_{0,T} = \left(1,\,S^1(x),\,S^2(x),\,\dots\right) \in T((\mathbb{R}^d))3 (per-level) Sig⁡(x)0,T=(1, S1(x), S2(x), … )∈T((Rd))\operatorname{Sig}(x)_{0,T} = \left(1,\,S^1(x),\,S^2(x),\,\dots\right) \in T((\mathbb{R}^d))4
Tensor Random Projection (TRP) Sig⁡(x)0,T=(1, S1(x), S2(x), … )∈T((Rd))\operatorname{Sig}(x)_{0,T} = \left(1,\,S^1(x),\,S^2(x),\,\dots\right) \in T((\mathbb{R}^d))5 Sig⁡(x)0,T=(1, S1(x), S2(x), … )∈T((Rd))\operatorname{Sig}(x)_{0,T} = \left(1,\,S^1(x),\,S^2(x),\,\dots\right) \in T((\mathbb{R}^d))6
  • DP averages only the diagonal of the RFF tensor product, reducing features at the expense of slower concentration.
  • TRP sketches each tensor by a CP-rank-1 Gaussian map, yielding sub-exponential convergence with respect to the number of projections.

Empirical results demonstrate negligible accuracy loss for moderate Sig⁡(x)0,T=(1, S1(x), S2(x), … )∈T((Rd))\operatorname{Sig}(x)_{0,T} = \left(1,\,S^1(x),\,S^2(x),\,\dots\right) \in T((\mathbb{R}^d))7 and Sig⁡(x)0,T=(1, S1(x), S2(x), … )∈T((Rd))\operatorname{Sig}(x)_{0,T} = \left(1,\,S^1(x),\,S^2(x),\,\dots\right) \in T((\mathbb{R}^d))8—with both DP and TRP matching the full signature kernel on time series benchmarks and scaling efficiently to Sig⁡(x)0,T=(1, S1(x), S2(x), … )∈T((Rd))\operatorname{Sig}(x)_{0,T} = \left(1,\,S^1(x),\,S^2(x),\,\dots\right) \in T((\mathbb{R}^d))9 sequences (Toth et al., 2023).

5. Dynamic Lifting via RF-CDE

A complementary approach uses dynamic random-feature reservoirs via random Fourier controlled differential equations (RF-CDEs):

  • At each time Sk(x)=∫0<t1<⋯<tk<Tdxt1⊗⋯⊗dxtk∈(Rd)⊗k.S^k(x) = \int_{0 < t_1 < \cdots < t_k < T} dx_{t_1} \otimes \cdots \otimes dx_{t_k} \in (\mathbb{R}^d)^{\otimes k}.0, map Sk(x)=∫0<t1<⋯<tk<Tdxt1⊗⋯⊗dxtk∈(Rd)⊗k.S^k(x) = \int_{0 < t_1 < \cdots < t_k < T} dx_{t_1} \otimes \cdots \otimes dx_{t_k} \in (\mathbb{R}^d)^{\otimes k}.1 to RFFs as Sk(x)=∫0<t1<⋯<tk<Tdxt1⊗⋯⊗dxtk∈(Rd)⊗k.S^k(x) = \int_{0 < t_1 < \cdots < t_k < T} dx_{t_1} \otimes \cdots \otimes dx_{t_k} \in (\mathbb{R}^d)^{\otimes k}.2.
  • Feed this lifted signal to a random linear CDE: Sk(x)=∫0<t1<⋯<tk<Tdxt1⊗⋯⊗dxtk∈(Rd)⊗k.S^k(x) = \int_{0 < t_1 < \cdots < t_k < T} dx_{t_1} \otimes \cdots \otimes dx_{t_k} \in (\mathbb{R}^d)^{\otimes k}.3 yielding a feature vector Sk(x)=∫0<t1<⋯<tk<Tdxt1⊗⋯⊗dxtk∈(Rd)⊗k.S^k(x) = \int_{0 < t_1 < \cdots < t_k < T} dx_{t_1} \otimes \cdots \otimes dx_{t_k} \in (\mathbb{R}^d)^{\otimes k}.4. Only a linear readout on top of Sk(x)=∫0<t1<⋯<tk<Tdxt1⊗⋯⊗dxtk∈(Rd)⊗k.S^k(x) = \int_{0 < t_1 < \cdots < t_k < T} dx_{t_1} \otimes \cdots \otimes dx_{t_k} \in (\mathbb{R}^d)^{\otimes k}.5 is trained.

In the infinite-width limit Sk(x)=∫0<t1<⋯<tk<Tdxt1⊗⋯⊗dxtk∈(Rd)⊗k.S^k(x) = \int_{0 < t_1 < \cdots < t_k < T} dx_{t_1} \otimes \cdots \otimes dx_{t_k} \in (\mathbb{R}^d)^{\otimes k}.6, Sk(x)=∫0<t1<⋯<tk<Tdxt1⊗⋯⊗dxtk∈(Rd)⊗k.S^k(x) = \int_{0 < t_1 < \cdots < t_k < T} dx_{t_1} \otimes \cdots \otimes dx_{t_k} \in (\mathbb{R}^d)^{\otimes k}.7, rigorously establishing that RF-CDEs realize the RBF-lifted signature kernel as their limiting covariance (Piatti et al., 29 Dec 2025).

6. Theoretical Properties

Sk(x)=∫0<t1<⋯<tk<Tdxt1⊗⋯⊗dxtk∈(Rd)⊗k.S^k(x) = \int_{0 < t_1 < \cdots < t_k < T} dx_{t_1} \otimes \cdots \otimes dx_{t_k} \in (\mathbb{R}^d)^{\otimes k}.8 is positive definite (Mercer kernel), universal, and characteristic. Universality and characteristicness are inherited from the classical signature kernel and the RBF kernel on Sk(x)=∫0<t1<⋯<tk<Tdxt1⊗⋯⊗dxtk∈(Rd)⊗k.S^k(x) = \int_{0 < t_1 < \cdots < t_k < T} dx_{t_1} \otimes \cdots \otimes dx_{t_k} \in (\mathbb{R}^d)^{\otimes k}.9 (Piatti et al., 29 Dec 2025).

  • Reparameterization invariance results from the time invariance of path signatures.
  • Stability under κσ(u,v)=exp⁡(−12σ2∥u−v∥2)\kappa_\sigma(u, v) = \exp\Big(-\frac{1}{2\sigma^2}\|u - v\|^2\Big)0-variation: small perturbations in path yield small changes in κσ(u,v)=exp⁡(−12σ2∥u−v∥2)\kappa_\sigma(u, v) = \exp\Big(-\frac{1}{2\sigma^2}\|u - v\|^2\Big)1.
  • Equivariance with respect to space translation and rotation is inherited from the RBF base kernel.

This kernel provides a continuous-time, non-Euclidean analogue of RBF feature learning in sequence learning.

Selection of hyperparameters follows empirical trade-offs:

  • κσ(u,v)=exp⁡(−12σ2∥u−v∥2)\kappa_\sigma(u, v) = \exp\Big(-\frac{1}{2\sigma^2}\|u - v\|^2\Big)2 suffices for κσ(u,v)=exp⁡(−12σ2∥u−v∥2)\kappa_\sigma(u, v) = \exp\Big(-\frac{1}{2\sigma^2}\|u - v\|^2\Big)3 kernel error.
  • Use the Gaussian spectral measure or improvements such as quasi-Monte Carlo or orthogonal realizations.
  • Truncation κσ(u,v)=exp⁡(−12σ2∥u−v∥2)\kappa_\sigma(u, v) = \exp\Big(-\frac{1}{2\sigma^2}\|u - v\|^2\Big)4 captures principal path interactions with exponential cost beyond.
  • For large-dimensional data and small κσ(u,v)=exp⁡(−12σ2∥u−v∥2)\kappa_\sigma(u, v) = \exp\Big(-\frac{1}{2\sigma^2}\|u - v\|^2\Big)5, DP is recommended; for moderate κσ(u,v)=exp⁡(−12σ2∥u−v∥2)\kappa_\sigma(u, v) = \exp\Big(-\frac{1}{2\sigma^2}\|u - v\|^2\Big)6 and κσ(u,v)=exp⁡(−12σ2∥u−v∥2)\kappa_\sigma(u, v) = \exp\Big(-\frac{1}{2\sigma^2}\|u - v\|^2\Big)7, TRP reduces memory; vanilla RFSF has superior concentration but balloons exponentially in κσ(u,v)=exp⁡(−12σ2∥u−v∥2)\kappa_\sigma(u, v) = \exp\Big(-\frac{1}{2\sigma^2}\|u - v\|^2\Big)8.

The rough signature kernel is the limiting case when no nonlinear RBF warping is applied, corresponding to linear (log-)signature propagation in a random reservoir (Piatti et al., 29 Dec 2025). RF-CDE and R-RDE offer two complementary methods whose infinite-width limits recover the RBF-lifted and rough signature kernels, respectively, unifying perspectives on random-feature reservoirs and continuous-time deep sequence models.

References

Definition Search Book Streamline Icon: https://streamlinehq.com
References (2)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to RBF-Lifted Signature Kernel.