---
title: Expected Signature Kernel
url: https://www.emergentmind.com/topics/expected-signature-kernel
type: topic
---

# Expected Signature Kernel

The expected signature kernel is a kernel on path space or on probability measures over path space built from signatures, the iterated-integral coordinates of paths in the tensor algebra. In its law-space form, it is the inner product of mean signatures, \(K(P,Q)=\langle \mathbb E_{X\sim P}[S(X)],\mathbb E_{Y\sim Q}[S(Y)]\rangle\), and therefore inherits positive-definiteness from the Hilbert structure of the tensor algebra [2506.17634]. Closely related work studies kernels obtained by averaging randomised developments of deterministic paths into matrix Lie groups; in the large-\(N\) limit, these constructions recover the ordinary signature kernel for \(GL(N)\) and yield a unitary analogue described by free probability and a quadratic functional equation [2402.12311].

## 1. Signature-theoretic foundation

Let \(V\simeq \mathbb R^d\) and let \(\gamma:[a,b]\to V\) be a continuous path of finite \(1\)-variation. Its signature \(S(\gamma)\in T((V))\) is defined by the controlled differential equation
\[
d z_t = z_t \otimes d\gamma_t,\qquad z_a = 1\in T((V)),
\]
with \(S(\gamma)=z_b\). Equivalently, the degree-\(m\) coordinate is the iterated integral
\[
S^{I}_{s,t}(\gamma) = \int_{s<t_1<\cdots<t_m<t} d\gamma_{t_1}^{I_1}\cdots d\gamma_{t_m}^{I_m},\qquad I\in\{1,\dots,d\}^m
\]
[2402.12311].

The classical signature kernel on path space is the Hilbert–Schmidt inner product
\[
K_{\mathrm{sig}}(\gamma,\sigma)=\langle S(\gamma),S(\sigma)\rangle
=\sum_{m=0}^\infty \langle S^m(\gamma),S^m(\sigma)\rangle_{HS}.
\]
In the Hilbert-space formulation, if \(H\) is a real Hilbert space and \(x:[0,1]\to H\) is of bounded variation, then each \(S^m(x)\in H^{\otimes m}\) is an “ordered” analogue of the \(m\)-th sample moment of the infinitesimal increments, and
\[
K(x,y)=\langle S(x),S(y)\rangle_{T(H)}
=\sum_{m=0}^\infty \langle S^m(x),S^m(y)\rangle_{H^{\otimes m}}
\]
[1601.08169].

This kernel is manifestly positive-definite because it is an inner product in feature space. For bounded-variation paths, the associated RKHS is universal on compact sets in the sense described via Stone–Weierstrass, and the signature is invariant to reparameterization up to tree-like equivalence; augmenting a path by time, \(x\mapsto \hat x_t=(t,x_t)\), restores sensitivity to explicit timing when required [1601.08169] [2506.17634].

A weighted generalisation replaces the unweighted tensor inner product by
\[
K^\phi(\gamma,\sigma)
=\sum_{k=0}^\infty \phi(k)\,\langle S(\gamma)^k,S(\sigma)^k\rangle_k,
\]
for nonnegative weights \(\phi(k)\) satisfying the factorial-type summability condition
\[
\sum_{k=0}^\infty C^k\,\phi(k)\,(k!)^{-2}<\infty\qquad \forall\,C>0.
\]
The two-parameter version \(K^\phi(\gamma,\sigma)(s,t)=\langle S(\gamma)_{0,s},S(\sigma)_{0,t}\rangle_\phi\) is central in PDE-based formulations [2107.00447].

## 2. From deterministic paths to probability measures

For probability distributions \(P,Q\) on \(Paths(V)\), the Expected Signature Kernel is defined by
\[
K(P,Q)=\left\langle \mathbb E_{X\sim P}[S(X)],\,\mathbb E_{Y\sim Q}[S(Y)]\right\rangle_{T(V)}.
\]
Writing \(\mu_P=\mathbb E_{X\sim P}[S(X)]\) and \(\mu_Q=\mathbb E_{Y\sim Q}[S(Y)]\), this becomes
\[
K(P,Q)=\langle \mu_P,\mu_Q\rangle_{T(V)}
=\sum_{m=0}^\infty \left\langle \mathbb E[S_m(X)],\,\mathbb E[S_m(Y)]\right\rangle_{V^{\otimes m}}.
\]
In discrete time, with a static kernel lift \(k:X\times X\to\mathbb R\) and truncation \(m\le M\), one obtains the finite-level approximation
\[
K(P,Q)\approx \sum_{m=0}^M \langle E[\mathrm{signature}_m(X)],E[\mathrm{signature}_m(Y)]\rangle_{H_k^{\otimes m}}
\]
[2506.17634].

This law-space kernel is closely tied to kernel-MMD constructions. For two measures \(P,Q\), the signature-MMD satisfies
\[
\mathrm{MMD}^2_{\mathrm{Sig}}(P,Q)
=\sum_{w\in\mathcal W}\bigl(m^w(P)-m^w(Q)\bigr)^2,
\]
with \(m^w(\mu)=\mathbb E_{\gamma\sim\mu}[\mathrm{Sig}^w(\gamma)]\), and at level \(N\),
\[
\mathrm{MMD}_N(\mu,\nu)^2
=\bigl\|\mathbb E_\mu[S^{(N)}]-\mathbb E_\nu[S^{(N)}]\bigr\|^2
\]
[2509.07893] [2606.28869].

A terminological point is important. Some treatments distinguish the expectation of the pathwise signature kernel,
\[
\mathbb E[K(X,Y)] = \mathbb E[\langle S(X),S(Y)\rangle],
\]
from the inner product of expected signatures,
\[
K_{\mathrm{exp}}(X,Y)=\langle \mathbb E[S(X)],\mathbb E[S(Y)]\rangle.
\]
The latter is the standard law-space expected signature kernel in the definitions above [2606.28869].

The expected signature itself is not generally group-like. If \(X\) is a random weakly-geometric path, then
\[
\mathbb E[S(X)]_u\,\mathbb E[S(X)]_v-\mathbb E[S(X)]_{u\shuffle v}
=-\,\mathrm{Cov}\bigl(S(X)_u,S(X)_v\bigr).
\]
Accordingly, averaging introduces a covariance defect in the shuffle identity, and \(\mathbb E[S(X)]\) is group-like if and only if \(X\) is almost surely deterministic, described there as a single tree-like path [2606.28869]. A common misconception is therefore to treat the mean signature as though it retained the exact multiplicativity of each individual realisation; the covariance identity shows that this fails except in the degenerate case.

## 3. Randomised developments and universal large-\(N\) limits

A second major construction starts from random developments of a deterministic path into a matrix Lie group. One chooses a random linear map \(M:V\to \mathrm{Lie}(G_N)\), writes \(M(v)=\sum_i A_i v^i\), and solves the matrix-valued CDE
\[
dZ^N_t = Z^N_t\,M(d\gamma_t),\qquad Z^N_a=I_N,
\]
with \(G_N=GL(N;\mathbb R)\) or \(U(N;\mathbb C)\). This yields feature maps
\[
\mathcal M_G(\gamma;M)=Z_b^N,\qquad \mathcal M_U(\gamma;M)=Z_b^N,
\]
and after averaging over a law \(\xi_N\) on \(M\), kernels
\[
\kappa_N(\gamma,\sigma)
=\mathbb E_{M\sim \xi_N}\bigl[\langle \mathcal M(\gamma;M),\mathcal M(\sigma;M)\rangle_{HS}\bigr]/N
\]
[2402.12311].

For \(GL(N)\), under very mild moment and weak-dependence assumptions on the entries of \(\{A_i^N\}\),
\[
\lim_{N\to\infty}\kappa_N(\gamma,\sigma)=K_{\mathrm{sig}}(\gamma,\sigma),
\]
independently of the precise distribution \(\xi_N\). In this sense, random \(GL(N)\)-based kernels converge to the ordinary signature kernel in the \(N\to\infty\) limit [2402.12311].

In the unitary case, the large-\(N\) limit
\[
K_{SD}(\gamma,\sigma):=
\lim_{N\to\infty}\mathbb E_M[\mathrm{tr}(Z^N_\gamma Z^{N*}_\sigma)]/N
\]
exists. If \(y=\gamma * \mathrm{reverse}(\sigma)\), then
\[
K_{SD}(\gamma,\sigma)
=\sum_{k=0}^\infty (-1)^k \sum_{|I|=k}\phi(X_I)\,S^I(y),
\]
where \(\{X_1,\dots,X_d\}\) are freely-independent semicircular operators in a non-commutative probability space \((\mathcal A,\phi)\) [2402.12311].

The key structural conclusion is that the limiting kernels depend on the choice of Lie group but are universal limits with respect to how the development map is randomised. This places the ordinary signature kernel and the unitary free-probability kernel in a common large-\(N\) framework [2402.12311].

## 4. Free probability, integral equations, and PDE formulations

For unitary developments, the free-semicircle moment functional \(\phi\) satisfies Schwinger–Dyson relations. These translate into a quadratic functional equation for
\[
K(s,t):=K_{SD}(\gamma_\cdot,\sigma_\cdot)(s,t),
\qquad y=\gamma * \mathrm{reverse}(\sigma),
\]
namely
\[
K(s,t)
=1-\int_s^t\int_s^r K(s,u)\,K(u,r)\,\langle dy_u,dy_r\rangle,
\]
with boundary conditions \(K(s,s)=K(t,t)=1\). Setting \(s=0\), \(t=T\) recovers the full kernel \(K_{SD}(\gamma,\sigma)=K(0,T)\) [2402.12311].

This integral-equation viewpoint parallels earlier PDE formulations for signature kernels. In the unweighted case \(\phi(k)\equiv 1\), the two-parameter kernel
\[
u(s,t)=K(\gamma,\sigma)(s,t)
\]
is the unique smooth solution of the Goursat PDE
\[
\frac{\partial^2u}{\partial s\,\partial t}(s,t)
=u(s,t)\,\langle \gamma'_s,\sigma'_t\rangle_V,
\qquad u(s,0)=u(0,t)=1.
\]
The weighted theory also admits integral-transform and randomisation representations: if \(\phi(k)=\mathbb E[\pi^k]\), then
\[
K^\phi(\gamma,\sigma)=\mathbb E_\pi[K^{\pi\gamma,\sigma}],
\]
so computing \(K^\phi\) reduces to averaging PDE solutions for rescaled paths [2107.00447].

For stochastic models, analogous PDEs appear at the level of expected signatures. For inhomogeneous Lévy processes with absolutely continuous characteristics, the expected signature is the free development of the characteristic velocity
\[
\mathfrak y(t)= b(t)+\tfrac12 a(t)
+\int_V\bigl(\exp_\otimes(x)-1-x\,\mathbf 1_{|x|\le 1}\bigr)\,K_t(dx),
\]
and the kernel \(u(s,t)=\langle \eta(s),\widetilde\eta(t)\rangle_{T^2(V)}\) satisfies a hyperbolic Goursat-type PDE coupled to two linear ODEs. In the special case of zero-drift continuous Gaussian martingales with deterministic covariation rates \(a(s)\) and \(\widetilde a(t)\), this closes to
\[
\frac{\partial^2}{\partial s\,\partial t}u(s,t)
=u(s,t)\,\frac14\,\langle a(s),\widetilde a(t)\rangle,
\qquad u(s,0)=1,\quad u(0,t)=1
\]
[2509.07893].

The Brownian case admits further closed-form structure. The expected Stratonovich signature of a \(d\)-dimensional Brownian motion over \([0,s]\) is
\[
\mathbb E[S(\circ B)_{0,s}]
=\exp\!\Bigl(\tfrac s2\,(e_1\otimes e_1+\cdots+e_d\otimes e_d)\Bigr),
\]
and for the factorial weighting \(\phi(k)=\Gamma(\tfrac k2+1)\),
\[
\bigl\langle \mathbb E[S(\circ B)_{0,s}],S(\gamma)_{0,t}\bigr\rangle_\phi
=\cosh\!\bigl(\rho_{\sqrt{s/2}\,\gamma}(t)\bigr),
\]
where \(\rho_{\lambda\gamma}(t)\) is the hyperbolic distance of the Cartan development of the rescaled path in the hyperboloid model [2107.00447].

## 5. Discretisation, dynamic programming, and scalable approximation

For discrete sequences, the signature kernel admits efficient dynamic programming. If \(x(s_1),\dots,x(s_L)\) and \(y(t_1),\dots,y(t_{L'})\) are observed sequences with increments \(\Delta x[i]\) and \(\Delta y[j]\), then the discrete signature is
\[
S(x)=\prod_{i=1}^L (1+\Delta x[i])\in T(H),
\]
and the truncated kernel can be computed by a Horner-type recursion in \(O(LL'M)\) time, instead of the direct \(O(L^M L'^M)\) cost of the double sum over subsequences [1601.08169].

The unitary large-\(N\) kernel from the Schwinger–Dyson equation also admits a direct discretisation that avoids computing full signatures. If \([0,T]=\{t_i\}\) is a partition and \(\gamma^\pi\) is a piecewise-constant or piecewise-linear approximation, then on the grid \(\Pi=\{(t_i,t_j):i\le j\}\) the left-point Riemann–Stieltjes recursion is
\[
K^\pi_{i,j}
=1-\sum_{k=i}^{j-1} K^\pi_{i,k}\,K^\pi_{k,j-1}\,
\langle \Delta \gamma^\pi_{k+1},\Delta \gamma^\pi_j\rangle,
\qquad K^\pi_{i,i}=1.
\]
Because each anti-diagonal depends only on earlier anti-diagonals, the scheme can be filled in \(O(d\,n^2)\) time and parallelised on GPU. Its mesh-size error satisfies
\[
|K(0,T)-K^\pi(0,T)|
\le C\,\|\gamma\|_{1;[0,T]}\exp\!\bigl(C\|\gamma\|_{1;[0,T]}\bigr)\,
\max_i \|\gamma\|_{1;[t_i,t_{i+1}]},
\]
so the discrete solution converges uniformly to the true kernel as the mesh tends to zero [2402.12311].

For large datasets, Random Fourier Signature Features provide an unbiased approximation to truncated signature kernels. With a translation-invariant static kernel \(k\) and independent Random Fourier maps \(\phi_m\), the random kernel \(\hat k_{\le M}(x,y)\) satisfies
\[
\mathbb E[\hat k_{\le M}(x,y)] = k_{\le M}(x,y),
\]
and under Bernstein or sub-Gaussian assumptions one obtains uniform approximation guarantees over bounded-variation sequences [2311.12214]. In complexity terms, exact dynamic-programming kernel evaluation requires \(\mathcal O(N^2L^2M)\) time for a Gram matrix of \(N\) sequences of maximum length \(L\), whereas the feature-based variants reduce the \(L^2\)-dependence to \(\mathcal O(L)\) and replace the \(\mathcal O(N^2)\) Gram-matrix cost by an \(\mathcal O(N)\) primal cost [2311.12214]. For the expected signature kernel on distributions, feature maps \(\Psi(P)\) can be built by averaging pathwise random signature features, yielding \(K(P,Q)\approx \Psi(P)^\top\Psi(Q)\) with uniform high-probability bounds that decay exponentially in the sketch dimension \(D'\) [2506.17634].

## 6. Stochastic models, examples, and uses

For Gaussian processes with strictly regular kernels, the expected signature admits an explicit pairing formula. If
\[
W_t=\int_0^t K(t,r)\,dB_r
\]
has covariance
\[
R(s,t)=\int_0^{s\wedge t} K(s,r)K(t,r)\,dr
=\int_0^s\int_0^t f(u,v)\,du\,dv,
\]
then for any even multi-index \((i_1,\dots,i_{2n})\in E_{2n}\),
\[
\left\langle \mathbb E[S(W)_{0,T}],
e_{i_1}\otimes\cdots\otimes e_{i_{2n}}\right\rangle
=
\sum_{\pi\in\Pi(2n):\,(j,\ell)\in\pi\Rightarrow i_j=i_\ell}
\int_{0<u_1<\cdots<u_{2n}<T}
\prod_{(j,\ell)\in\pi} f(u_j,u_\ell)\,du_1\cdots du_{2n}.
\]
This induces a positive-definite kernel
\[
K(W,\widetilde W)
=\left\langle \mathbb E[S(W)_{0,T}],\,\mathbb E[S(\widetilde W)_{0,T}]\right\rangle
\]
between Gaussian processes [1304.4930].

Fractional Brownian motion with Hurst index \(H>1/2\) is a concrete example. With Volterra kernel \(K_H\), the corresponding density is
\[
f(u,v)=H(2H-1)|u-v|^{2H-2},
\]
and the expected signature coordinates are given by the same pairing formula with this choice of \(f\) [1304.4930].

The expected signature kernel is also used as a covariance on the space of probability measures. Given training distributions \(P_1,\dots,P_N\) and a test distribution \(P^*\), one may place a Gaussian process prior
\[
f\sim GP(0,K),
\]
with posterior mean and covariance determined by the Gram matrix \(K(P_i,P_j)\). Replacing the true expected signatures by empirical averages \(\bar S_i=\hat E[S(X)\mid X\sim P_i]\) yields a GP on finite samples [2506.17634].

Beyond Gaussian settings, the law-space kernel extends to inhomogeneous Lévy processes via the characteristic-velocity PDE system, offering an analytic alternative to Monte Carlo for signature-MMD computation [2509.07893]. For affine and exponential Hawkes processes, truncated expected signatures admit finite-dimensional linear closures after state-weight augmentation, so the corresponding truncated expected signature kernels can be assembled from matrix exponentials [2606.28869].

Taken together, these formulations show that the expected signature kernel occupies two linked roles. It is, first, a positive-definite kernel on laws of paths built from mean signatures; second, it is connected to pathwise signature kernels through randomisation, free probability, PDEs, and numerical schemes that avoid explicit high-order tensor computation [2506.17634] [2402.12311].

Source: https://www.emergentmind.com/topics/expected-signature-kernel