---
title: Reproducing Kernel Hilbert Space
url: https://www.emergentmind.com/topics/reproducing-kernel-hilbert-space
type: topic
---

# Reproducing Kernel Hilbert Space

Reproducing kernel Hilbert space (RKHS) is a Hilbert space of functions on a set \(X\) in which every point-evaluation functional \(f\mapsto f(x)\) is continuous. By the Riesz representation theorem, for each \(x\in X\) there is a unique kernel section \(K(\cdot,x)\) such that
\[
f(x)=\langle f,K(\cdot,x)\rangle_H.
\]
The function \(K:X\times X\to\mathbb C\), called the reproducing kernel, encodes simultaneously the Hilbert-space inner product, the embedding of \(H\) into the function space on \(X\), and the geometry of evaluations. RKHSs are in one-to-one correspondence with positive-semidefinite kernels through the Moore–Aronszajn construction. They provide canonical feature representations, minimum-norm interpolation, kernel-based regularization, covariance and Gaussian-process models, operator-valued observations, geometric embeddings, and function spaces adapted to differential, integral, stochastic, and representation-theoretic structures.

## 1. Definition, reproducing property, and kernel correspondence

Let \(X\) be a set and let \(H\subseteq\mathbb C^X\) be a Hilbert space. For each \(x\in X\), define the evaluation functional
\[
\delta_x:H\to\mathbb C,\qquad \delta_x(f)=f(x).
\]
The space \(H\) is an RKHS if every \(\delta_x\) is continuous. Equivalently, for every \(x\in X\), there is a unique \(K(\cdot,x)\in H\) satisfying
\[
f(x)=\langle f,K(\cdot,x)\rangle_H.
\]
The element \(K(\cdot,x)\) is the Riesz representer of evaluation at \(x\), and the reproducing kernel is
\[
K(x,y)=\langle K(\cdot,y),K(\cdot,x)\rangle_H,
\]
with the order of variables determined by the convention for linearity of the inner product.

The reproducing identity yields the evaluation estimate
\[
|f(x)|\leq \|f\|_H\sqrt{K(x,x)}.
\]
Thus norm convergence in \(H\) implies pointwise convergence. This distinguishes RKHSs from general Hilbert spaces of functions, such as many \(L^2\)-spaces, in which point evaluation need not be continuous [1408.0952].

The kernel is Hermitian and positive semidefinite. For every finite collection \(x_1,\ldots,x_n\in X\) and scalars \(c_1,\ldots,c_n\),
\[
\sum_{i,j=1}^n c_i\overline{c_j}K(x_i,x_j)\geq 0.
\]
Indeed,
\[
\sum_{i,j=1}^n c_i\overline{c_j}K(x_i,x_j)
=
\left\|\sum_{i=1}^n c_iK(\cdot,x_i)\right\|_H^2.
\]

Conversely, every positive-semidefinite kernel determines a unique RKHS. One begins with
\[
H_0(K)=\operatorname{span}\{K(\cdot,x):x\in X\},
\]
defines
\[
\left\langle \sum_i a_iK(\cdot,x_i),\sum_j b_jK(\cdot,y_j)\right\rangle
=
\sum_{i,j}a_i\overline{b_j}K(y_j,x_i),
\]
and completes the resulting pre-Hilbert space. This is the Moore–Aronszajn construction. The kernel sections form a canonical, generally overcomplete coordinate system:
\[
H=\overline{\operatorname{span}\{K(\cdot,x):x\in X\}}.
\]

A positive-semidefinite kernel also admits a feature representation. There are a Hilbert space \(W\) and a map \(\Phi:X\to W\) such that
\[
K(x,y)=\langle \Phi(y),\Phi(x)\rangle_W.
\]
The associated RKHS consists of functions
\[
f(x)=\langle h,\Phi(x)\rangle_W
\]
with the quotient norm induced by the feature representation. In the canonical realization, \(\Phi(x)=K(\cdot,x)\).

## 2. Feature geometry, distances, and regularity

The canonical feature map induces the kernel pseudometric
\[
d_K(x,y)=\|K(\cdot,x)-K(\cdot,y)\|_H,
\]
whose squared form is
\[
d_K^2(x,y)
=
K(x,x)+K(y,y)-2\operatorname{Re}K(x,y).
\]
For a real symmetric kernel,
\[
d_K^2(x,y)=K(x,x)+K(y,y)-2K(x,y).
\]
It is a genuine metric when the feature map is injective; otherwise distinct points with identical kernel sections are identified.

Every \(f\in H\) is automatically Lipschitz with respect to this intrinsic metric:
\[
|f(x)-f(y)|
=
|\langle f,K(\cdot,x)-K(\cdot,y)\rangle_H|
\leq
\|f\|_H d_K(x,y).
\]
This property does not require an externally specified topology on \(X\). It is a direct consequence of the Hilbert-space geometry of kernel sections [2310.18078].

The intrinsic metric differs from regularity measured in a pre-existing metric \(d\) on \(X\). If kernel sections are locally \(\alpha\)-Hölder with respect to \(d\), then, generally,
\[
d_K(x,y)\lesssim d(x,y)^{\alpha/2},
\]
and every RKHS function is \(\alpha/2\)-Hölder with a bound proportional to its RKHS norm. In particular, ordinary Lipschitz continuity of kernel sections generally yields only \(1/2\)-Hölder continuity of all RKHS functions. A feature map that is itself \(\alpha\)-Hölder in Hilbert-space norm avoids this exponent-halving phenomenon and gives
\[
|f(x)-f(y)|\leq L_\Phi\|f\|_H d(x,y)^\alpha.
\]

If the kernel is bounded,
\[
\sup_{x\in X}K(x,x)<\infty,
\]
then every RKHS function is bounded:
\[
\sup_{x\in X}|f(x)|
\leq
\|f\|_H\sup_{x\in X}\sqrt{K(x,x)}.
\]
Bounded kernels therefore produce RKHSs continuously embedded into \(\ell_\infty(X)\).

The feature-space distance is distinct from the projective distance obtained by normalizing kernel sections. The ordinary feature-space distance is
\[
d_H(x,y)=\|K(\cdot,x)-K(\cdot,y)\|_H.
\]
When \(K(\cdot,x)\neq0\), define
\[
\widehat K_x=\frac{K(\cdot,x)}{\|K(\cdot,x)\|_H}.
\]
The projective sine distance is
\[
\delta_H(x,y)
=
\sqrt{1-\left|\langle \widehat K_x,\widehat K_y\rangle_H\right|^2}
=
\sqrt{1-\frac{|K(x,y)|^2}{K(x,x)K(y,y)}}.
\]
It measures the angle between normalized kernel sections and is invariant under nonzero rescaling of those sections. It also equals the operator norm distance between the rank-one projections \(P_x\) and \(P_y\):
\[
\delta_H(x,y)=\|P_x-P_y\|.
\]
Related Skwarczyński and Fubini–Study distances have the same infinitesimal geometry, although their global formulas differ [1010.0136].

## 3. Interpolation, approximation, and inclusion relations

Given distinct points \(x_1,\ldots,x_N\), the finite-dimensional kernel space is
\[
\mathcal S_X=\operatorname{span}\{K(\cdot,x_j):1\leq j\leq N\}.
\]
An interpolant
\[
s(x)=\sum_{j=1}^N c_jK(x,x_j)
\]
satisfying \(s(x_i)=f(x_i)\) is obtained from the Gram system
\[
A_Xc=f_X,\qquad
(A_X)_{ij}=K(x_i,x_j).
\]
If the kernel is strictly positive definite, \(A_X\) is nonsingular and the interpolant is unique.

The cardinal functions \(h_1,\ldots,h_N\in\mathcal S_X\) satisfy
\[
h_j(x_i)=\delta_{ij},
\]
so the interpolant is
\[
I_Xf(x)=\sum_{j=1}^N f(x_j)h_j(x).
\]
The orthogonal decomposition
\[
H=\mathcal S_X\oplus\mathcal S_X^\perp
\]
has
\[
\mathcal S_X^\perp=\{f\in H:f(x_i)=0,\ i=1,\ldots,N\}.
\]
Consequently, the kernel interpolant is the orthogonal projection of \(f\) onto \(\mathcal S_X\) and preserves the prescribed nodal values [1705.01364].

The kernel also determines minimum-norm interpolation. For constraints \(f(x_i)=c_i\), the minimum-norm solution belongs to \(\mathcal S_X\). For a single constraint \(f(x_0)=c\),
\[
f=\frac{c}{K(x_0,x_0)}K(\cdot,x_0),
\]
provided \(K(x_0,x_0)>0\).

Inclusion of RKHSs is governed by kernel domination rather than pointwise comparison. For positive-semidefinite kernels \(K\) and \(G\),
\[
\mathcal H_K\subseteq\mathcal H_G
\quad\Longleftrightarrow\quad
\exists\,\lambda>0\text{ such that }K\preceq\lambda G,
\]
where \(G-K\) is positive semidefinite. The optimal constant
\[
\lambda(K,G)=\inf\{\lambda>0:K\preceq\lambda G\}
\]
determines the norm of the inclusion:
\[
\|f\|_{\mathcal H_G}
\leq
\sqrt{\lambda(K,G)}\,\|f\|_{\mathcal H_K}.
\]

Equivalently, if
\[
K(x,y)=\langle\Phi_1(x),\Phi_1(y)\rangle_{W_1},
\qquad
G(x,y)=\langle\Phi_2(x),\Phi_2(y)\rangle_{W_2},
\]
then
\[
\mathcal H_K\subseteq\mathcal H_G
\]
if and only if there exists a bounded operator \(T:W_2\to W_1\) such that
\[
T\Phi_2(x)=\Phi_1(x).
\]
Thus the feature representation of the smaller RKHS must be obtainable from that of the larger one by a bounded linear transformation [1106.4075].

For translation-invariant kernels on \(\mathbb R^d\), Bochner’s theorem gives
\[
K(x,y)=\int_{\mathbb R^d}e^{i(x-y)\cdot\xi}\,d\mu(\xi).
\]
Then
\[
\mathcal H_K\subseteq\mathcal H_G
\]
if and only if
\[
\mu\ll\nu,
\qquad
\frac{d\mu}{d\nu}\in L^\infty(\nu),
\]
where \(\nu\) is the representing measure of \(G\). The optimal inclusion constant is
\[
\lambda(K,G)=\left\|\frac{d\mu}{d\nu}\right\|_{L^\infty}.
\]
For Hilbert–Schmidt kernels, the analogous condition is coefficientwise domination:
\[
\mathcal H_{K_a}\subseteq\mathcal H_{K_b}
\quad\Longleftrightarrow\quad
a_n\leq \lambda b_n
\]
for all \(n\). Equivalent norms require two-sided domination rather than one-sided inclusion.

The inclusion order is stable under several kernel operations. Sums, positive scalar multiplication, products through the Schur product theorem, uniformly bounded pointwise limits, and analytic functions with nonnegative power-series coefficients preserve appropriate forms of inclusion.

## 4. Analytic, geometric, stochastic, and operator-valued RKHSs

RKHSs arise naturally from differential operators and geometric structures. On a complete \(n\)-dimensional Riemannian manifold of bounded geometry, the Sobolev space \(H^s(M)\) is an RKHS when
\[
s>\frac n2.
\]
Its kernel is the integral kernel of the elliptic smoothing operator
\[
(I+\Delta)^{-s},
\]
where \(\Delta\) is the Laplace–Beltrami operator:
\[
J_sJ_s^*=(I+\Delta)^{-s}=L_{K_s}.
\]
On a compact manifold with Laplacian eigenfunctions \(\{f_k\}\) and eigenvalues \(\{\lambda_k\}\),
\[
K_s(m,m')
=
\sum_k(1+\lambda_k)^{-s}f_k(m)f_k(m').
\]
Diffusion spaces
\[
\mathcal H^t=e^{-t\Delta/2}L^2(M)
\]
are RKHSs for every \(t>0\), with reproducing kernel equal to the heat kernel
\[
K_{\mathcal H^t}(m,m')=p(m,m',t).
\]
They impose exponential spectral decay and are smoother than every fixed-order Sobolev space [1905.10913].

Reachable spaces of boundary-controlled heat equations also possess RKHS structures. If a state is represented as
\[
w(z,T)=(u,h_z)_{L^2(0,T)},
\]
then its range, equipped with the quotient norm, is an RKHS with kernel
\[
K(z,w)=(h_w,h_z)_{L^2(0,T)}.
\]
For a finite rod, analytically continued reachable states form RKHSs on squares; for the half-line, the reachable space is a sum of Gaussian-weighted pullbacks of Bergman and Hardy spaces on the right half-plane [1910.03765].

The operator reproducing kernel Hilbert space (ORKHS) generalizes point evaluation. Let \(H\) and \(Y\) be Hilbert spaces and let
\[
L_a:H\to Y
\]
be a family of continuous linear operators. An ORKHS is a Hilbert space for which the operators \(L_a\) are continuous and separate elements of \(H\). Its operator reproducing kernel is
\[
K(a)=L_a^*\in B(Y,H),
\]
and
\[
(L_a(f),\xi)_Y=(f,K(a)\xi)_H.
\]
The induced operator-valued kernel is
\[
\mathcal K(a,b)\xi=L_b(K(a)\xi).
\]
This framework includes Fourier coefficients, wavelet coefficients, integral measurements, local averages, vector-valued observations, and operator-valued data. Perfect ORKHSs reproduce both ordinary point evaluations and another family of operators. Regularized learning from operator-valued observations has minimizers of the form
\[
f_0=\sum_{j=1}^mK(a_j)\eta_j,
\]
an operator-valued representer theorem [1512.05923].

Positive-definite kernels are also covariance kernels. Every such kernel is the covariance kernel of a centered Gaussian process:
\[
K(x,y)=\mathbb E[\overline{W_x}W_y].
\]
The associated RKHS is the deterministic Hilbert space generated by the same covariance structure. Kernel transforms map measures or bounded linear functionals into RKHS elements:
\[
(T_K\mu)(t)=\int K(s,t)\,d\mu(s),
\]
with
\[
\|T_K\mu\|_{\mathcal H(K)}^2
=
\iint K(s,t)\,d\mu(s)\,d\mu(t).
\]
This identifies RKHS geometry with moment, measure, and stochastic-process geometry [2209.03801].

Group and groupoid representations produce kernels through matrix coefficients. For a unitary representation \(\mathcal U\) of a group \(G\) and \(v\) in the representation space,
\[
K(h,g)=\langle\mathcal U(g)v,\mathcal U(h)v\rangle.
\]
Conversely, a positive-definite kernel satisfying the appropriate covariance identity generates a unitary representation by translating kernel sections. The groupoid version decomposes the RKHS over range fibers and connects kernel geometry with convolution, harmonic analysis, complex transformations, and representation theory [2102.09585].

## 5. Kernel methods, distributions, and mean-field limits

In machine learning, a datum \(x\) is mapped to its kernel section
\[
x\longmapsto K(\cdot,x).
\]
Linear functionals in the RKHS correspond to nonlinear functions of the original input. Minimum-norm interpolation, kernel ridge regression, support-vector methods, and regularized empirical risk minimization exploit this feature-space representation.

For a loss depending only on values at training points, the representer theorem reduces an infinite-dimensional optimization problem to the finite span of training sections:
\[
\widehat f(\cdot)=\sum_{j=1}^n\widehat\alpha_jK(\cdot,x_j).
\]
For squared loss with RKHS penalty,
\[
\widehat\alpha=(C+\lambda I)^{-1}Y,
\]
where \(C\) is the Gram matrix.

Probability measures can themselves serve as inputs. For a probability measure \(\mu\), the kernel mean element is
\[
m_\mu=\int K(\cdot,x)\,d\mu(x).
\]
A kernel is characteristic if \(\mu\mapsto m_\mu\) is injective. In one-dimensional distribution regression, the quadratic Wasserstein distance is the \(L^2\)-distance between quantile functions:
\[
W_2^2(\mu,\nu)
=
\int_0^1\left(F_\mu^{-1}(t)-F_\nu^{-1}(t)\right)^2dt.
\]
The kernel
\[
k_\Theta(\mu,\nu)
=
\gamma^2\exp\left(-\frac{W_2^{2H}(\mu,\nu)}{l}\right),
\qquad 0<H\leq1,
\]
is positive definite. For \(H=1\), under compactness and regularity assumptions, it is universal on \(\mathcal W_2(\Omega)\), allowing uniform approximation of continuous functions of probability distributions [1806.10493].

Mean-field limits extend this construction to systems with many measurement variables. A symmetric configuration
\[
\vec x=(x_1,\ldots,x_M)\in X^M
\]
is represented by its empirical measure
\[
[\vec x]=\frac1M\sum_{i=1}^M\delta_{x_i}.
\]
If a sequence of kernels \(k^{[M]}\) is permutation invariant, uniformly bounded, and equicontinuous through empirical measures, then a subsequence converges uniformly to a continuous positive-definite kernel
\[
k:\mathcal P(X)\times\mathcal P(X)\to\mathbb R.
\]
The associated finite-\(M\) RKHS functions have subsequences converging uniformly, on empirical configurations, to functions in the limiting RKHS. Double-sum kernels provide an important example:
\[
k^{[M]}(\vec x,\vec x')
=
\frac1{M^2}\sum_{m,m'=1}^M k_0(x_m,x_{m'}'),
\]
whose limit is
\[
k(\mu,\mu')
=
\iint k_0(x,x')\,d\mu(x)\,d\mu'(x').
\]
The convergence is subsequential and function-wise; it does not assert convergence of the entire sequence of RKHSs in operator norm [2302.14446].

RKHSs also arise in meshless collocation. Cardinal functions constructed from kernel sections yield differentiation matrices
\[
L_{ij}=\mathcal L h_j(x_i),
\]
and pointwise differentiation errors satisfy
\[
\left|
\mathcal Lu(z)-\sum_k u(x_k)\mathcal Lh_k(z)
\right|
\leq
\|u\|_H\|\varepsilon_X(z)\|_H,
\]
where \(\varepsilon_X(z)\) is a kernel-determined interpolation error. Sobolev RKHS trial spaces can incorporate homogeneous boundary conditions directly, producing meshless collocation schemes for boundary-value, parabolic, and Burgers equations [1705.01364].

## 6. Algebraic structure, stability, and limitations

Pointwise multiplication does not generally preserve an RKHS. A reproducing kernel Hilbert algebra (RKHA) is an RKHS for which the diagonal map
\[
\Delta(k_x)=k_x\otimes k_x
\]
extends boundedly from \(H\) to \(H\otimes H\). Its adjoint
\[
m=\Delta^*
\]
is pointwise multiplication:
\[
m(f\otimes g)=fg.
\]
Hence
\[
\|fg\|_H\leq \|\Delta\|_{\mathrm{op}}\|f\|_H\|g\|_H.
\]
The comultiplication is coassociative and cocommutative, and the RKHA structure is closed under Hilbert-space tensor products and pullbacks [2401.01295].

For applications to system identification, an RKHS of impulse responses is called stable when every element is absolutely integrable:
\[
H\subseteq L_1(\mathbb R_+)
\]
in continuous time, or
\[
H\subseteq\ell_1(\mathbb N)
\]
in discrete time. If \(L_K\) is the kernel-induced operator
\[
L_K[u](t)=\int_0^\infty K(t,\tau)u(\tau)\,d\tau,
\]
then stability is equivalent to boundedness
\[
L_K:L_\infty\to L_1.
\]
For continuous Mercer kernels under the stated integrability assumptions, this can be tested using only sign-valued functions \(u\) satisfying \(|u|=1\) almost everywhere:
\[
\sup_{\|u\|_\infty\leq1}\|L_K[u]\|_1
=
\sup_{|u|=1\ \mathrm{a.e.}}\|L_K[u]\|_1.
\]
The discrete-time analogue holds for the entire class of kernels as a boundedness test for \(K:\ell_\infty\to\ell_1\) [2305.02213].

The existence of an RKHS between two Banach function spaces is constrained by Hilbertian geometry. For proper Banach function spaces \(E\subset F\), there exists an intermediate RKHS
\[
E\subset H\subset F
\]
if and only if the inclusion operator \(E\to F\) factors through a Hilbert space. Such 2-factorability implies type \(2\) and cotype \(2\) conditions. For classical smoothness spaces on a \(d\)-dimensional domain, the resulting thresholds involve dimension-dependent smoothness gaps. For example, an RKHS with bounded kernel containing a Sobolev-type space of smoothness \(s\) requires, under the stated hypotheses,
\[
s\geq \frac d p+\frac d2.
\]
This obstruction reflects the abundance of localized, nearly independent functions in high-dimensional domains rather than pointwise smoothness alone [2312.14711].

Finally, RKHS inclusion and metric properties have important degeneracies. Kernel-induced distances may be pseudometrics if kernel sections fail to separate points. A zero kernel section makes the corresponding point invisible to the RKHS. Kernel inclusion is not determined by pointwise kernel magnitude, and equivalent norms are stronger than set inclusion. Regularity of a kernel in an external metric does not automatically give the same regularity for all RKHS functions without exponent loss or additional feature-map structure. Likewise, mean-field convergence generally requires compactness, symmetry, uniform boundedness, equicontinuity, and often passage to a subsequence. These qualifications are intrinsic to the theory: an RKHS is determined not merely by a similarity formula, but by the complete interaction among positive definiteness, Hilbert geometry, evaluation continuity, feature representations, and the function-space topology.

Source: https://www.emergentmind.com/topics/reproducing-kernel-hilbert-space