---
title: Radon–Nikodym Theorem for Kernels
url: https://www.emergentmind.com/topics/radon-nikodym-theorem-for-kernels
type: topic
---

# Radon–Nikodym Theorem for Kernels

The Radon–Nikodym theorem for kernels is a family of representation results that extend the classical measure-theoretic formula $d\mu = f\, d\nu$ to settings where the basic positive object is a kernel rather than a measure. In current usage, the term *kernel* covers at least three distinct structures: measurable kernels $\kappa:X\rightsquigarrow Y$ in probability and semantics, scalar or operator-valued positive definite kernels $K:X\times X\to \mathcal L(H)$ in RKHS and dilation theory, and completely positive matrix-valued kernels arising from Stinespring- and KSGNS-type constructions. Across these settings, domination is converted into a derivative: either a measurable density $d\kappa/d\lambda$, or a positive contraction acting on a Kolmogorov, reproducing-kernel, or Stinespring space [1810.01837] [2509.20496] [2405.09315].

## 1. Terminological scope and basic structures

A first technical point is that the phrase *kernel* is genuinely overloaded. In the measure-theoretic literature, a kernel from $(X,\Sigma_X)$ to $(Y,\Sigma_Y)$ is a map
\[
\kappa:X\times \Sigma_Y \to [0,\infty]
\]
such that $A\mapsto \kappa(x,A)$ is a measure for each $x$, and $x\mapsto \kappa(x,A)$ is measurable for each $A$ [1810.01837]. In the RKHS and operator-theoretic literature, a kernel is instead a function on a Cartesian square. A scalar kernel $K:X\times X\to \mathbb C$ is positive definite if
\[
\sum_{i,j} c_i \overline{c_j}\, K(x_i,x_j)\ge 0,
\]
while an operator-valued kernel $K:X\times X\to \mathcal L(H)$ is positive definite if
\[
\sum_{i,j}\langle u_i, K(x_i,x_j)u_j\rangle_H \ge 0
\]
for all finite choices of points and vectors [2509.20496].

For positive definite kernels, the basic structural tool is the Kolmogorov–Aronszajn decomposition. In the scalar case one has a Hilbert space $\mathcal H_K$ and feature map $\xi_K:X\to \mathcal H_K$ such that
\[
K(x,y)=\langle \xi_K(x),\xi_K(y)\rangle_{\mathcal H_K},
\]
with dense span. In the operator-valued case one has
\[
K(s,t)=V_K(s)^*V_K(t),\qquad V_K(t):H\to \mathcal K_K,
\]
again with a minimality condition given by density of $\operatorname{span}\{V_K(t)h\}$ [2509.20496]. The same construction appears in a canonical scalarized form in operator-valued kernel theory: if
\[
\tilde K((s,a),(t,b)):=\langle a, K(s,t)b\rangle_H,
\]
then $\tilde K$ is a scalar positive definite kernel with RKHS $H_{\tilde K}$, and the feature map is $V(s)a=\tilde K_{(s,a)}$ [2405.09315].

The relevant order relation also depends on context. For measurable kernels, domination is absolute continuity:
\[
\kappa \ll \lambda \iff \lambda(x,A)=0 \Rightarrow \kappa(x,A)=0
\]
for all $x$ and measurable $A$ [1810.01837]. For positive definite kernels, domination is the Loewner-type order
\[
K\preceq L \iff L-K \text{ is positive definite},
\]
equivalently, every finite Gram block for $L-K$ is positive semidefinite [2509.20496]. A persistent source of confusion is to treat these notions as interchangeable; they are not. They live on different categories of kernels and produce different kinds of Radon–Nikodym derivatives.

## 2. Positive definite kernels and operator-valued densities

For positive definite kernels, the core Radon–Nikodym statement is an operator factorization theorem. If $K,L:X\times X\to \mathcal L(H)$ are positive definite and $K\preceq L$, and if
\[
L(s,t)=V_L(s)^*V_L(t)
\]
is a minimal Kolmogorov decomposition, then there exists a unique operator
\[
A\in \mathcal L(\mathcal K_L),\qquad 0\le A\le I,
\]
such that
\[
K(s,t)=V_L(s)^* A\, V_L(t),\qquad \forall s,t\in X.
\]
This is the kernel Radon–Nikodym theorem in the form stated in Theorem 3.1 of "Kernel Radon-Nikodym Derivatives for Random Matrix Ensembles" [2509.20496]. The same functional-analytic pattern appears in the general operator-valued kernel theorem of "Operator-Valued Kernels, Machine Learning, and Dynamical Systems," where the derivative is denoted $T$ and acts on the scalarized feature space $H_{\tilde L}$:
\[
K(s,t)=V_L(s)^*\, T\, V_L(t),\qquad 0\le T\le I,
\]
with the equivalent feature-map relation
\[
V_K(t)=T^{1/2}V_L(t)
\]
[2405.09315].

The factorization admits an equivalent contraction form. If
\[
K(s,t)=\widetilde V_K(s)^*\widetilde V_K(t)
\]
is any minimal decomposition of $K$, then there is a unique contraction
\[
W:\mathcal K_L\to \mathcal K_K
\]
such that
\[
\widetilde V_K(\cdot)=W V_L(\cdot),\qquad A=W^*W.
\]
This shows that the derivative is not a pointwise multiplier on $X$, but an operator on the feature space of the dominating kernel [2509.20496]. In the scalar case, this reduces to the familiar Aronszajn inclusion principle: $K\preceq L$ if and only if $\mathcal H_K$ embeds contractively into $\mathcal H_L$, and the derivative is the operator implementing that inclusion [2405.09315].

A further structural refinement appears when the kernel is equivariant under a representation of a $C^*$-algebra. If a nondegenerate representation $\rho:\mathcal B\to \mathcal L(\mathcal K_L)$ satisfies
\[
V_L(t)(bh)=\rho(b)V_L(t)h,
\]
and similarly for $K$, then the derivative lies in the commutant:
\[
A\in \rho(\mathcal B)'.
\]
This aligns the kernel theorem with Radon–Nikodym theorems for completely positive maps and Stinespring dilations [2509.20496]. A common misconception is that the kernel derivative should be a scalar density analogous to $d\mu/d\nu$; in the positive-definite setting, the derivative is typically operator-valued because the order relation is encoded in feature-space geometry rather than pointwise integration.

## 3. Shifted kernels, moment-ratio tests, and random matrix ensembles

A distinctive development of the positive-definite theory is the shifted-kernel formalism on the free semigroup $\mathbb F_d^+$. If
\[
K_{\mathcal B}:\mathbb F_d^+\times \mathbb F_d^+\to \mathcal L(H)
\]
is positive definite, its shift is defined by
\[
(K_{\mathcal B})_\Sigma(\alpha,\beta):=\sum_{i=1}^d K_{\mathcal B}(\alpha i,\beta i).
\]
Given a minimal Kolmogorov decomposition
\[
K_{\mathcal B}(\alpha,\beta)=V(\alpha)^*V(\beta),
\]
one defines abstract shift operators on the dense span $\mathcal D=\operatorname{span}\{V(\alpha)u\}$ by
\[
B_i\,V(\alpha)u:=V(\alpha i)u,\qquad i=1,\dots,d.
\]
The Radon–Nikodym derivative of the shifted kernel is then
\[
\frac{d (K_{\mathcal B})_\Sigma}{d K_{\mathcal B}}=\sum_{i=1}^d B_i^*B_i,
\]
and
\[
(K_{\mathcal B})_\Sigma \preceq K_{\mathcal B}\quad \Longleftrightarrow\quad \sum_{i=1}^d B_i^*B_i\le I.
\]
This is the content of Lemma 4.1 and Theorem 4.2 in [2509.20496].

In the one-variable bi-unitarily invariant case, the derivative becomes diagonal and the domination criterion reduces to a moment-ratio test. If
\[
K(m,n)=\mathbb E[A^m(A^n)^*]=\delta_{mn}\, c_m\, I_N,
\]
where
\[
c_m:=\frac1N \operatorname{Tr}\,\mathbb E\big[A^m(A^m)^*\big],
\]
then the shift satisfies
\[
\frac{dK_\Sigma}{dK}=B^*B=\sum_{m\ge 0}\frac{c_{m+1}}{c_m}\, |e_m\rangle\langle e_m|,
\]
and hence
\[
K_\Sigma\preceq K \quad \Longleftrightarrow\quad \frac{c_{m+1}}{c_m}\le 1 \quad \forall m\ge 0
\]
[2509.20496]. For Haar unitary matrices, $c_m\equiv 1$, so $K_\Sigma=K$. For complex Ginibre ensembles with variance $\sigma_N^2=\tau/N$, one has $c_m^{(N)}\to \tau^m$, and therefore
\[
\frac{c_{m+1}^{(N)}}{c_m^{(N)}}\to \tau,
\]
so the shifted-kernel domination criterion is asymptotically equivalent to $\tau\le 1$ [2509.20496].

The same mechanism extends blockwise. If $H=\bigoplus_{r=1}^R \mathbb C^{k_r}$ and block invariance yields
\[
K_{\mathcal B}(m,n)=\delta_{mn}\,\bigoplus_{r=1}^R c_m^{(r)} I_{k_r},
\]
then
\[
\frac{d (K_{\mathcal B})_\Sigma}{d K_{\mathcal B}}
=
\bigoplus_{r=1}^R \sum_{m\ge 0}\frac{c_{m+1}^{(r)}}{c_m^{(r)}}\, P_m^{(r)},
\]
and domination holds if and only if
\[
\frac{c_{m+1}^{(r)}}{c_m^{(r)}}\le 1 \qquad \forall m,\ \forall r
\]
[2509.20496]. The paper also states mean-square von Neumann inequalities under shifted domination:
\[
E_{\mathcal B}\!\left(\mathbb E[f(A)f(A)^*]\right)\le \|f(L)\|^2 I_H,
\]
and for $Y\ge 0$ in $\mathcal B$,
\[
E_{\mathcal B}\!\big(Y^{1/2}\mathbb E[f(A)f(A)^*]Y^{1/2}\big)\le \|f(L)\|^2 Y.
\]
These formulas connect kernel domination to operator inequalities for noncommutative polynomials [2509.20496].

## 4. Measurable kernels, s-finiteness, and strengthened absolute continuity

For measurable kernels, the Radon–Nikodym theorem takes the classical integral form. If $\kappa,\lambda:X\rightsquigarrow Y$ are $\sigma$-finite kernels and $\kappa\ll\lambda$, then there exists a density
\[
f:X\times Y\to [0,\infty)
\]
such that
\[
\kappa(x,A)=\int_A f(x,y)\,\lambda(x,dy),
\]
with uniqueness $\lambda(x)$-a.e. for each $x$ [1810.01837]. Under the mild structural assumptions that $X$ is countable discrete or $Y$ is standard Borel, the density can be chosen jointly measurable in $(x,y)$ [1810.01837].

The s-finite extension is subtler. A kernel is s-finite if it is a countable sum of finite kernels, with no mutual singularity requirement. The paper "On S-Finite Measures and Kernels" shows that for s-finite kernels $\nu,\nu':X\rightsquigarrow Y$, ordinary absolute continuity $\nu'\ll\nu$ is not sufficient for a Radon–Nikodym density. One needs a strengthened relation, denoted in the paper by $\nu' \,\nu$, meaning $\nu'\ll\nu$ and preservation of $\nu(x)$-0–$\infty$ sets for each $x$ [1810.01837]. Under this hypothesis there exists
\[
\frac{d\nu'}{d\nu}:X\times Y\to [0,\infty]
\]
with
\[
\nu'(x,A)=\int_A \frac{d\nu'}{d\nu}(x,y)\,\nu(x,dy).
\]

The strengthened condition is tied to the notion of a top $0$–$\infty$ set $\infty[\mu]$ for an s-finite measure $\mu$. For s-finite kernels, under the same countability or standard-Borel assumptions, there is a measurable set $\infty[k]\subseteq X\times Y$ whose vertical sections are the pointwise top $0$–$\infty$ sets, and $k$ restricted to the complement of $\infty[k]$ is $\sigma$-finite [1810.01837]. Uniqueness of densities is correspondingly refined: derivatives are “a.e. $\infty$-unique,” meaning standard a.e. uniqueness on the $\sigma$-finite part and a weaker zero/nonzero agreement condition on the infinite part.

A common misconception is that the classical hypothesis $\nu'\ll\nu$ should remain sufficient in the s-finite world. The counterexample in [1810.01837] shows otherwise: if
\[
\nu=\sum_{n\in\mathbb N} U_{\mathbb R},\qquad \nu'=\operatorname{Normal}(0,1),
\]
then $\nu'\ll\nu$ holds, but no measurable density $f$ can satisfy $\nu'=\#_1(\nu)\{f\}$. The obstruction is precisely the mismatch of $0$–$\infty$ structure.

The same paper also proves an s-finite Lebesgue decomposition for kernels:
\[
k = k_a + k_\infty + k_s,
\]
where $k_a$ is absolutely continuous in the strengthened sense, $k_\infty$ is $\infty$-singular but still dominated, and $k_s$ is singular [1810.01837]. This decomposition has no analogue in the classical $\sigma$-finite setting, where the $\infty$-singular term vanishes.

## 5. RKHS reformulations and Lebesgue decomposition

A different kernel-theoretic route to Radon–Nikodym theory is to encode measure relations inside reproducing-kernel Hilbert spaces. For finite positive regular Borel measures $\mu,\nu$ on the unit circle $\mathbb T$, "A reproducing kernel approach to Lebesgue decomposition" constructs the RKHS $H_+(\mu)$ of $\mu$-Cauchy transforms on the unit disk $\mathbb D$, with reproducing kernel
\[
K_\mu(z,w)
=
\int_{\mathbb T}
\frac{1}{1-z\overline\zeta}\,
\frac{1}{1-\overline w\,\zeta}\,
d\mu(\zeta)
=
\frac{H_\mu(z)+H_\mu(w)}{2(1-z\overline w)}
\]
[2312.01961]. Here $H_\mu$ is the Herglotz–Riesz transform of $\mu$.

In this framework, domination is equivalent to RKHS containment. Theorem 3 of [2312.01961] states that
\[
\mu \le t^2 \nu
\quad\Longleftrightarrow\quad
K_\mu \preceq t^2 K_\nu,
\]
equivalently $H_+(\mu)$ embeds in $H_+(\nu)$ with norm at most $t$. Absolute continuity is likewise characterized by intersection density: if
\[
\operatorname{int}(\mu,\nu):=H_+(\mu)\cap H_+(\nu),
\]
then
\[
\mu\ll \nu
\quad\Longleftrightarrow\quad
\operatorname{int}(\mu,\nu)\text{ is norm-dense in }H_+(\mu)
\]
(Theorem 6) [2312.01961].

The Radon–Nikodym derivative appears here as a Toeplitz symbol. When $\mu=f\,\nu$ with $f\ge 0$, the kernel identity is
\[
K_\mu(z,w)=\int_{\mathbb T} k_z(\zeta)\,\overline{k_w(\zeta)}\, f(\zeta)\, d\nu(\zeta),
\qquad
k_z(\zeta)=\frac{1}{1-z\overline\zeta}.
\]
The associated operator on $H^2(\nu)$ is
\[
T_\mu = E_{\mu,\nu}^*E_{\mu,\nu}
      = P_{H^2(\nu)}\, M_f\,\big|_{H^2(\nu)},
\]
so the derivative is realized as the multiplication symbol of a Toeplitz operator implementing the inclusion [2312.01961]. This is structurally parallel to the operator-density formulas $K=V_L^*AV_L$ of the abstract positive-definite theory.

The same paper gives a kernel-theoretic construction of Lebesgue decomposition via positive quadratic forms. A significant limitation is explicitly identified: the Simon-Lebesgue decomposition of forms need not coincide with the measure-theoretic decomposition unless the intersection space is invariant, equivalently $V_\mu$-reducing. Example 3 in [2312.01961] shows that one may have $\mu_a=0$ relative to $\nu$ while the form-theoretic absolutely continuous part is nonzero. This is an important caution against assuming that every RKHS decomposition automatically matches measure decomposition without an invariance hypothesis.

## 6. Noncommutative extensions, stochastic realizations, and computation

The operator-density form of the kernel Radon–Nikodym theorem has several extensions. Every operator-valued positive definite kernel $K:S\times S\to \mathcal L(H)$ admits a Gaussian realization:
\[
K(s,t)=\int_\Omega |W(s)\rangle \langle W(t)|\, d\mathbb P,
\]
and, if $(\varphi_i)$ is an orthonormal basis of the feature space, one may take
\[
W(t)=\sum_i \big(V(t)^*\varphi_i\big)\, Z_i,
\]
where $(Z_i)$ are i.i.d. standard Gaussians [2405.09315]. If $K\le L$ with Radon–Nikodym derivative $T$, then the process with covariance $K$ is obtained from the process for $L$ by inserting $T^{1/2}$ in feature space:
\[
W_K(t)=\sum_i \big(V_L(t)^*T^{1/2}\varphi_i\big)\, Z_i.
\]
This realizes domination of kernels as covariance compression [2405.09315].

In noncommutative probability, the same pattern becomes a Radon–Nikodym theorem for completely positive maps. For a unital CP map $\psi:\mathfrak A\to \mathcal L(H)$ one defines
\[
K(A,B):=\psi(A^*B),
\]
builds the minimal Stinespring representation $(\pi_\psi,V_\psi,\mathcal K_\psi)$, and then, for another CP map $\varphi\le \psi$, there exists a unique positive contraction
\[
T\in \pi_\psi(\mathfrak A)'
\]
such that
\[
\varphi(A)=V_\psi^*\, T^{1/2}\, \pi_\psi(A)\, T^{1/2}\, V_\psi
\]
[2405.09315]. The finite-index Hilbert-module analogue replaces a single CP map by a completely positive $n\times n$ matrix of maps and proves that domination corresponds bijectively to a positive contraction in the commutant of the matrix KSGNS dilation; this is the content of the order-isomorphism theorem in [1608.01672].

Computation in the positive-definite setting proceeds through finite Gram data. Given sample points $\{x_i\}$, one forms block Gram matrices
\[
G_K=[K(x_i,x_j)],\qquad G_L=[L(x_i,x_j)],
\]
tests $K\preceq L$ by checking
\[
G_L-G_K\ge 0
\]
in the Loewner sense, constructs feature maps by Cholesky or spectral factorization, and then solves for a contraction $W$ mapping the $L$-features to the $K$-features, finally setting
\[
D=W^*W
\]
as the finite-dimensional Radon–Nikodym density on the Kolmogorov span [2509.20496]. For shifted kernels on $\mathbb F_d^+$, one may compute the abstract shift matrices $B_i$ directly and set
\[
D=\sum_i B_i^*B_i.
\]
If empirical violations are small and attributable to noise, the paper recommends regularization by adding $\alpha I$ and retesting domination for
\[
K^{(\alpha)}:=K+\alpha I_X
\]
[2509.20496].

A different computational line arises in RKHS estimation of classical Radon–Nikodym derivatives. If $q\ll p$ with density $\beta=dq/dp$, then in the RKHS framework of [2308.07887] the derivative solves
\[
J_p^*J_p\,\beta = J_q^*J_q\, \mathbf 1,
\]
and is estimated by a spectral filter:
\[
\beta_{\mathbf X}^\lambda
=
g_\lambda(X_p^*X_p)\, X_q^*X_q\, \mathbf 1.
\]
For the iterated Lavrentiev filter,
\[
g_{\lambda,k}(t)=\frac{1-\left(\frac{\lambda}{\lambda+t}\right)^k}{t},
\]
the paper gives an explicit recursion in terms of the Gram matrix $\mathbf K$ and proves global and pointwise error bounds controlled by regularized Christoffel functions [2308.07887]. This is not the same theorem as the operator-density result for positive-definite kernels, but it shows how Radon–Nikodym differentiation interacts with kernel methods algorithmically.

Taken together, these developments show that the phrase *Radon–Nikodym theorem for kernels* does not designate a single theorem. It denotes a cluster of closely related representation principles: measurable densities for dominated measurable kernels, positive-contraction densities on Kolmogorov or RKHS feature spaces for dominated positive definite kernels, and commutant-valued densities in Stinespring or KSGNS dilations for noncommutative kernels [1810.01837] [2509.20496] [2405.09315] [1608.01672].

Source: https://www.emergentmind.com/topics/radon-nikodym-theorem-for-kernels