---
title: Hellinger Discrepancy Tensor in HK Geometry
url: https://www.emergentmind.com/topics/hellinger-discrepancy-tensor
type: topic
---

# Hellinger Discrepancy Tensor in HK Geometry

Searching arXiv for the cited paper and closely related work to ground the article.
The **Hellinger discrepancy tensor** denotes the Riemannian metric tensor associated with the local linearization of the Hellinger–Kantorovich (HK) distance at a reference measure $\mu_0$. In the terminology of "The Linearized Hellinger--Kantorovich Distance" [2102.08807], if “Hellinger discrepancy tensor” is not standard terminology, it precisely refers to this HK Riemannian metric tensor at $\mu_0$. The construction combines transport and mass variation in a single inner product, supports logarithmic and exponential maps, and yields a local Hilbertian embedding of measures around $\mu_0$. The resulting framework interpolates between Wasserstein-$2$ geometry and Fisher–Rao/Hellinger geometry through an intrinsic length scale $\kappa$, while retaining an exact static and dynamic interpretation of tangent directions, geodesics, and measure creation.

## 1. Geometric setting and dynamic formulation

The framework is posed on a domain $\Omega \subset \mathbb{R}^d$ that is convex, closed, bounded, and has non-empty interior. Lebesgue measure on $\Omega$ is denoted by $L$. The relevant spaces are signed Radon measures $\mathrm{meas}(\Omega)$, nonnegative Radon measures $\mathrm{meas}^+(\Omega)$, and absolutely continuous nonnegative reference measures $\mu_0$ with respect to $L$.

The HK distance is defined dynamically through a Benamou–Brenier-type action. For a time-dependent measure $\rho_t \in \mathrm{meas}^+(\Omega)$, momentum $\omega_t \in \mathrm{meas}(\Omega)^d$, and source $\zeta_t \in \mathrm{meas}(\Omega)$, one imposes the continuity equation with source
\[
\partial_t \rho_t + \nabla\!\cdot(\omega_t) = \zeta_t
\quad\text{in the sense of distributions,}
\]
or, when $\omega_t = v_t \rho_t$ and $\zeta_t = \alpha_t \rho_t$,
\[
\partial_t \rho_t + \nabla\!\cdot(\rho_t v_t) = \alpha_t\,\rho_t.
\]
The canonical HK action is
\[
J_{HK}(\rho,\omega,\zeta)
=
\int_0^1 \int_{\Omega} \Big( |v_t(x)|^2 + \tfrac{1}{4}\,\alpha_t(x)^2 \Big)\, d\rho_t(x)\,dt,
\]
and the squared HK distance is the infimum of this action over admissible triples connecting $\mu_0$ and $\mu_1$ [2102.08807].

A scaled family replaces the factor $\tfrac14$ by $\tfrac{\kappa^2}{4}$:
\[
J_{HK,\kappa}(\rho,\omega,\zeta)
=
\int_0^1 \int_{\Omega} \Big( |v_t|^2 + \tfrac{\kappa^2}{4}\,\alpha_t^2 \Big)\, d\rho_t\,dt,
\qquad
\operatorname{HK}_\kappa(\mu_0,\mu_1)^2
=
\inf J_{HK,\kappa}(\rho,\omega,\zeta).
\]
The parameter $\kappa>0$ is the intrinsic length scale that balances transport and mass change. Precisely,
\[
\operatorname{HK}_\kappa(\mu_0,\mu_1) \to W_2(\mu_0,\mu_1)\quad\text{as}\ \kappa\to\infty,
\qquad
\frac{1}{\kappa}\operatorname{HK}_\kappa(\mu_0,\mu_1) \to \operatorname{Hell}(\mu_0,\mu_1)\quad\text{as}\ \kappa\to 0.
\]
For $\kappa=1$, transported mass never travels farther than distance $\pi/2$, a global transport bound that later reappears as a cut-locus phenomenon.

## 2. Tangent representation and the metric tensor

Informally, the tangent space at $\mu_0$ consists of pairs $(v,\alpha)$ with
\[
v \in L^2(\Omega,\mu_0)^d,\qquad \alpha \in L^2(\Omega,\mu_0),
\]
representing transport velocities and relative mass growth rates. For a smooth path $t\mapsto \rho_t$ with $\rho_0=\mu_0$, the infinitesimal constraint at $t=0$ is
\[
\partial_t \rho_t\big|_{t=0} + \nabla\!\cdot(\mu_0 v_0) = \alpha_0\,\mu_0,
\]
together with a singular creation part at points where mass appears. Accordingly, the tangent representation contains a third component that encodes creation from nothing.

The logarithmic map is defined by
\[
\log_{HK}(\mu_0;\mu_1)\equiv \mathrm{Log}_{HK}(\mu_0;\mu_1):=(v_0,\alpha_0,\sqrt{\mu_1^\perp}),
\]
where $\mu_1^\perp$ is the singular “created” part and $(v_0,\alpha_0)$ are the initial transport and reaction components on the moving part. The induced inner product at $\mu_0$ is
\[
\big\langle (v_0,\alpha_0,\sqrt{\nu}),\,(\tilde v_0,\tilde \alpha_0,\sqrt{\tilde \nu}) \big\rangle_{HK,\mu_0}
:=
\int_{\Omega} \Big( \langle v_0,\tilde v_0\rangle + \tfrac{1}{4}\,\alpha_0 \tilde\alpha_0 \Big)\, d\mu_0
\;+\;
\int_{\Omega} \sqrt{\frac{d\nu}{d\lambda}\,\frac{d\tilde \nu}{d\lambda}}\, d\lambda,
\]
for any dominating measure $\lambda$ with $\nu,\tilde\nu \ll \lambda$. With scale parameter $\kappa$, the term $\tfrac14$ is replaced by $\tfrac{\kappa^2}{4}$ [2102.08807].

This tensor measures transport by the $\mu_0$-weighted Euclidean norm of $v$ and mass change by the $\mu_0$-weighted Fisher–Rao norm of $\alpha$, with the precise $\tfrac14$ or $\tfrac{\kappa^2}{4}$ factor fixed by the HK action. The associated local norm is
\[
\big\|(v,\alpha,\sqrt{\nu})\big\|_{HK,\mu_0}^2
=
\int_{\Omega}\Big(|v|^2+\tfrac{1}{4}\alpha^2\Big)\,d\mu_0
\;+\;
\int_{\Omega}\frac{d\nu}{d\lambda}\,d\lambda,
\]
again with $\tfrac14$ replaced by $\tfrac{\kappa^2}{4}$ in the scaled case. For a small displacement $\rho=\mu_0+\varepsilon\cdot(\cdots)$, the HK distance admits the first-order approximation
\[
\operatorname{HK}(\mu_0,\rho)\approx \big\|(v,\alpha,\sqrt{\nu})\big\|_{HK,\mu_0}.
\]
When $\mu_1^\perp=0$, the third term vanishes and the embedding is Hilbertian on $L^2(\mu_0;\mathbb{R}^d)\times L^2(\mu_0;\mathbb{R})$.

## 3. Static formulation, tangent components, and exact norm identities

The same geometry has a static Kantorovich-type formulation with soft marginals:
\[
\operatorname{HK}(\mu_0,\mu_1)^2
=
\inf_{\pi\in\mathrm{meas}^+(\Omega^2)}
\left\{
\int_{\Omega^2} c(x_0,x_1)\,d\pi
+
\sum_{i\in\{0,1\}}
\mathrm{KL}\!\big(P_{i\sharp}\pi\,\big|\,\mu_i\big)
\right\},
\]
with
\[
c(x_0,x_1):= -2\log\big(\cos(\|x_0-x_1\|)\big)
\quad\text{for}\ \|x_0-x_1\|<\pi/2,
\]
and $c=+\infty$ otherwise. If $\mu_0$ is nonnegative and absolutely continuous with respect to $L$, the minimizer is unique and is induced by a Monge map:
\[
\pi=(\mathrm{id},T)_\sharp \sigma,
\qquad \sigma=P_{0\sharp}\pi.
\]

Writing the Lebesgue decomposition with respect to the marginals $\nu_i=P_{i\sharp}\pi$ as
\[
\mu_i=u_i\,\nu_i+\mu_i^\perp,
\]
one obtains explicit formulas for the tangent data on the moving part:
\[
v_0(x)
=
\frac{T(x)-x}{\|T(x)-x\|}
\cdot
\sqrt{\frac{u_1(T(x))}{u_0(x)}\,\sin\big(\|T(x)-x\|\big)},
\]
with the convention $v_0(x)=0$ if $T(x)=x$, and
\[
\alpha_0(x)
=
2\Big(\sqrt{\frac{u_1(T(x))}{u_0(x)}\,\cos\big(\|T(x)-x\|\big)} - 1\Big).
\]
On the disappearing part, $\mu_0^\perp$-almost everywhere,
\[
\alpha_0(x)=-2,\qquad v_0(x)=0,
\]
while the remaining appearing mass is $\mu_1^\perp$ [2102.08807].

These quantities satisfy the exact identity
\[
\operatorname{HK}(\mu_0,\mu_1)^2
=
\int_{\Omega}\Big(|v_0|^2 + \tfrac{1}{4}\alpha_0^2\Big)\,d\mu_0
\;+\; \|\mu_1^\perp\|.
\]
This identity is the exact counterpart of “distance equals squared Riemannian norm of the logarithmic map” and motivates the metric tensor. A plausible implication is that the tensor is not merely a heuristic quadratic surrogate: it is directly induced by the exact decomposition of an HK geodesic into moving, disappearing, and appearing components.

## 4. Logarithmic and exponential maps

The logarithmic map at $\mu_0$ is defined from the optimal soft-marginal coupling by
\[
\mathrm{Log}_{HK}(\mu_0;\mu_1):=(v_0,\alpha_0,\sqrt{\mu_1^\perp}).
\]
Along the constant-speed HK geodesic $(\rho_\tau)_{\tau\in[0,1]}$ from $\mu_0$ to $\mu_1$, it scales exactly as
\[
\mathrm{Log}_{HK}(\mu_0;\rho_\tau)
=
\big(\tau\,v_0,\ \tau\,\alpha_0,\ \tau\,\sqrt{\mu_1^\perp}\big).
\]
This linear scaling is the central structural fact underlying the local linearization.

The exponential map reconstructs a measure from tangent data. Given
\[
\xi=(v_0,\alpha_0,\sqrt{\mu_1^\perp}),
\]
define
\[
a(x):=\|v_0(x)\|,\qquad b(x):=\frac{\alpha_0(x)}{2}+1,\qquad
S:=\{x\in\Omega\mid a(x)=0,\ b(x)=0\},
\]
\[
q(x):=\sqrt{a(x)^2+b(x)^2},
\]
and choose $\varphi(x)\in[0,\pi/2]$ such that
\[
(a,b)=q\cdot(\sin\varphi,\cos\varphi).
\]
Then
\[
\widehat{T}(x):=x+\begin{cases}
\frac{v_0(x)}{\|v_0(x)\|}\,\varphi(x), & v_0(x)\neq 0,\\
0, & v_0(x)=0,
\end{cases}
\]
and
\[
\mathrm{Exp}_{HK}(\mu_0;v_0,\alpha_0,\sqrt{\mu_1^\perp})
=
\widehat{T}_{\sharp}\big(q^2\,\mu_0\!\restr(\Omega\setminus S)\big) + \mu_1^\perp.
\]
The constant-speed geodesic is
\[
t\mapsto \mathrm{Exp}_{HK}(\mu_0;t\,v_0,t\,\alpha_0,t\,\sqrt{\mu_1^\perp}),
\]
and this matches the geodesic constructed from the optimal soft-marginal coupling [2102.08807].

If $\mu_0$ is nonnegative and absolutely continuous, the optimal soft-marginal coupling is unique and induced by a Monge map; hence the logarithmic map is unique. The exponential map reconstructs $\mu_1$ exactly when fed the HK logarithm. This suggests that the local tensor, the logarithmic map, and the exponential map form a coherent Riemannian package, although the paper emphasizes local Euclidean behavior rather than a fully global smooth geometry.

## 5. Interpolation between Wasserstein and Fisher–Rao/Hellinger geometries

The scaled family $\operatorname{HK}_\kappa$ interpolates between transport-dominated and reaction-dominated regimes. As $\kappa\to\infty$, transport dominates and $\operatorname{HK}_\kappa$ converges to $W_2$. At the level of tensors, the HK metric reduces to
\[
\int_{\Omega}\langle v,\tilde v\rangle\,d\mu_0,
\]
which is the Wasserstein metric tensor at $\mu_0$ on tangent velocities.

As $\kappa\to 0$, after dividing distances by $\kappa$, $\operatorname{HK}_\kappa/\kappa$ converges to the Hellinger–Kakutani metric. Correspondingly, the tensor reduces to the Fisher–Rao inner product on reaction rates,
\[
\int_{\Omega} \tfrac{1}{4}\,\alpha\,\tilde\alpha\,d\mu_0,
\]
plus the pure Hellinger term on measure creation, namely the square-root inner product on measures [2102.08807].

This places the Hellinger discrepancy tensor in a comparative triad:

\[
\langle v,\tilde v\rangle_{W_2,\mu_0}
=
\int_{\Omega}\langle v,\tilde v\rangle\,d\mu_0,
\]
\[
\langle \alpha,\tilde\alpha\rangle_{FR,\mu_0}
=
\int_{\Omega}\tfrac{1}{4}\,\alpha\,\tilde\alpha\,d\mu_0,
\]
and
\[
\langle (v,\alpha,\sqrt{\nu}),(\tilde v,\tilde\alpha,\sqrt{\tilde\nu})\rangle_{HK,\mu_0}
=
\int_{\Omega}\Big(\langle v,\tilde v\rangle+\tfrac{1}{4}\alpha\tilde\alpha\Big)\,d\mu_0
+
\int_{\Omega}\sqrt{\frac{d\nu}{d\lambda}\frac{d\tilde\nu}{d\lambda}}\,d\lambda.
\]

Transport and mass change are therefore measured in commensurate units. The paper states that this explains HK’s robustness to local mass fluctuations. A common misconception is to treat HK as merely a convex combination of Wasserstein and Hellinger metrics; the tensorial description shows instead that HK has its own Riemannian structure with an exact coupling of transport, relative reaction, and ex nihilo creation.

## 6. Geodesics, cut locus, and discrete embeddings

For Dirac masses, the geometry becomes explicit. If one connects $m_0\delta_{x_0}$ to $m_1\delta_{x_1}$ with $d:=\|x_1-x_0\|<\pi/2$, the HK geodesic remains a single moving Dirac,
\[
\rho_t=\delta_{X(t)}\,M(t),
\]
with
\[
M(t)=(1-t)^2 m_0 + t^2 m_1 + 2t(1-t)\sqrt{m_0 m_1}\cos d,
\]
and
\[
\big(|\dot X(t)|^2 + \tfrac{1}{4}(\dot M(t)/M(t))^2\big)\,M(t)
\equiv
\operatorname{HK}(m_0\delta_{x_0},m_1\delta_{x_1})^2.
\]
At $t=0$,
\[
\dot X(0)=\tfrac{x_1-x_0}{\|x_1-x_0\|}\,\sqrt{\tfrac{m_1}{m_0}\,\sin d},
\qquad
\frac{\dot M(0)}{M(0)}=2\Big(\sqrt{\tfrac{m_1}{m_0}\,\cos d}-1\Big),
\]
in agreement with the formulas for $(v_0,\alpha_0)$. When $d\ge \pi/2$, transport switches off and mass teleports via a pure Fisher–Rao geodesic. The cut locus at $\pi/2$ causes discontinuities in the logarithmic and exponential maps across that threshold; away from it, the Riemannian structure is smooth [2102.08807].

The same formalism yields a practical linearization pipeline for discrete data. For a discrete reference
\[
\mu_0=\sum_i m_i^0 \delta_{x_i}
\]
and samples $\mu_k$, one solves the soft-marginal HK problem by entropic unbalanced Sinkhorn to obtain optimal plans $\pi^{(k)}$. If $\mu_0\ll L$ is approximated on a grid, barycentric projection yields an approximate Monge map $T^{(k)}$ and disintegrations $\{\pi^{(k)}_{x_i}\}$. One then computes
\[
u_1\big(T^{(k)}(x_i)\big)\ \approx\ \int_\Omega \frac{d\mu_k}{d(P_{1\sharp}\pi^{(k)})}(x_1)\,d\pi^{(k)}_{x_i}(x_1),
\]
followed by
\[
v_0^{(k)}(x_i)
=
\frac{T^{(k)}(x_i)-x_i}{\|T^{(k)}(x_i)-x_i\|}
\sqrt{\frac{u_1(T^{(k)}(x_i))}{u_0(x_i)}\,\sin\big(\|T^{(k)}(x_i)-x_i\|\big)},
\]
\[
\alpha_0^{(k)}(x_i)
=
2\Big(\sqrt{\frac{u_1(T^{(k)}(x_i))}{u_0(x_i)}\,\cos\big(\|T^{(k)}(x_i)-x_i\|\big)}-1\Big),
\]
with creation component $\mu_k^\perp$. The linearized embedding of $\mu_k$ is the vector field–scalar pair $x_i\mapsto (v_0^{(k)}(x_i),\alpha_0^{(k)}(x_i))$ equipped with inner product
\[
\sum_i \Big(\langle v_0^{(k)}(x_i), v_0^{(\ell)}(x_i)\rangle + \tfrac{1}{4}\alpha_0^{(k)}(x_i)\alpha_0^{(\ell)}(x_i)\Big)\,m_i^0,
\]
plus the Hellinger inner product of the creation parts if present.

The stated computational advantage is that embedding $n$ samples requires $n$ HK solves rather than $O(n^2)$ pairwise distances, after which PCA, LDA, and SVM can be applied in Euclidean space. The paper further notes that a wide-support $\mu_0$, such as a uniform grid or HK barycenter, typically yields $\mu_k^\perp=0$ and hence a purely Hilbert embedding. Accuracy is best when samples lie near $\mu_0$ along HK geodesics; large displacements, especially those crossing the $\pi/2$ cut locus, degrade linearization fidelity.

The Hellinger discrepancy tensor is therefore the local object that makes this pipeline possible: it converts HK geometry into a weighted Euclidean structure while preserving, to first order and in several cases exactly, the combined effects of transport, disappearance, and appearance of mass.

Source: https://www.emergentmind.com/topics/hellinger-discrepancy-tensor