---
title: Generalized Wasserstein-2 Distance
url: https://www.emergentmind.com/topics/generalized-wasserstein-2-distance
type: topic
---

# Generalized Wasserstein-2 Distance

Searching arXiv for the cited works to ground the article in current arXiv metadata and related formulations.
arxiv_search query: 1304.7014 Generalized Wasserstein distance Piccoli Rossi

arxiv_search({"query":"1304.7014 Piccoli Rossi generalized Wasserstein distance", "max_results": 5})
The generalized Wasserstein-2 distance denotes a family of extensions of the classical quadratic Wasserstein distance \(W_2\) in which one alters either the admissible measures, the ground geometry, the comparison class, or the ambient state space. In the literature surveyed here, the expression covers unbalanced transport on finite measures, kernelized Gaussian \(W_2\) in reproducing-kernel Hilbert spaces, restricted-potential approximations, translation-invariant quotient metrics, Hausdorff-type distances on sets of measures, path-space and local-in-time constructions for stochastic dynamics, geodesic restrictions to submanifolds of Wasserstein space, and Bures–Wasserstein linearizations in the quantum setting [1304.7014], [1905.09314], [1902.07197], [2409.02416], [1505.04954], [2401.11354], [2311.08549], [2410.04937]. This suggests that “generalized Wasserstein-2 distance” is not a single canonical object, but a family of quadratic optimal-transport geometries adapted to different structural constraints.

## 1. Unbalanced quadratic transport on finite measures

The most direct generalization removes the equal-mass restriction of classical Wasserstein distance. In the Piccoli–Rossi formulation, the underlying space is
\[
\mathcal{M}:=\{\text{positive Borel regular measures on }\mathbb{R}^d\text{ with finite mass}\},
\]
and for \(p\ge 1\), \(a,b>0\), one defines
\[
T_{a,b}(\mu,\nu)
:= \inf_{\tilde\mu,\tilde\nu\in\mathcal{M},\;|\tilde\mu|=|\tilde\nu|}
\Big(
a^p\big(|\mu-\tilde\mu|+|\nu-\tilde\nu|\big)^p
+b^pW_p^p(\tilde\mu,\tilde\nu)
\Big),
\]
followed by
\[
W_{a,b}(\mu,\nu):=\big(T_{a,b}(\mu,\nu)\big)^{1/p}.
\]
For \(p=2\), this yields the generalized Wasserstein-2 distance \(W_2^{a,b}\): mass may be removed from \(\mu\) and \(\nu\) at total-variation cost and the surviving equal masses are then transported at quadratic Wasserstein cost [1304.7014].

A related earlier unbalanced construction writes
\[
W_{p,a,b}(\mu,\nu)
:=\inf_{\tilde\mu,\tilde\nu\in\mathcal M_p,\ |\tilde\mu|=|\tilde\nu|}
\Bigl(a|\mu-\tilde\mu|+a|\nu-\tilde\nu|+bW_p(\tilde\mu,\tilde\nu)\Bigr).
\]
It has the same conceptual interpretation: one first equalizes masses by removing mass, then transports the remaining common mass [1206.3219].

On equal-mass measures, the generalized construction reduces to classical transport in the mass-preserving regime: an admissible choice is \(\tilde\mu=\mu\), \(\tilde\nu=\nu\), so one obtains \(W_{a,b}(\mu,\nu)=bW_p(\mu,\nu)\) when the minimizer keeps all mass [1304.7014]. For Dirac masses of equal weight, the competition between transport and deletion is explicit:
\[
T_{a,b}(\delta_x,\delta_y)=\min\{b^2|x-y|^2,\;4a^2\},
\]
showing that short displacements favor transport while large displacements may favor annihilation and recreation [1304.7014].

| Construction | Domain | Defining feature |
|---|---|---|
| \(W_2^{a,b}\) / \(W_{a,b}\) | finite positive Borel measures | transport plus total variation [1304.7014] |
| Kernel \(W_2\) | Gaussian measures in an RKHS | Gaussian \(W_2\) in feature space [1905.09314] |
| \(W_{2,\mathcal F}\) | probability measures on \(\mathbb R^d\) | restriction to convex potentials in \(\mathcal F\) [1902.07197] |
| \(RW_2\) | \(\mathcal P_2(\mathbb R^n)/\sim\) | optimization over translations [2409.02416] |
| \(\mathcal W_2\) | weakly compact convex sets of measures | Hausdorff-type max-sup-inf of \(W_2\) [1505.04954] |
| \(W_\Lambda\) | submanifolds of \(\mathcal P_{\mathrm{a.c.}}(\Omega)\) | geodesic restriction of \(W_2\) [2311.08549] |

## 2. Dynamic, dual, and topological structure

For the unbalanced quadratic theory, the classical Benamou–Brenier formulation is replaced by a continuity equation with source,
\[
\partial_t\mu_t+\nabla\cdot(\mu_t v_t)=h_t,
\]
and the action functional becomes
\[
\mathcal B_{a,b}[\mu,v,h]
=
a^2\int_0^1\!\!\int d|h_t|
+
b^2\int_0^1\!\!\int |v_t|^2\,d\mu_t.
\]
The generalized Benamou–Brenier formula states
\[
\big(W_{a,b}(\mu_0,\mu_1)\big)^2
=
\inf_{(\mu,v,h)}\mathcal B_{a,b}[\mu,v,h],
\]
so \(W_2^{a,b}\) is the minimum cost of combining kinetic transport and source terms. When \(h\equiv0\) and \(a\to+\infty\), the formula reduces to the classical mass-preserving Benamou–Brenier identity [1304.7014].

The same paper establishes a complementary dual result at \(p=1\): \(W_1^{1,1}\) coincides with the flat metric,
\[
W_{1,1}(\mu,\nu)
=
\sup\left\{
\int f\,d(\mu-\nu)\ \middle|\ f\in C_0,\ \|f\|_\infty\le1,\ \mathrm{Lip}(f)\le1
\right\}.
\]
Although this is not a \(W_2\) formula, it identifies the dual structure associated with the same generalized family and clarifies how transport and mass variation appear simultaneously in the dual constraints [1304.7014].

Topologically, the unbalanced metric is complete and metrizes weak convergence for tight sequences. More precisely,
\[
W_{a,b}(\mu_n,\mu)\to0
\]
is equivalent to weak convergence \(\mu_n\rightharpoonup\mu\) together with tightness, and \((\mathcal M,W_{a,b})\) is complete [1304.7014], [1206.3219]. A useful estimate is
\[
\left|\int f\,d\mu-\int f\,d\nu\right|
\le
\sqrt{2}\,\max\left\{\|f\|_\infty,\frac{\mathrm{Lip}(f)}{b}\right\}W_{a,b}(\mu,\nu),
\]
for \(f\in \mathrm{Lip}(\mathbb R^d)\cap L^\infty(\mathbb R^d)\), which makes the bounded-Lipschitz control explicit [1304.7014].

## 3. Measure dynamics, stochastic processes, and mixed random fields

The unbalanced metric was introduced in part to study continuity equations with source,
\[
\partial_t\mu_t+\nabla\cdot(v_t\mu_t)=h_t,
\]
for which classical \(W_p(\mu_t,\mu_s)\) may be undefined when masses differ. Under Lipschitz assumptions on the velocity field and source with respect to the generalized distance, one obtains existence and uniqueness of solutions to the Cauchy problem, together with stability estimates of the form
\[
W_{p,a,b}(\mu_t,\nu_t)\le e^{Ct}W_{p,a,b}(\mu_0,\nu_0)
\]
and flow estimates under pushforwards by Lipschitz vector fields [1206.3219].

For stochastic differential equations, one generalization acts on path laws. On the Hilbert space \(C([0,T];\mathbb R^d)\) with norm
\[
\|\mathbf X\|=\Big(\int_0^T\sum_{i=1}^d |X_i(t)|^2\,dt\Big)^{1/2},
\]
the path-space quadratic Wasserstein distance is
\[
W_2^2(\mu,\hat\mu)
=
\inf_{\pi(\mu,\hat\mu)}
\mathbb E\big[\|\mathbf X-\hat{\mathbf X}\|^2\big].
\]
Because this object is computationally demanding, the paper introduces the time-decoupled functional
\[
\tilde W_2^2(\mu,\hat\mu)
:=
\int_0^T W_2^2(\mu(t),\hat\mu(t))\,dt
\le
W_2^2(\mu,\hat\mu),
\]
which serves as an efficient generalized Wasserstein-2 loss for reconstructing SDEs from noisy data [2401.11354].

A further mixed-type generalization replaces the Euclidean cost by a custom cost on \(\mathbb R^{d_1}\times S_{d-d_1}\), where the first \(d_1\) coordinates are continuous and the remaining ones are categorical. The corresponding distance,
\[
\hat W_2(f_{\mathbf x},\hat f_{\mathbf x})
=
\inf_{\pi_{f_{\mathbf x},\hat f_{\mathbf x}}}
\Big(
\mathbb E[\|\mathbf y_{\mathbf x}-\hat{\mathbf y}_{\mathbf x}\|^2]
\Big)^{1/2},
\]
is then integrated over the input domain,
\[
\hat W_2^2(\mathbf y_{\mathbf x},\hat{\mathbf y}_{\mathbf x})
=
\int_D \hat W_2^2(f_{\mathbf x},\hat f_{\mathbf x})\,\nu(d\mathbf x).
\]
Its empirical approximation is the generalized local squared Wasserstein-2 loss, built from neighborhood-wise empirical measures and used to train stochastic neural networks for classification, mixed random-variable reconstruction, and noisy dynamical systems [2507.05143].

## 4. Kernel, restricted, and optimization-based redefinitions

A different line of work keeps mass preservation but changes the ground geometry. In the kernel construction, empirical distributions are mapped into an RKHS \(\mathcal H\) by a feature map \(\phi\), approximated by Gaussian measures \(N(\boldsymbol\mu_i,\boldsymbol\Sigma_i)\) in \(\mathcal H\), and compared by the Gaussian \(W_2\) formula
\[
W_2(k\nu_1,k\nu_2)^2
=
\|\boldsymbol\mu_1-\boldsymbol\mu_2\|_{\mathcal H}^2
+
\mathrm{tr}\!\Big(
\boldsymbol\Sigma_1+\boldsymbol\Sigma_2
-
2(\boldsymbol\Sigma_1^{1/2}\boldsymbol\Sigma_2\boldsymbol\Sigma_1^{1/2})^{1/2}
\Big).
\]
The mean term is exactly the empirical squared MMD, while the covariance term is expressed entirely through kernel Gram matrices. For the linear kernel, the construction reduces to the usual Gaussian \(W_2\) in \(\mathbb R^d\) [1905.09314].

Another generalization restricts the Kantorovich dual to a convex class \(\mathcal F\subset cvx(\mathbb R^d)\). The restricted distance is
\[
W_{2,\mathcal F}^2(\mu,\nu)
=
\int \frac12\|x\|_2^2\,d\mu
+
\int \frac12\|y\|_2^2\,d\nu
-
\inf_{\theta\in\Theta}J_{\mu,\nu}(\theta),
\]
with \(J_{\mu,\nu}(\theta)=\int f(x;\theta)\,d\mu+\int f^\star(y;\theta)\,d\nu\). Because \(\mathcal F\subset cvx\), one always has \(W_{2,\mathcal F}(\mu,\nu)\le W_2(\mu,\nu)\). In general it is not symmetric, and for conic classes it becomes a pseudo-metric characterized by restricted moment matching. Input-convex neural networks provide an explicit parametrization of \(\mathcal F\) and an approximate transport map \(T_{\mathcal F}=\nabla f^\star\) [1902.07197].

The cost itself may also be generalized. In a GAN setting, one may replace Euclidean or squared Euclidean cost by any continuous transportation cost \(c(x,y)\); for \(c(x,y)=\|x-y\|^2\), one recovers a genuine \(W_2\)-type geometry, while image-centered costs such as SSIM induce different transport-based discrepancies. The associated “assigner” network represents a Kantorovich potential whose \(c\)-transform induces a transport map between generated and real samples [1910.00535].

Optimization-oriented formulations further reinterpret \(W_2\) as training geometry. One approach pulls back optimal transport structures from probability space to parameter space and defines a parametrization-invariant natural gradient together with Wasserstein proximal regularizers for generator updates [2102.06862]. Another constructs a distribution-dependent ODE whose dynamics involves the Kantorovich potential between the current estimate and the true data distribution; the time-marginal laws form a gradient flow for the \(W_2\) loss and converge exponentially to the true data distribution [2406.13619].

## 5. Quotients, ambiguity sets, and intrinsic submanifolds

A quotient-space generalization removes global translations. For probability measures on \(\mathbb R^n\), define \(\mu\sim\nu\) if \(\nu=T_s\#\mu\) for some shift \(s\in\mathbb R^n\), and set
\[
RW_p([\mu],[\nu])
:=
\min_{\mu'\in[\mu],\nu'\in[\nu]}W_p(\mu',\nu').
\]
For \(p=2\),
\[
RW_2^2(\mu,\nu)
=
\min_{s\in\mathbb R^n}\min_{P\in\Pi(a,b)}
\sum_{i,j}\|x_i-y_j-s\|_2^2P_{ij},
\]
and one obtains the exact decomposition
\[
W_2^2(\mu,\nu)
=
RW_2^2(\mu,\nu)+\|\bar\mu-\bar\nu\|_2^2.
\]
The first term captures shape difference after optimal recentering, while the second isolates mean translation [2409.02416].

A set-valued generalization replaces single measures by weakly compact convex sets of measures. For \(\mathcal P_1,\mathcal P_2\subset\mathcal P(\Omega)\),
\[
\mathcal W_2(\mathcal P_1,\mathcal P_2)
=
\max\Big(
\sup_{\mu\in\mathcal P_1}\inf_{\nu\in\mathcal P_2}W_2(\mu,\nu),
\sup_{\nu\in\mathcal P_2}\inf_{\mu\in\mathcal P_1}W_2(\mu,\nu)
\Big).
\]
This is the Hausdorff distance induced by \(W_2\) on ambiguity sets, and on \(\mathscr P_2(\Omega)\) it metrizes weak convergence of the associated sublinear expectations together with the natural second-moment tail condition [1505.04954].

A submanifold construction restricts the ambient Wasserstein geometry to a finite-dimensional embedded manifold \(\Lambda\subset\mathcal P_{\mathrm{a.c.}}(\Omega)\). If \(E:S\to\Lambda\) is an embedding from a compact Riemannian manifold \(S\), the intrinsic metric is
\[
W_\Lambda(\mu,\nu)
=
\inf\{\mathrm{energy}(\rho)^{1/2}:\rho\in \mathrm{Lip}([0,1];(\Lambda,W)),\ \rho(0)=\mu,\ \rho(1)=\nu\},
\]
that is, the geodesic restriction of the ambient \(W_2\). The resulting geometry is not necessarily flat, but it admits local linearizations of Riemannian type, and the latent metric space \((\Lambda,W_\Lambda)\) can be asymptotically recovered from samples and pairwise extrinsic Wasserstein distances in the sense of Gromov–Wasserstein [2311.08549]. Related linearized Gromov–Wasserstein constructions use tangent-space or barycentric-projection representations to approximate pairwise GW distances between metric-measure spaces [2112.11964].

## 6. Quantum Bures–Wasserstein geometry and the scope of the term

In the noncommutative setting, the manifold of positive definite matrices \(\mathbb P_d\) carries the Bures–Wasserstein metric, whose geodesic distance is
\[
d_{\mathrm{BW}}^2(P,Q)
=
\operatorname{Tr}[P+Q]-2\operatorname{F}^{\mathrm U}(P,Q).
\]
A base-point-dependent linearization defines the generalized fidelity
\[
\operatorname{F}_R(P,Q)
:=
\operatorname{Tr}\Bigl[
\sqrt{R^{1/2}PR^{1/2}\;R^{-1}\;\sqrt{R^{1/2}QR^{1/2}}}
\Bigr]
\]
and the generalized Bures–Wasserstein squared distance
\[
\operatorname{B}_R(P,Q)
:=
\operatorname{Tr}[P+Q]-2\,\Re\operatorname{F}_R(P,Q).
\]
Its defining geometric identity is
\[
\operatorname{B}_R(P,Q)
=
\|\mathrm{Log}_R[P]-\mathrm{Log}_R[Q]\|_R^2,
\]
so the generalized distance is the BW tangent-space norm between the logarithmic images of \(P\) and \(Q\) at an arbitrary base \(R\) [2410.04937].

This construction recovers standard quantum fidelities as special base choices: \(R=P\) or \(R=Q\) yields Uhlmann fidelity, \(R=\mathbb I\) yields Holevo fidelity, and \(R=P^{-1}\) or \(R=Q^{-1}\) yields Matsumoto fidelity [2410.04937]. A common misconception is that generalized Wasserstein-2 is synonymous with a single unbalanced transport metric. In the surveyed literature, the same expression is used for unbalanced finite measures, kernelized feature-space Gaussians, restricted dual classes, translation quotients, ambiguity sets, stochastic-process laws, intrinsic submanifolds, and quantum BW linearizations. What these constructions share is quadratic transport geometry; what differs is the structural constraint under which that geometry is defined.

Source: https://www.emergentmind.com/topics/generalized-wasserstein-2-distance