---
title: Generalized Wasserstein Geometries
url: https://www.emergentmind.com/topics/generalized-wasserstein-geometries
type: topic
---

# Generalized Wasserstein Geometries

Generalized Wasserstein geometries consist of a broad family of metric structures and optimization frameworks that extend the classical Wasserstein space of probability measures. These generalizations encompass unbalanced transport (allowing for mass creation/destruction), Bregman divergences, slicing-based metrics that exploit nonlinear projections, Gromov–Wasserstein frameworks that are invariant to isometries and reflect structural information, geometric analysis on spaces of SPD matrices and operators, synthetic and barycentric curvature-dimension conditions on abstract measure-metric spaces—including infinite-dimensional and non-smooth settings—as well as iterated and hierarchical constructions. This article surveys major constructions, geometric and variational properties, and recent advances in such generalized optimal transport geometries.

## 1. Unbalanced and Source-Regularized Wasserstein Geometries

A primary generalization over classical Wasserstein space addresses the limitation that $W_p$ is only defined between measures with equal mass. The Piccoli–Rossi metric $W_p^{a,b}$ on finite measures introduces two nonnegative weights $a$ (mass creation/removal) and $b$ (transport) and, for $\mu,\nu \in \mathcal M(\mathbb{R}^d)$,
\[
W_p^{a,b}(\mu,\nu) = \left[ \inf_{\substack{\tilde\mu \leq \mu,\,\tilde\nu \leq \nu\\ |\tilde\mu|=|\tilde\nu|}}\, a^p\big(\|\mu-\tilde\mu\|_{\rm TV}+\|\nu-\tilde\nu\|_{\rm TV}\big)^p + b^p W_p(\tilde\mu,\tilde\nu)^p\, \right]^{1/p}.
\]
This metric interpolates between strict mass-conservation transport ($a\to\infty$) and total variation ($b\to\infty$), and is a complete metric on measure spaces, admitting a generalized Benamou–Brenier formula for $W_2^{a,b}$ involving both velocities and signed source measures $s_t$ in the continuity equation,
\[
\partial_t\mu_t + \nabla\!\cdot(\mu_t v_t) = s_t.
\]
The corresponding action is
\[
\mathcal A_{a,b}[\mu,v,s] = a^2 \left( \int_0^1 \! \int_{\mathbb{R}^d} d|s_t|(x)\,dt \right)^2 + b^2 \int_0^1\! \int_{\mathbb{R}^d} \tfrac12 |v_t(x)|^2\,d\mu_t(x)\,dt.
\]
For $p=1$ and $a=b=1$, $W_1^{1,1}$ coincides with the flat (bounded-Lipschitz) metric, i.e.,
\[
W_1^{1,1}(\mu,\nu) = \sup \left\{ \int f\,d(\mu-\nu): \|f\|_\infty \leq 1,\, \mathrm{Lip}(f) \leq 1 \right\}.
\]
This structure underlies well-posedness for nonlinear transport equations with source terms and extends classical duality and stability—the metric space $(\mathcal M(X), W_p^{a,b})$ is geodesic and stable under Gromov–Hausdorff convergence [1304.7014, 1904.12461].

## 2. Sliced, Generalized Sliced, and Differentiable Sliced Wasserstein Distances

Metrics based on slicing project high-dimensional distributions onto lower-dimensional spaces to exploit the efficiency of 1D transport. The sliced Wasserstein metric,
\[
\mathrm{SW}_2(\mu,\nu) = \left[ \int_{\mathbb S^{d-1}} W_2^2 \left( P_\sharp^\theta\mu, P_\sharp^\theta\nu \right) d\theta \right]^{1/2},
\]
where $P^\theta$ is projection onto direction $\theta$, is generalized by replacing $P^\theta$ with nonlinear or learnable functions $g(x,\theta)$. The generalized sliced Wasserstein distance (GSW) for nonlinear $g$ is
\[
\mathrm{GSW}_p(\mu,\nu) = \left[ \int W_p^p\left( g^\theta_\sharp\mu,\,g^\theta_\sharp\nu \right) d\gamma_d(\theta) \right]^{1/p}.
\]
Recent developments provide deterministic and learnable function approximations (polynomials, neural nets), exploiting concentration of high-dimensional random projections, yielding scalability for high $d$ and facilitating moment-based approximations. Differentiable generalized sliced Wasserstein plans (DGSWP) employ a bilevel scheme to select optimal nonlinear projections $\phi$,
\[
\min_\theta\,\,\, h(\theta) = \langle C_{\mu\nu}, \pi^\theta \rangle,\quad \text{subject to}\; \pi^\theta \in \arg\min_{\pi \in U(a,b)} \sum_{ij} \pi_{ij} |\phi(x_i,\theta) - \phi(y_j,\theta)|^p,
\]
with gradients efficiently estimated by Gaussian smoothing over $\theta$ [2210.10268, 2505.22049]. Both Sliced and Generalized Sliced Wasserstein metrics are true metrics (injectivity conditions on $g$) and admit low-complexity computation.

The min-SWGG proxy leverages generalized Wasserstein geodesics with a line-supported pivot, computing
\[
\min\text{-SWGG}(\mu,\nu) = \min_{\theta \in \mathbb{S}^{d-1}} W_2^2 \left( P^\theta_\sharp\mu,\, P^\theta_\sharp\nu \right),
\]
which provides a metric that metrizes weak convergence and yields efficient computation and explicit couplings [2307.01770].

## 3. Gromov–Wasserstein and Linearized/Inner-Product Generalizations

Gromov–Wasserstein (GW) distances generalize $W_p$ to compare metric-measure spaces up to measure-preserving isometry, defined by
\[
GW_p^p(\mu_X, \mu_Y) = \min_{\pi \in \Pi(\mu_X, \mu_Y)} \iint_{X \times Y} \iint_{X \times Y} |d_X(x,x') - d_Y(y,y')|^p\, d\pi(x,y)\, d\pi(x',y').
\]
Linearized GW (LGW) and inner-product GW (IGW) geometries are proposed for computational tractability, e.g., LGW leverages barycentric projections and lot-based tangent embeddings, reducing quadratic complexity and retaining isometry-invariance properties [2112.11964]. The IGW metric,
\[
\mathrm{IGW}(\mu,\nu) = \inf_{\pi \in \Pi(\mu, \nu)} \iint \left| \langle x,x'\rangle - \langle y, y'\rangle \right|^2\, d(\pi \otimes \pi)(x,y,x',y')^{1/2},
\]
is analyzed with an associated gradient flow and Riemannian structure; the induced mobility operator modifies the local Wasserstein gradients to encode global structure, with a Benamou–Brenier-like dynamic reformulation and an Otto-calculus-type gradient [2407.11800].

## 4. Generalized Wasserstein Geometries on Metric-Measure and Infinite-Dimensional Spaces

The classical $W_2$ geometry on probability measures can be generalized to extended metric-measure spaces $(X,d,m)$, including abstract Wiener spaces and configuration spaces over Riemannian manifolds. A new barycentric curvature-dimension condition BCD$(K,N)$ is imposed via Jensen-type variational inequalities for entropy at barycenters,
\[
BCD(K,N):\quad U_N(\bar{\mu}) \geq \sum_i \lambda_i\, \sigma_{K/N}^{(1-\lambda_i)}(W_2(\mu_i, \bar{\mu}))\,U_N(\mu_i),
\]
with $U_N(\mu) = \exp(-Ent_m(\mu)/N)$. This condition encompasses curvature-dimension properties of the Lott–Sturm–Villani and Ambrosio–Gigli–Savaré theories but is designed to handle branching/non-geodesic and infinite-dimensional settings. Stability under measured Gromov–Hausdorff convergence holds, and existence, uniqueness, and absolute continuity of barycenters are established under mild integrability conditions. Geometric and functional inequalities, including multi-marginal Brunn–Minkowski and functional Blaschke–Santaló inequalities, are obtained directly from the barycentric Jensen inequalities [2412.01190].

Variational structures are lifted via category-theoretic functors to iterated Wasserstein spaces $P_2^{(n)}(M)$, with velocity plans and geodesics at every level, and gradient flows defined for suitable functionals [2512.03726]. In spaces of all signed Radon measures, the quotient structure under group actions descends naturally to Wasserstein distances and is compatible with Gromov–Hausdorff stability [1904.12461].

## 5. Bregman–Wasserstein and Dualistic Information-Geometric Extensions

The Bregman–Wasserstein divergence arises from replacing the quadratic cost in classical OT by a Bregman divergence (generated by a strictly convex function $\Omega$) on $M$:
\[
\mathscr{B}(\mu_0, \mu_1) = \inf_{\pi \in \Pi(\mu_0, \mu_1)} \int_{M \times M} \mathbf{B}(p, q)\,d\pi(p, q),
\]
where $\mathbf{B}(p, q)$ is the canonical Bregman divergence. This framework induces displacement interpolations corresponding to so-called primal and dual geodesics in information geometry, recovers the classical Wasserstein geometry for $\Omega(x) = \frac12 |x|^2$, and transports the dualistic (Amari) geometric structure to infinite-dimensional statistical manifolds. An associated generalized Pythagorean theorem, dual connections (primal, dual, Levi-Civita), and corresponding JKO gradient flows are established [2302.05833].

## 6. SPD Matrix and Operator-Valued Wasserstein Geometries

On the manifold of symmetric positive-definite (SPD) matrices $S_{++}^n$, the $2$-Wasserstein geometry is given by the metric tensor
\[
g_W|_A(X, Y) = \mathrm{tr}(\Gamma_A[Y]\,A\,\Gamma_A[X]) = \frac12\, \mathrm{tr}(\Gamma_A[Y] X),
\]
where $\Gamma_A[X]$ solves $A Y + Y A = X$. The geodesics, exponential and logarithm maps, and explicit positive curvature properties are available [2012.07106]. The Bures–Wasserstein geometry, further generalized as GBW, introduces a metric tensor parameterized by $M \in S_{++}^n$, leading to geodesics and distances incorporating a Mahalanobis-type precision weighting. This anisotropic structure permits improved statistical efficiency and conditioning, and is generalized to infinite-dimensional and operator-valued settings—e.g., covariance operators acting on Hilbert spaces—via unitized Hilbert–Schmidt operators and an extended Mahalanobis norm. Operator-valued Procrustes geodesics, learnable regularization parameters, and tractable computational schemes for high-dimensional inference are provided [2110.10464, 2511.09801].

## 7. Geometry on Special Structures and Embeddings

In ultrametric spaces, the $L^p$ Wasserstein geometry collapses to an affine form, admitting an isometric embedding into a convex subset of an $\ell^1$ Banach space. Geodesics exist only for $p=1$, and otherwise connectivity is via Hölder arcs of exponent $1/p$ [1304.5219]. For time series and signed measures, the generalized Wasserstein geometry leverages Jordan decompositions and signed Cumulative Distribution Transforms (SCDT), embedding signals into a flat Hilbert space $(L^2 \times \mathbb R)^2$ and providing interpretability and straight-line geodesics in classes generated by template deformations [2206.01984].

## 8. Outlook and Applications

Generalized Wasserstein geometries enable the handling of mass transfer beyond strict conservation, greater expressivity in comparing complex distributions (e.g., via nonlinear projections or matrix-valued data), and robust analysis in non-Euclidean, infinite-dimensional, or branching settings. They are now central to algorithm design in scalable transport, functional data analysis, statistical learning with manifold constraints, configuration and Wiener spaces, and information-geometric optimization. Recent advances unify these constructions into a broader synthetic and categorical framework, facilitating further extension to abstract settings while retaining the core variational, geometric, and computational underpinnings of optimal transport [1304.7014, 2412.01190, 2302.05833, 2511.09801, 2505.22049, 2210.10268].

Source: https://www.emergentmind.com/topics/generalized-wasserstein-geometries