---
title: Wasserstein Distance Indices
url: https://www.emergentmind.com/topics/wasserstein-distance-indices
type: topic
---

# Wasserstein Distance Indices

Wasserstein distance indices are quantitative summaries and derived metrics built on the Wasserstein transport distance, widely used to compare and analyze probability measures in mathematical statistics, optimal transport, computer science, and related fields. These indices include test statistics, proxies, dependence coefficients, and efficient computational surrogates, each exploiting the geometric, statistical, or operational structure of the Wasserstein metric. This article surveys core theoretical formulations, relaxations, surrogate indices, statistical properties, practical applications, and their implementation as developed in the literature.

## 1. Foundational Definitions and Theoretical Basis

The $p$-Wasserstein distance on a Polish metric space $(\mathcal X, d)$ between measures $\mu,\nu$ with finite $p$th moments is defined by
\[
W_p(\mu,\nu) = \left(\inf_{\gamma\in \Gamma(\mu,\nu)}\int_{\mathcal X\times \mathcal X} d(x, y)^p\, d\gamma(x, y)\right)^{1/p},
\]
where $\Gamma(\mu,\nu)$ denotes the set of couplings of $\mu$ and $\nu$. For $p=1$, the dual formulation is given by
\[
W_1(\mu, \nu) = \sup_{f \in \operatorname{Lip}_1(\mathcal X)} \left\{\int f\, d\mu - \int f\, d\nu\right\},
\]
with $\operatorname{Lip}_1(\mathcal X)$ the class of 1-Lipschitz functions. Fundamental properties include nonnegativity, symmetry, the triangle inequality, and metrization of weak convergence plus moment convergence for the probability space $\mathcal{W}_p(\mathcal X)$ [1806.05500].

Wasserstein distances seamlessly incorporate the geometry of the domain, which is central to their utility as indices in shape analysis, imaging, and probabilistic modeling.

## 2. Statistical Wasserstein Indices: Estimation and Testing

Wasserstein distances play a pivotal role as statistical indices for hypothesis testing, clustering, and model evaluation:

### 2.1. Empirical Estimation  
The plug-in estimator $W_p(\mu_n, \nu_n)$, where $\mu_n, \nu_n$ are empirical measures from i.i.d. samples, satisfies $W_p(\mu_n,\nu_n)\to W_p(\mu,\nu)$ almost surely (rates are dimension-dependent; $O(n^{-1/d})$ for $d > 2p$ and $O(n^{-1/2})$ for discrete/finitely supported laws) [1806.05500].

### 2.2. Goodness-of-fit and Clustering  
- **Goodness-of-fit**: $T_n = W_p(\mu_n, \mu_0)$ is used for one-sample tests.
- **Two-sample testing**: $T_{n,m} = W_p(\mu_n, \nu_m)$ serves as a nonparametric test statistic.
- **Clustering**: Minimizing within-cluster Wasserstein dispersion yields robust clustering procedures [1806.05500].

### 2.3. Disparity and Inequality  
The Wasserstein–Gini index is defined via the $L^1$ distance between quantile functions: $G_W(\mu) = 2\int_{0 < u < v < 1}|F^{-1}(u) - F^{-1}(v)|\, du\, dv$ [1806.05500].

### 2.4. Explicit Null-Model Indices  
On the simplex $\Omega_n$, with $W_1$ computed via the cumulative coordinate trick, exact closed-form moments are available:
\[
\mathbb{E}[W_1] = \frac{2^{2n-3}(n-1)}{(2n-1)!}(n-1)!^2, \qquad \mathbb{E}[W_1^2] = \frac{(n-1)(7n-4)}{30n},
\]
characterizing the typical scale and variability when comparing random discrete distributions [1912.04945].

## 3. Relaxation, Surrogate, and Proxy Indices

To address computational or statistical challenges, several effective indices serve as proxies for exact Wasserstein distances.

### 3.1. Sliced and Projected Indices  
- **Sliced Wasserstein** ($SW_p$): Defined as $SW_p(P,Q) = \int_{S^{d-1}} W_p(\theta^*_{\#}P, \theta^*_{\#}Q) d\sigma(\theta)$, enjoying dimension-free sample complexity and sub-Gaussian concentration [2205.14624].
- **Maximum Sliced Wasserstein** ($MSW_p$): $MSW_1(P,Q) = \sup_{\theta \in S^{d-1}} W_1(\theta^*_{\#}P, \theta^*_{\#}Q)$, with uniform tail bounds and nonparametric Donsker-theorem-based CLTs [2205.14624].
- **Observable Wasserstein Distance**: $\theta_p(\mu, \nu) = \sup_{f \in \operatorname{Lip}_1(X)} w_p(f_{\#}\mu, f_{\#}\nu)$, with computationally tractable lower-bounding pseudo-metrics $\theta_{p, n}$ given by maximizing over a finite anchor-set hierarchy. These proxies form a tunable hierarchy trading sharpness for efficiency [2605.09916].
- **Distance-Matrix Wasserstein** ($\mathrm{DMW}_{n,p}$): For metric-measure spaces, compares the distribution of finite random distance matrices; $\mathrm{DMW}_{n,p}(X, Y)\leq \mathrm{GW}_p(X, Y)$ (GW = Gromov–Wasserstein), converging to GW as $n\to\infty$, with tight finite-sample and dimension-adaptive bounds [2605.14981].

### 3.2. Proxies via Discrepancies and KS Distance  
On $[0,1]^d$, $W_p$ can be sharply upper bounded by powers of box discrepancies:
\[
W_p(\mu, \nu) \leq K_{p, d} D(\mu, \nu)^{1/\max(d,p)},
\]
where $D(\mu, \nu)$ is the uniform box discrepancy; or, in $\mathbb{R}^d$, by KS distance-weighted moments:
\[
W_p(\mu, \nu) \leq C_{p,d} \left(m_p(\mu)+m_p(\nu)\right)^{1/p} d_{KS}(\mu, \nu),
\]
offering practical computable indices for model assessment and sampling [2605.03528].

## 4. Indices in Complex Structures: Dependence and Graphs

### 4.1. Dependence Indices  
- **Wasserstein Correlation Coefficient**: For a coupling $\pi$ between $\mu,\nu$,
\[
\overrightarrow{\mathcal{W}}(\pi) = \frac{\int W(\pi_{x_1}, \nu)\, \mu(dx_1)}{\iint d(y, z) \nu(dy)\nu(dz)},
\]
which is $0$ at independence ($\pi = \mu \otimes \nu$), $1$ under functional dependence, convex in $\pi$, and admits nonparametric plug-in estimation and independence testing procedures [2102.00356].  
- **Wasserstein Index of Dependence** for random measures: $I_W(\mathcal{L}) = 1-W_*^2(\nu, \nu)/(W_*^2(\nu^\perp, \nu))$, quantifies the dependence structure within vectors of completely random measures, normalizing the Wasserstein distance between full dependence and independence. This index is intrinsically multivariate, nonpairwise, numerically tractable, and directly enables prior specification and model selection in Bayesian nonparametrics [2109.06646].

### 4.2. Graph, Matrix, and Kernel-Induced Indices  
- **Graph Distance via GMMs**: Nodes of graphs are mapped to (probabilistic) embeddings, fitted as Gaussian mixtures; the resulting Wasserstein distance between mixtures quantifies graph dissimilarity and admits fast closed-form or OT-based computation under tied/diagonal covariance structure, scaling to large graphs [2401.03913].
- **Tree-Wasserstein Distance (Supervised)**: For document or feature-collection comparison, fast computation uses parent–child summation formulas on trees, with end-to-end differentiable relaxations for metric learning, outperforming exact OT on large corpora [2101.11520].
- **Quasi-Manhattan Wasserstein Distance**: Combines three 1D Wasserstein computations on linearized, rotated, and transposed matrix representations, providing a linear-time, sub-5% error proxy for high-dimensional matrix data [2310.12498].
- **Kernel Wasserstein**: Embeds measures into a reproducing kernel Hilbert space; $W_2^2(P, Q) = \|\mu_P - \mu_Q\|_H^2$ serves as a fast, closed-form Wasserstein-type index, enabling robust distributional comparisons in applications like imaging and anomaly detection [1905.09314].

## 5. Generalizations, Unbalanced Cases, and Dualities

Classical Wasserstein distances are only defined for equal total mass. The generalized Wasserstein distance
\[
W_p^{a, b}(\mu, \nu) = \left(\inf_{\substack{\mu' \leq \mu,\, \nu' \leq \nu}} a^p(|\mu - \mu'| + |\nu - \nu'|)^p + b^p W_p^p(\mu', \nu')\right)^{1/p}
\]
admits source terms (mass creation/destruction) at cost $a$, and transport at cost $b$, remaining a metric for all nonnegative measures. In special cases, e.g. $W_1^{1,1}$, it coincides with the flat (bounded–Lipschitz) distance, providing a dual characterization built on test functions bounded both in norm and Lipschitz seminorm [1304.7014].

The corresponding dynamic (Benamou–Brenier) formulations include penalties for both kinetic energy and total source mass, extending applicability to unbalanced data and source-driven PDEs.

## 6. Numerical Methods, Concentration, and Algorithmic Aspects

Efficient computation of Wasserstein indices leverages diverse algorithmic approaches:
- **Exact OT**: Solved via linear programming or the Sinkhorn (entropic regularization) algorithm, with respective $O(n^3)$ and $O(n^2/\epsilon^2)$ complexity for $n$ support points.
- **Sliced, Observable, and DMW**: Lower computational cost via random projections or anchor-based subspaces; error decays as $O(1/\sqrt{L})$ for $L$ projections.
- **Tree/Supervised/KD-tree/Quasi-linear proxies**: For document, image, or graph comparisons, sacrifices exactness for scalability with proven $<5\%$ or empirical error guarantees [2101.11520, 2310.12498, 2605.14981].  
- **Statistical consistency**: Many indices admit nonparametric CLTs, sub-Gaussian tail inequalities, and explicit rate guarantees; rates can be optimal or dimension-dependent, highlighting trade-offs for high-dimensional data [2205.14624, 2605.09916].

## 7. Applications and Impact in Practice

Wasserstein distance indices are pervasive across theoretical and applied contexts:
- **Hypothesis testing**: Two-sample and independence tests in high-dimensional and complex domains [1806.05500, 2102.00356].
- **Fair benchmarking**: Using scaled or normalized Wasserstein indices as null distributions or performance thresholds for random or expected scenarios [1912.04945].
- **Model selection**: Informativity and dependence control in structured Bayesian models via dependence indices [2109.06646].
- **Computational geometry and imaging**: Proxies (sliced, observable, kernel) drive applications in anomaly detection, image retrieval, generative modeling, and shape analysis [2605.09916, 1905.09314].
- **Numerical methods and QMC**: Proxy indices facilitate efficient error control in sampling, approximation, and integration, e.g., via Proinov’s theorem or box-discrepancy bounds [2605.03528].

In sum, Wasserstein distance indices, in their many forms, provide a fundamental and flexible toolbox for quantifying, comparing, and understanding probability laws, structures, and models in modern applied mathematics and statistics. Their evolving proxies, computational relaxations, and derivative indices address the inherent complexity of the metric, enabling robust analysis, inference, and decision-making across diverse domains.

Source: https://www.emergentmind.com/topics/wasserstein-distance-indices