Papers
Topics
Authors
Recent
Search
2000 character limit reached

Concentration Inequalities in Wasserstein Distance

Updated 23 January 2026
  • Wasserstein concentration inequalities are statistical bounds that quantify how empirical distributions deviate from underlying laws using transport metrics.
  • They leverage transport–entropy and metric entropy techniques to derive sub-Gaussian tail bounds and optimal finite-sample rates across various regimes.
  • Applications span statistical learning, random matrix theory, quantum information, and stochastic processes, providing rigorous error guarantees in complex systems.

Concentration inequalities in Wasserstein distance quantify the probability that random probability measures—typically empirical distributions arising from finite samples—deviate from their underlying population law, with the deviation measured in the Wasserstein metric. This framework, originating in the study of the “concentration of measure” phenomenon, underpins theoretical guarantees and finite-sample error control across probability, statistics, learning theory, random matrix theory, quantum information, and statistical physics. Multiple regimes, model classes, and metrics (including W1W_1, WpW_p for p1p\geq 1, sliced/projected Wasserstein, and quantum generalizations) exhibit sharp non-asymptotic bounds, which can depend on ambient or intrinsic dimension, support geometry, tail behavior, and process structure.

1. Definitions and Frameworks

The pp–Wasserstein distance between probability measures μ\mu and ν\nu on a Polish metric space (E,d)(E, d) is defined as

Wp(μ,ν):=(infπC(μ,ν)E×Ed(x,y)pπ(dx,dy))1/pW_p(\mu, \nu) := \left( \inf_{\pi \in \mathcal{C}(\mu, \nu)} \int_{E \times E} d(x, y)^p\, \pi(dx, dy) \right)^{1/p}

for p1p \geq 1, where C(μ,ν)\mathcal{C}(\mu, \nu) denotes the set of couplings of WpW_p0 and WpW_p1 (Fournier et al., 2013, Dedecker et al., 16 Jan 2026, Chafai et al., 2016). The WpW_p2-Wasserstein admits a dual formulation via the Kantorovich–Rubinstein theorem, making it an integral metric over 1-Lipschitz functions.

Transport–entropy inequalities, such as the WpW_p3 (Talagrand) inequality:

WpW_p4

where WpW_p5 denotes relative entropy, serve as the pivot for deriving measure concentration—sub-Gaussian tail bounds for Lipschitz observables under WpW_p6 (Khoshnevisan et al., 2017, Boissard, 2011, Park, 25 Jul 2025).

2. Classical Concentration Bounds: Rates, Regimes, and Dimensionality

For the empirical measure WpW_p7 from i.i.d. samples WpW_p8, concentration inequalities in WpW_p9 fall into several regimes, with rates determined by moment conditions and (ambient or intrinsic) dimension (Fournier et al., 2013, Dedecker et al., 16 Jan 2026, Lei, 2018):

  • Sub-Gaussian regime (p1p\geq 10): p1p\geq 11 up to logarithmic factors.
  • Critical regime (p1p\geq 12): p1p\geq 13.
  • Curse-of-dimensionality regime (p1p\geq 14): p1p\geq 15 (the classical quantization rate).

High-probability (tail) inequalities mirror these rates, yielding

p1p\geq 16

where the exponent p1p\geq 17 interpolates between p1p\geq 18 and p1p\geq 19 depending on pp0 and pp1 (Fournier et al., 2013, Chafai et al., 2016).

By refining metric entropy arguments, these bounds extend to intrinsic (covering/Hausdorff) dimension pp2, so that—for empirical measures supported on sets with covering number pp3—the same pp4 rates and corresponding concentration inequalities hold (Dedecker et al., 16 Jan 2026).

3. Concentration Under Functional and Transport-Entropy Inequalities

When the law pp5 satisfies a transport–entropy inequality (classically pp6 or pp7), exponential concentration of pp8 arises. For example, Boissard (Boissard, 2011) proves:

pp9

assuming only μ\mu0 and exponential integrability of μ\mu1. The constant μ\mu2 is explicitly related to the transport-entropy constant.

For measures on bounded domains or with strong exponential tails, the bound holds globally; for more general settings, additional double-exponential prefactors may enter but the sub-Gaussian exponent in μ\mu3 persists.

Tensorization arguments and Laplace functional techniques (Herbst's argument) link μ\mu4 inequalities to concentration of 1-Lipschitz functionals and empirical measures. Such strategies underpin many advanced bounds, including those for Gaussian, product, and Markov chain measures (Park, 25 Jul 2025, Boissard, 2011, Barbour et al., 2019).

4. Advanced Variants: High Dimension, Intrinsic Geometry, and Non-Euclidean/Wasserstein Variants

Several refinements and variants address settings where classical bounds are suboptimal.

  • Intrinsic Dimension: For measures supported on lower-dimensional (e.g., μ\mu5-dim Riemannian manifold, fractal, or with covering dimension μ\mu6), μ\mu7 rates are sharp, and all concentration regimes (small/large deviation, moderate/large deviations, almost-sure convergence) persist with μ\mu8 replacing μ\mu9 (Dedecker et al., 16 Jan 2026).
  • Projected/Sliced Wasserstein: Projected or sliced Wasserstein distances bypass curse-of-dimensionality rates; for instance, for the Sliced ν\nu0, one has ν\nu1 concentration with dimension-independent exponents under second moment assumptions (Xu et al., 2022, Wang et al., 2020). Projected Wasserstein distances in ν\nu2-dimensional subspaces allow interpolation between high-dimensional and low-dimensional rates, with explicit trade-offs between ν\nu3 and ν\nu4.
  • Occupation Measures and Markov Chains: For ergodic Markov chains with contractivity in Wasserstein, empirical laws concentrate sharply about the invariant distribution, with the contraction rate Propagating directly into sub-Gaussian/Poissonian tail bounds (Barbour et al., 2019, Boissard, 2011).
  • Infinite-Dimensional/Functional Data: Extension to Banach/Hilbert space-valued data is accomplished via telescoping block decompositions and hierarchical coupling, yielding tight rates for functional classes with polynomial or exponential decay (Lei, 2018). For Gaussian processes/ellipsoidal moment classes and their empirical measures, the mean ν\nu5 bias decays at rates determined by the coordinate decay.
  • Quantum Wasserstein: Quantum Markov semigroups admit analogues of classical ν\nu6, ν\nu7, logarithmic Sobolev, and Poincaré inequalities in the quantum setting, with the corresponding concentration bounds for quantum states (e.g., depolarizing semigroup) controlled via quantum Wasserstein metrics and noncommutative Lipschitz norms (Rouzé et al., 2017).

5. Functional Inequalities, Stein Discrepancy, and Information Geometry

Improved concentration inequalities, which relate entropy, Fisher information, and Stein discrepancy to Wasserstein distance, have been established—e.g., the HSI (entropy–Stein discrepancy–information) and WS (Wasserstein–Stein) inequalities (Cheng et al., 2021). These can yield strictly sharper bounds compared to classical Talagrand/log-Sobolev inequalities:

  • For a measure ν\nu8 on a Riemannian manifold ν\nu9,

(E,d)(E, d)0

where (E,d)(E, d)1 is the Stein discrepancy and (E,d)(E, d)2 encodes curvature. Additional HWSI inequalities improve upon (E,d)(E, d)3 (Talagrand) by exploiting the nontrivial geometry of (E,d)(E, d)4 and (E,d)(E, d)5 (Cheng et al., 2021).

6. Applications in Statistical Learning, High-Dimensional Inference, and Random Matrices

  • Statistical Learning: Concentration in Wasserstein is foundational to statistical consistency and finite-sample precision of learning algorithms based on integral probability metrics (e.g., WGANs), with generalization bounds scaling as either (E,d)(E, d)6 (bounded metric-entropy), (E,d)(E, d)7 (finite-moment), and associated exponential tails controlled by Rademacher complexity (Birrell, 2024).
  • Gaussian Approximation: Recent advances use Stein’s method and exchangeable pairs to produce computable, non-asymptotic (E,d)(E, d)8 bounds between sample mean and its Gaussian target, achieving sub-Gaussian tails and optimal (E,d)(E, d)9 rates uniformly in Wp(μ,ν):=(infπC(μ,ν)E×Ed(x,y)pπ(dx,dy))1/pW_p(\mu, \nu) := \left( \inf_{\pi \in \mathcal{C}(\mu, \nu)} \int_{E \times E} d(x, y)^p\, \pi(dx, dy) \right)^{1/p}0 (Austern et al., 2022).
  • Random Matrix Theory and Coulomb Gases: In multi-particle Coulomb systems, the empirical spectral law exhibits sub-Gaussian concentration in Wp(μ,ν):=(infπC(μ,ν)E×Ed(x,y)pπ(dx,dy))1/pW_p(\mu, \nu) := \left( \inf_{\pi \in \mathcal{C}(\mu, \nu)} \int_{E \times E} d(x, y)^p\, \pi(dx, dy) \right)^{1/p}1 at rate Wp(μ,ν):=(infπC(μ,ν)E×Ed(x,y)pπ(dx,dy))1/pW_p(\mu, \nu) := \left( \inf_{\pi \in \mathcal{C}(\mu, \nu)} \int_{E \times E} d(x, y)^p\, \pi(dx, dy) \right)^{1/p}2, matching nonasymptotic large deviations and improving earlier results with Wp(μ,ν):=(infπC(μ,ν)E×Ed(x,y)pπ(dx,dy))1/pW_p(\mu, \nu) := \left( \inf_{\pi \in \mathcal{C}(\mu, \nu)} \int_{E \times E} d(x, y)^p\, \pi(dx, dy) \right)^{1/p}3 scaling (Chafai et al., 2016).

7. Extensions, Limitations, and Future Directions

  • Infinite-Dimensional Processes and SPDEs: Extension of Wp(μ,ν):=(infπC(μ,ν)E×Ed(x,y)pπ(dx,dy))1/pW_p(\mu, \nu) := \left( \inf_{\pi \in \mathcal{C}(\mu, \nu)} \int_{E \times E} d(x, y)^p\, \pi(dx, dy) \right)^{1/p}4 inequalities and concentration in Wasserstein to measure-valued laws of SPDEs is established for the 1D parabolic case with space-time white noise. Coupling and Girsanov techniques replace classical log-Sobolev functional arguments (Khoshnevisan et al., 2017).
  • Quantum and Noncommutative Regimes: Quantum analogues of Wasserstein distance, transport, and concentration inequalities rely on recent noncommutative metric and entropy constructs, and are active areas of investigation for quantum state tomography and parameter estimation (Rouzé et al., 2017).
  • Adapted Wasserstein and Stochastic Processes: For discrete-time stochastic processes, adapted Wasserstein distances and their transport-entropy inequalities extend concentration to path-space, with process-level causal constraints resulting in optimal Wp(μ,ν):=(infπC(μ,ν)E×Ed(x,y)pπ(dx,dy))1/pW_p(\mu, \nu) := \left( \inf_{\pi \in \mathcal{C}(\mu, \nu)} \int_{E \times E} d(x, y)^p\, \pi(dx, dy) \right)^{1/p}5 dependence on time horizon (Park, 25 Jul 2025).

A plausible implication is that concentration in Wasserstein distance—when properly localized to intrinsic geometry, support regularity, or process structure—achieves sub-Gaussian or optimal sample-complexity rates in a diverse array of models, encompassing both classical, modern high-dimensional, quantum, and infinite-dimensional regimes. The machinery is thus central to understanding the behavior of empirical measures, statistical estimators, Markov processes, and many-body systems across mathematical and applied disciplines.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Concentration Inequalities in Wasserstein Distance.