Papers
Topics
Authors
Recent
Search
2000 character limit reached

Wasserstein Metric Certificates

Updated 10 July 2026
  • Wasserstein metric certificates are rigorous guarantees derived from Wasserstein geometry that certify optimal values, worst-case risks, convergence, and structural relations across diverse settings.
  • They cover a range of applications including statistical minimax risk estimation, explicit analytic evaluations, variational dynamics in gradient flows, and robust adversarial defenses in deep networks.
  • Combining geometric, analytic, variational, and information-theoretic methods, these certificates provide practical and actionable insights into estimation accuracy and performance bounds.

Across several strands of optimal transport, statistics, information geometry, and adversarial ML, Wasserstein metric certificates are rigorous statements expressed directly in Wasserstein geometry that verify an optimal value, a worst-case risk, a convergence property, an efficiency bound, or a structural relation between measures. In the literature, the term appears in multiple concrete senses: minimax certificates for distribution estimation under Wasserstein loss, analytic certificates giving explicit formulas for WpW_p, variational certificates for gradient-flow dynamics, information-theoretic certificates such as Wasserstein–Cramér–Rao bounds, and robustness certificates for perturbations constrained by Wasserstein balls (Singh et al., 2018, Abdellatif et al., 2023, Li et al., 2019, Levine et al., 2019).

1. Statistical minimax certificates

A central statistical use of Wasserstein certificates is to characterize the best achievable estimation error under Wasserstein loss. For distribution estimation from nn IID samples on a metric space (Ω,ρ)(\Omega,\rho), the minimax risk is

M(P,r)=infP^supPPEX1nP ⁣[Wrr ⁣(P,P^(X1n))].M(\mathcal{P}, r) = \inf_{\hat P} \sup_{P \in \mathcal{P}} \mathbb{E}_{X_1^n \sim P}\!\left[ W_r^r\!\big(P,\hat P(X_1^n)\big) \right].

The analysis in "Minimax Distribution Estimation in Wasserstein Distance" identifies the governing geometric quantities as covering numbers N(Ω,ϵ)N(\Omega,\epsilon), packing numbers M(Ω,ϵ)M(\Omega,\epsilon), and the packing radius R(Ω,n)R(\Omega,n), together with the moment condition

m,x0(P)=(EYP[ρ(x0,Y)])1/μ,r.m_{\ell,x_0}(P) = \left(\mathbb{E}_{Y\sim P}[\rho(x_0,Y)^\ell]\right)^{1/\ell} \le \mu, \qquad \ell \ge r.

The empirical distribution PnP_n satisfies a multiscale upper bound, while lower bounds are obtained from packing arguments and heavy-tail constructions; in regular settings these bounds match up to constants and sometimes log-factors (Singh et al., 2018).

The resulting certificates are both geometric and probabilistic. In bounded spaces with N(Ω,ϵ)ϵsN(\Omega,\epsilon)\asymp \epsilon^{-s}, the rate simplifies to nn0, with nn1 acting as an intrinsic dimension. In unbounded spaces, finite-nn2 moments impose the rate limitation nn3. The lower bound

nn4

certifies that the packing structure of the sample space itself limits estimation accuracy, while the heavy-tail lower bound

nn5

certifies unavoidable bias from mass at large radius. Under only these geometric and moment assumptions, the empirical measure is minimax rate-optimal, so more elaborate estimators do not improve the Wasserstein risk unless stronger assumptions are imposed (Singh et al., 2018).

A closely related certificate concerns estimation of the metric itself rather than the underlying distributions. "On the Minimax Optimality of Estimating the Wasserstein Metric" proves that estimating nn6 between two unknown measures on nn7 is not significantly easier than estimating the measures under Wasserstein loss. For nn8-Hölder densities on nn9, (Ω,ρ)(\Omega,\rho)0, the minimax rate is

(Ω,ρ)(\Omega,\rho)1

and the smoothed plug-in estimator is minimax optimal up to logarithmic factors (1908.10324). A frequent misconception is therefore that functional estimation of (Ω,ρ)(\Omega,\rho)2 should be substantially easier than distribution estimation; the cited result rules this out in the stated smooth setting.

2. Analytic certificates for explicit Wasserstein values

Another meaning of certificate is an explicit analytic representation of the optimal transport value. In one dimension, the Wasserstein distance admits such a certificate through the comonotonicity copula (Ω,ρ)(\Omega,\rho)3. For probability measures (Ω,ρ)(\Omega,\rho)4 on (Ω,ρ)(\Omega,\rho)5 with CDFs (Ω,ρ)(\Omega,\rho)6 and quantile functions (Ω,ρ)(\Omega,\rho)7,

(Ω,ρ)(\Omega,\rho)8

and for (Ω,ρ)(\Omega,\rho)9,

M(P,r)=infP^supPPEX1nP ⁣[Wrr ⁣(P,P^(X1n))].M(\mathcal{P}, r) = \inf_{\hat P} \sup_{P \in \mathcal{P}} \mathbb{E}_{X_1^n \sim P}\!\left[ W_r^r\!\big(P,\hat P(X_1^n)\big) \right].0

These formulas certify the optimal value directly, because the infimum is realized by the comonotonic coupling (Abdellatif et al., 2023).

The multivariate case is substantially more restrictive. An explicit quantile-based formula in M(P,r)=infP^supPPEX1nP ⁣[Wrr ⁣(P,P^(X1n))].M(\mathcal{P}, r) = \inf_{\hat P} \sup_{P \in \mathcal{P}} \mathbb{E}_{X_1^n \sim P}\!\left[ W_r^r\!\big(P,\hat P(X_1^n)\big) \right].1 is available only when the two measures share the same copula. If both have the comonotonicity copula M(P,r)=infP^supPPEX1nP ⁣[Wrr ⁣(P,P^(X1n))].M(\mathcal{P}, r) = \inf_{\hat P} \sup_{P \in \mathcal{P}} \mathbb{E}_{X_1^n \sim P}\!\left[ W_r^r\!\big(P,\hat P(X_1^n)\big) \right].2, then

M(P,r)=infP^supPPEX1nP ⁣[Wrr ⁣(P,P^(X1n))].M(\mathcal{P}, r) = \inf_{\hat P} \sup_{P \in \mathcal{P}} \mathbb{E}_{X_1^n \sim P}\!\left[ W_r^r\!\big(P,\hat P(X_1^n)\big) \right].3

If they share a general copula M(P,r)=infP^supPPEX1nP ⁣[Wrr ⁣(P,P^(X1n))].M(\mathcal{P}, r) = \inf_{\hat P} \sup_{P \in \mathcal{P}} \mathbb{E}_{X_1^n \sim P}\!\left[ W_r^r\!\big(P,\hat P(X_1^n)\big) \right].4, then

M(P,r)=infP^supPPEX1nP ⁣[Wrr ⁣(P,P^(X1n))].M(\mathcal{P}, r) = \inf_{\hat P} \sup_{P \in \mathcal{P}} \mathbb{E}_{X_1^n \sim P}\!\left[ W_r^r\!\big(P,\hat P(X_1^n)\big) \right].5

The one-dimensional integral reduction is therefore not available for arbitrary multivariate dependence structures. This restriction is one of the main clarifications of the copula-based treatment: explicit certificates exist, but only under matched dependence, and especially in the comonotonic case (Abdellatif et al., 2023).

3. Variational and dynamical certificates

In Wasserstein gradient-flow theory, certificates are variational identities and inequalities that characterize discrete and continuous evolution. "The Exponential Formula for the Wasserstein Metric" develops the M(P,r)=infP^supPPEX1nP ⁣[Wrr ⁣(P,P^(X1n))].M(\mathcal{P}, r) = \inf_{\hat P} \sup_{P \in \mathcal{P}} \mathbb{E}_{X_1^n \sim P}\!\left[ W_r^r\!\big(P,\hat P(X_1^n)\big) \right].6-transport metric

M(P,r)=infP^supPPEX1nP ⁣[Wrr ⁣(P,P^(X1n))].M(\mathcal{P}, r) = \inf_{\hat P} \sup_{P \in \mathcal{P}} \mathbb{E}_{X_1^n \sim P}\!\left[ W_r^r\!\big(P,\hat P(X_1^n)\big) \right].7

under which squared distance is genuinely convex along generalized geodesics. This enables a Hilbertian-style analysis of the JKO proximal map

M(P,r)=infP^supPPEX1nP ⁣[Wrr ⁣(P,P^(X1n))].M(\mathcal{P}, r) = \inf_{\hat P} \sup_{P \in \mathcal{P}} \mathbb{E}_{X_1^n \sim P}\!\left[ W_r^r\!\big(P,\hat P(X_1^n)\big) \right].8

The Euler–Lagrange condition

M(P,r)=infP^supPPEX1nP ⁣[Wrr ⁣(P,P^(X1n))].M(\mathcal{P}, r) = \inf_{\hat P} \sup_{P \in \mathcal{P}} \mathbb{E}_{X_1^n \sim P}\!\left[ W_r^r\!\big(P,\hat P(X_1^n)\big) \right].9

acts as a local certificate that a measure is a JKO proximal point, while the exponential formula

N(Ω,ϵ)N(\Omega,\epsilon)0

certifies convergence of the discrete scheme to the gradient flow (Craig, 2013).

The same framework yields dynamical guarantees in the form of a semigroup contraction property and an energy dissipation inequality: N(Ω,ϵ)N(\Omega,\epsilon)1

N(Ω,ϵ)N(\Omega,\epsilon)2

These are certificates of stability and monotone energy decay for Wasserstein flows (Craig, 2013).

Distributed dynamics in Wasserstein space admit analogous certification. In "Network Consensus in the Wasserstein Metric Space of Probability Measures," each agent updates by computing a local Wasserstein barycenter,

N(Ω,ϵ)N(\Omega,\epsilon)3

Under weak joint connectivity, all agents converge to a common measure N(Ω,ϵ)N(\Omega,\epsilon)4, the weighted Wasserstein barycenter of the initial measures. In time-invariant connected networks, convergence is exponential; on the real line for N(Ω,ϵ)N(\Omega,\epsilon)5, the inverse CDFs satisfy a linear consensus recursion,

N(Ω,ϵ)N(\Omega,\epsilon)6

These results make the barycenter itself the certified consensus value (Bishop et al., 2014).

4. Information-geometric certificates

The Wasserstein information matrix (WIM) extends classical information geometry to the N(Ω,ϵ)N(\Omega,\epsilon)7-Wasserstein setting. For a parametric family N(Ω,ϵ)N(\Omega,\epsilon)8, the WIM is the pullback of the Wasserstein metric tensor N(Ω,ϵ)N(\Omega,\epsilon)9, with

M(Ω,ϵ)M(\Omega,\epsilon)0

The associated Wasserstein score functions

M(Ω,ϵ)M(\Omega,\epsilon)1

solve a Poisson equation and satisfy M(Ω,ϵ)M(\Omega,\epsilon)2. In one dimension,

M(Ω,ϵ)M(\Omega,\epsilon)3

This construction defines the metric certificate directly on parameter space rather than through the Fisher–Rao geometry (Li et al., 2019).

The principal inferential certificate is the Wasserstein–Cramér–Rao bound

M(Ω,ϵ)M(\Omega,\epsilon)4

where Wasserstein covariance is defined by

M(Ω,ϵ)M(\Omega,\epsilon)5

A statistic is Wasserstein-efficient if and only if it is a linear combination of Wasserstein score functions. The same framework yields asymptotic efficiency statements for Wasserstein natural gradient and a Poincaré efficiency for Wasserstein natural gradient of maximal likelihood estimation (Li et al., 2019).

The WIM is especially relevant where classical Fisher information is undefined or poorly adapted to the geometry. The cited analytical examples include location-scale families, independent families, and ReLU generative models. For the Gaussian location-scale family, the WIM is the identity matrix, while for the uniform and ReLU generative examples the paper emphasizes that the Fisher information does not exist but the WIM does. In this sense, the WIM functions as a metric certificate for models with singular or nonclassical transport geometry (Li et al., 2019).

5. Concave certificates for distributionally robust risk and complexity

In distributionally robust optimization, a Wasserstein certificate upper-bounds the worst-case expected loss over a Wasserstein uncertainty set. "Concave Certificates: Geometric Framework for Distributionally Robust Risk and Complexity Analysis" studies

M(Ω,ϵ)M(\Omega,\epsilon)6

and replaces global Lipschitz or first-order local certificates with a certificate based on the least concave majorant of the empirical maximal growth rate. The framework is designed for non-Lipschitz, non-differentiable, non-convex, and even unbounded losses, provided the majorant is finite. The main theorem bounds the robust risk increment by the concave certificate M(Ω,ϵ)M(\Omega,\epsilon)7, defined from the least concave majorant of the maximal empirical rate evaluated at scale M(Ω,ϵ)M(\Omega,\epsilon)8 (Chu, 4 Jan 2026).

This certificate is geometric rather than purely differential. It adapts to the actual growth rate of the loss under perturbation and reduces to gradient-based certificates for small M(Ω,ϵ)M(\Omega,\epsilon)9 when differentiability holds, but remains valid globally for any R(Ω,n)R(\Omega,n)0. The same paper introduces concave complexity,

R(Ω,n)R(\Omega,n)1

which provides a deterministic generalization bound, and proves that the gap between adversarial and empirical Rademacher complexity is controlled by concave complexity. The stated significance is that dependencies on input diameter, network width, and depth can be eliminated in the resulting bound (Chu, 4 Jan 2026).

For deep networks, exact majorants are replaced by the adversarial score, a tractable compositional upper bound computed layerwise. The paper reports that on traffic regression the adversarial score gave strictly tighter and less volatile robustness certificates than gradient and Lipschitz certificates, and that on MNIST the empirical adversarial–empirical complexity gap was dimension-free. A common misunderstanding is that Wasserstein DR certification must be tied either to a global Lipschitz constant or to differentiability; the concave-certificate framework is explicitly constructed to avoid both requirements (Chu, 4 Jan 2026).

6. Adversarial robustness certificates under Wasserstein threat models

For image classification, Wasserstein robustness certificates convert transport-based perturbation sets into more tractable domains. "Wasserstein Smoothing: Certified Robustness against Wasserstein Adversarial Attacks" constructs the first certifiable defense against Wasserstein adversarial attacks by representing images in a local flow space. If R(Ω,n)R(\Omega,n)2 is obtained by a local flow plan R(Ω,n)R(\Omega,n)3, then

R(Ω,n)R(\Omega,n)4

Randomized smoothing is then applied in flow space rather than pixel space, with Laplace noise on flow coordinates. Because Wasserstein distance is upper-bounded by R(Ω,n)R(\Omega,n)5 distance in that representation, existing R(Ω,n)R(\Omega,n)6 randomized smoothing certificates become Wasserstein certificates. The reported guarantee is that if the smoothed class score for the correct class dominates by the prescribed factor, then every image within the certified R(Ω,n)R(\Omega,n)7 radius is assigned the same class (Levine et al., 2019).

A more general verification framework is developed in "A Framework for Verification of Wasserstein Adversarial Robustness." There, Wasserstein balls in image space are mapped to R(Ω,n)R(\Omega,n)8-balls in a flow domain relative to a reference distribution R(Ω,n)R(\Omega,n)9, with affine map m,x0(P)=(EYP[ρ(x0,Y)])1/μ,r.m_{\ell,x_0}(P) = \left(\mathbb{E}_{Y\sim P}[\rho(x_0,Y)^\ell]\right)^{1/\ell} \le \mu, \qquad \ell \ge r.0. Two certification regimes are distinguished. The vanilla method certifies on an m,x0(P)=(EYP[ρ(x0,Y)])1/μ,r.m_{\ell,x_0}(P) = \left(\mathbb{E}_{Y\sim P}[\rho(x_0,Y)^\ell]\right)^{1/\ell} \le \mu, \qquad \ell \ge r.1-ball in flow space and is incomplete because infeasible flows may not correspond to valid images. The fine-tuned method restricts the flow to the feasible polytope

m,x0(P)=(EYP[ρ(x0,Y)])1/μ,r.m_{\ell,x_0}(P) = \left(\mathbb{E}_{Y\sim P}[\rho(x_0,Y)^\ell]\right)^{1/\ell} \le \mu, \qquad \ell \ge r.2

so that m,x0(P)=(EYP[ρ(x0,Y)])1/μ,r.m_{\ell,x_0}(P) = \left(\mathbb{E}_{Y\sim P}[\rho(x_0,Y)^\ell]\right)^{1/\ell} \le \mu, \qquad \ell \ge r.3 yields a complete certificate if the underlying certifier is complete. The same flow-domain viewpoint supports a projected gradient descent attack with substantially lower computational burden than image-space Wasserstein projection (Wegel et al., 2021).

Recent WDRO work tightens these certificates by matching upper and lower bounds more closely to network geometry. "Tight Robustness Certificates and Wasserstein Distributional Attacks for Deep Neural Networks" introduces a primal approach based on an exact Lipschitz certificate for ReLU networks, exploiting their piecewise-affine activation cells. For ReLU models, the certificate uses the maximum operator norm of cellwise Jacobians over reachable activation masks; a lower bound is constructed by moving mass along a worst-case margin direction inside the corresponding cell, and in many cases the lower and upper bounds coincide. The paper also introduces a Wasserstein distributional attack with flexible support and reports tighter certificates than existing methods (Le et al., 11 Oct 2025). This suggests a progression within the literature from transfer-based certificates to locally exact, activation-aware certificates.

7. Generalized and structural certificates beyond classical balanced transport

Wasserstein certificates also arise when the metric is generalized beyond balanced probability measures. "A Framework for Wasserstein-1-Type Metrics" introduces the discrepancy

m,x0(P)=(EYP[ρ(x0,Y)])1/μ,r.m_{\ell,x_0}(P) = \left(\mathbb{E}_{Y\sim P}[\rho(x_0,Y)^\ell]\right)^{1/\ell} \le \mu, \qquad \ell \ge r.4

for concave upper semicontinuous m,x0(P)=(EYP[ρ(x0,Y)])1/μ,r.m_{\ell,x_0}(P) = \left(\mathbb{E}_{Y\sim P}[\rho(x_0,Y)^\ell]\right)^{1/\ell} \le \mu, \qquad \ell \ge r.5 and convex m,x0(P)=(EYP[ρ(x0,Y)])1/μ,r.m_{\ell,x_0}(P) = \left(\mathbb{E}_{Y\sim P}[\rho(x_0,Y)^\ell]\right)^{1/\ell} \le \mu, \qquad \ell \ge r.6. The same family is shown to be equivalent to an infimal-convolution formulation m,x0(P)=(EYP[ρ(x0,Y)])1/μ,r.m_{\ell,x_0}(P) = \left(\mathbb{E}_{Y\sim P}[\rho(x_0,Y)^\ell]\right)^{1/\ell} \le \mu, \qquad \ell \ge r.7 combining m,x0(P)=(EYP[ρ(x0,Y)])1/μ,r.m_{\ell,x_0}(P) = \left(\mathbb{E}_{Y\sim P}[\rho(x_0,Y)^\ell]\right)^{1/\ell} \le \mu, \qquad \ell \ge r.8 transport with local discrepancies for mass creation and annihilation. The paper emphasizes convexity, locality, strong duality, and computational efficiency, so certificates derived in one formulation transfer to the other (Schmitzer et al., 2017).

For structured data, "Hausdorff and Wasserstein metrics on graphs and other structured data" lifts optimal transport to m,x0(P)=(EYP[ρ(x0,Y)])1/μ,r.m_{\ell,x_0}(P) = \left(\mathbb{E}_{Y\sim P}[\rho(x_0,Y)^\ell]\right)^{1/\ell} \le \mu, \qquad \ell \ge r.9-sets, including graphs, directed and undirected, labeled and unlabeled. The Wasserstein metric on PnP_n0-sets is a convex relaxation of the corresponding Hausdorff metric, and for finite PnP_n1-sets it is the value of a linear program. Infeasibility of the LP or an infinite optimal value serves as a certificate that no structure-preserving match exists, while deterministic optimal kernels can act as certificates of exact matching or homomorphism. The paper explicitly presents these as certificate-like guarantees for graph matching and related structured comparisons (Patterson, 2019).

At a more algebraic level, "Free complete Wasserstein algebras" axiomatizes Wasserstein barycentric structure on complete metric spaces. A Wasserstein algebra of order PnP_n2 is a barycentric algebra with metric satisfying

PnP_n3

The free complete algebra over a complete metric space PnP_n4 is PnP_n5, the Radon probability measures with finite PnP_n6-th moment, equipped with PnP_n7 and convex sum operations. The universal property of this construction provides an algebraic certificate that PnP_n8 is the canonical completion for probabilistic choice under Wasserstein geometry (Mardare et al., 2018).

Taken together, these extensions show that Wasserstein metric certificates are not confined to explicit transport costs between probability distributions. They also formalize guarantees for unbalanced transport, categorical structure, and algebraic probabilistic semantics, provided the certificate is stated in a geometry compatible with the relevant Wasserstein-type discrepancy (Schmitzer et al., 2017, Patterson, 2019, Mardare et al., 2018).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Wasserstein Metric Certificates.