Papers
Topics
Authors
Recent
Search
2000 character limit reached

Excess Wasserstein Gap in Optimal Transport

Updated 14 July 2026
  • Excess Wasserstein gap is a residual measure that quantifies the discrepancy when replacing exact Wasserstein metrics with tractable approximations.
  • It arises in optimal transport when discretization, projection, or surrogate methods reduce the precision of continuous transport objects across various applications.
  • Understanding this gap provides actionable insights to improve variational inference, transport cost surrogation, and adversarial training in high-dimensional settings.

“Excess Wasserstein gap” is best understood as an Editor’s term for a family of residual quantities that measure how far a tractable object remains from an exact Wasserstein quantity. In the literature considered here, the expression is not introduced uniformly, but closely related discrepancies appear in several precise forms: the residual between a continuous-time Wasserstein gradient flow and its variational discretization, the difference between a surrogate Gaussian transport cost and the exact Gaussian Wasserstein cost, the excess of a distributional distance over a mean-gap summary, the deficit between full and sliced Wasserstein costs, and the mismatch between a conditional Wasserstein objective and its implementable adversarial formulation (Yi et al., 2023, Rostami et al., 30 Mar 2026, Sauerberg, 24 Aug 2025, Han, 25 May 2026, Martin, 2021).

1. Scope of the notion

Across these works, the relevant quantity is always a discrepancy between a more exact transport object and a reduced representation. The sign convention varies by context, but the structural role is stable: the gap isolates what is lost under discretization, projection, surrogation, or aggregation.

Context Precise quantity Interpretation
VI and gradient flows (Yi et al., 2023) continuous-time Wasserstein gradient flow vs. discrete-time BBVI update discretization and parameterization residual
Gaussian mixture flow matching (Rostami et al., 30 Mar 2026) CijW2,ij2C_{ij}-W_{2,ij}^2 excess of surrogate transport cost over exact Gaussian Wasserstein cost
Life tables (Sauerberg, 24 Aug 2025) W1(PA,PB)e0,Ae0,BW_1(P_A,P_B)-|e_{0,A}-e_{0,B}| distributional excess beyond life expectancy gap
Sliced Wasserstein (Han, 25 May 2026) D(μ,ν)=1dW22(μ,ν)SW22(μ,ν)\mathrm D(\mu,\nu)=\frac1dW_2^2(\mu,\nu)-SW_2^2(\mu,\nu) multidimensional transport cost not captured by slicing
Conditional WGANs (Martin, 2021) averaged conditional Wasserstein objective vs. practical discriminator objective objective mismatch if E\mathbb E and sup\sup are not exchangeable

This suggests that the phrase is not a single invariant of optimal transport, but rather a context-dependent descriptor for the residual left after replacing a full Wasserstein object by something computationally or analytically simpler.

2. Variational inference and Wasserstein gradient-flow residuals

In “Bridging the Gap Between Variational Inference and Wasserstein Gradient Flows” (Yi et al., 2023), the relevant gap is the mismatch between the continuous-time Wasserstein gradient flow and the discrete-time black-box variational inference update. The paper studies a Bures-Wasserstein gradient flow on a Gaussian variational family,

qθ(z)=N(z;μ,Σ),q_\theta(z)=\mathcal N(z;\mu,\Sigma),

and shows that, under suitable conditions, the Bures-Wasserstein gradient flow can be recast as a Euclidean gradient flow whose forward Euler discretization is exactly the standard black-box variational inference update.

The forward Euler step is written as

θk+1=θkηθL(θk),\theta_{k+1}=\theta_k-\eta\,\nabla_\theta \mathcal L(\theta_k),

or, equivalently for ELBO maximization,

θk+1=θk+ηθELBO(θk)^.\theta_{k+1}=\theta_k+\eta\,\widehat{\nabla_\theta \mathrm{ELBO}(\theta_k)}.

The paper’s key claim is that this discrete-time VI update can be interpreted as a forward Euler step for the Wasserstein/Bures-Wasserstein flow. In that sense, the excess gap is not a new divergence but the residual between ideal continuous transport dynamics and the algorithmic update used in practice.

A second key object is the path-derivative gradient estimator. In reparameterized form,

z=gθ(ε),εp(ε),z=g_\theta(\varepsilon), \qquad \varepsilon\sim p(\varepsilon),

and

θEε[f(gθ(ε))]=Eε[zf(gθ(ε))θgθ(ε)].\nabla_\theta \mathbb E_{\varepsilon}\big[f(g_\theta(\varepsilon))\big] = \mathbb E_{\varepsilon}\big[\nabla_z f(g_\theta(\varepsilon))\,\nabla_\theta g_\theta(\varepsilon)\big].

The paper identifies the vector field of the gradient flow with this path-derivative gradient estimator. It also frames the path-derivative gradient as a distillation procedure: the Wasserstein gradient flow is the “teacher” dynamics in probability space, while the variational family acts as a “student” that mimics that flow.

Within this framework, a natural excess Wasserstein gap is the deviation between the algorithmic update and the exact flow vector field, written in the synthesis as

W1(PA,PB)e0,Ae0,BW_1(P_A,P_B)-|e_{0,A}-e_{0,B}|0

The paper’s equivalence is strongest in the Gaussian case and relies on a Gaussian variational family, Bures-Wasserstein geometry, sufficient regularity of the target density and variational objective, a reparameterizable family, and the Euclidean recasting that is special to the Gaussian/Bures setting. The same distillation viewpoint is then extended to W1(PA,PB)e0,Ae0,BW_1(P_A,P_B)-|e_{0,A}-e_{0,B}|1-divergences and non-Gaussian variational families, yielding a new gradient estimator for W1(PA,PB)e0,Ae0,BW_1(P_A,P_B)-|e_{0,A}-e_{0,B}|2-divergences that is readily implementable in PyTorch or TensorFlow (Yi et al., 2023).

3. Surrogate-versus-exact Gaussian transport costs

In “An Explicit Surrogate for Gaussian Mixture Flow Matching with Wasserstein Gap Bounds” (Rostami et al., 30 Mar 2026), the excess Wasserstein gap is introduced explicitly as the difference between a surrogate kinetic transport cost and the exact Gaussian Wasserstein cost. The source and target are Gaussian mixture models,

W1(PA,PB)e0,Ae0,BW_1(P_A,P_B)-|e_{0,A}-e_{0,B}|3

and transport is built component-wise.

For each pair W1(PA,PB)e0,Ae0,BW_1(P_A,P_B)-|e_{0,A}-e_{0,B}|4, the paper uses the linear Gaussian path

W1(PA,PB)e0,Ae0,BW_1(P_A,P_B)-|e_{0,A}-e_{0,B}|5

A lemma states that if W1(PA,PB)e0,Ae0,BW_1(P_A,P_B)-|e_{0,A}-e_{0,B}|6 with differentiable W1(PA,PB)e0,Ae0,BW_1(P_A,P_B)-|e_{0,A}-e_{0,B}|7 and W1(PA,PB)e0,Ae0,BW_1(P_A,P_B)-|e_{0,A}-e_{0,B}|8, then the continuity equation admits the affine solution

W1(PA,PB)e0,Ae0,BW_1(P_A,P_B)-|e_{0,A}-e_{0,B}|9

Hence the component-wise surrogate field is

D(μ,ν)=1dW22(μ,ν)SW22(μ,ν)\mathrm D(\mu,\nu)=\frac1dW_2^2(\mu,\nu)-SW_2^2(\mu,\nu)0

The surrogate pairwise cost is the kinetic action

D(μ,ν)=1dW22(μ,ν)SW22(μ,ν)\mathrm D(\mu,\nu)=\frac1dW_2^2(\mu,\nu)-SW_2^2(\mu,\nu)1

and the exact comparison target is the Gaussian squared D(μ,ν)=1dW22(μ,ν)SW22(μ,ν)\mathrm D(\mu,\nu)=\frac1dW_2^2(\mu,\nu)-SW_2^2(\mu,\nu)2-Wasserstein distance

D(μ,ν)=1dW22(μ,ν)SW22(μ,ν)\mathrm D(\mu,\nu)=\frac1dW_2^2(\mu,\nu)-SW_2^2(\mu,\nu)3

The excess Wasserstein gap is then

D(μ,ν)=1dW22(μ,ν)SW22(μ,ν)\mathrm D(\mu,\nu)=\frac1dW_2^2(\mu,\nu)-SW_2^2(\mu,\nu)4

or its absolute value.

The analysis is organized by the whitened covariance perturbation size

D(μ,ν)=1dW22(μ,ν)SW22(μ,ν)\mathrm D(\mu,\nu)=\frac1dW_2^2(\mu,\nu)-SW_2^2(\mu,\nu)5

Under the local commuting regime D(μ,ν)=1dW22(μ,ν)SW22(μ,ν)\mathrm D(\mu,\nu)=\frac1dW_2^2(\mu,\nu)-SW_2^2(\mu,\nu)6, the surrogate and exact costs share the same mean term and the same second-order covariance term,

D(μ,ν)=1dW22(μ,ν)SW22(μ,ν)\mathrm D(\mu,\nu)=\frac1dW_2^2(\mu,\nu)-SW_2^2(\mu,\nu)7

so that

D(μ,ν)=1dW22(μ,ν)SW22(μ,ν)\mathrm D(\mu,\nu)=\frac1dW_2^2(\mu,\nu)-SW_2^2(\mu,\nu)8

The paper further gives an explicit cubic error bound

D(μ,ν)=1dW22(μ,ν)SW22(μ,ν)\mathrm D(\mu,\nu)=\frac1dW_2^2(\mu,\nu)-SW_2^2(\mu,\nu)9

and emphasizes that the constants deteriorate as E\mathbb E0 becomes ill-conditioned.

For nonlocal regimes, path splitting subdivides the covariance path into local segments. This restores locality of the comparison, but the paper makes an important qualification: path splitting fixes the locality requirement and does not remove the commuting assumption used in the cubic-gap theorem. The exact reference construction, called Method B, uses the Gaussian Wasserstein geodesic and has kinetic action exactly equal to the Gaussian Wasserstein cost. The practical regime map summarized in the paper is that the surrogate is best in local, well-conditioned, and approximately commuting regimes, whereas the exact method is preferable in highly nonlocal, strongly noncommuting, ill-conditioned, or near-boundary regimes (Rostami et al., 30 Mar 2026).

4. Distributional excess beyond mean-gap summaries

In “On the relationship between the Wasserstein distance and differences in life expectancy at birth” (Sauerberg, 24 Aug 2025), the excess Wasserstein gap is the amount by which a full distributional distance exceeds the absolute mean-gap summary. The paper treats a life table as an age-at-death distribution, with survivorship function E\mathbb E1, density E\mathbb E2, and distribution function E\mathbb E3. Life expectancy at birth satisfies

E\mathbb E4

and for two populations E\mathbb E5 and E\mathbb E6,

E\mathbb E7

For one-dimensional age-at-death distributions, the E\mathbb E8-Wasserstein distance simplifies to

E\mathbb E9

Thus sup\sup0 is the area between the survivorship curves, while the life-expectancy difference is the signed area between them.

The main theorem is that if the survivorship functions do not cross, then the sup\sup1-Wasserstein distance equals the difference in life expectancy at birth. If, for example,

sup\sup2

then

sup\sup3

and with reversed ordering sup\sup4.

The excess gap arises when curves cross. The synthesis defines it as

sup\sup5

Because sup\sup6 uses absolute area while sup\sup7 uses signed area, crossing generates cancellation in the mean but not in the transport distance. Therefore,

sup\sup8

with strict inequality whenever survivorship curves cross and cancellation occurs. Demographically, sup\sup9 captures the net mean lifespan difference, whereas qθ(z)=N(z;μ,Σ),q_\theta(z)=\mathcal N(z;\mu,\Sigma),0 captures the full distributional dissimilarity in ages at death.

The empirical sections reinforce this interpretation. For 5,000 sampled country pairs from the Human Mortality Database, the distributions of qθ(z)=N(z;μ,Σ),q_\theta(z)=\mathcal N(z;\mu,\Sigma),1 and qθ(z)=N(z;μ,Σ),q_\theta(z)=\mathcal N(z;\mu,\Sigma),2 differences overlap strongly, both range from qθ(z)=N(z;μ,Σ),q_\theta(z)=\mathcal N(z;\mu,\Sigma),3 to about qθ(z)=N(z;μ,Σ),q_\theta(z)=\mathcal N(z;\mu,\Sigma),4 years, the mean qθ(z)=N(z;μ,Σ),q_\theta(z)=\mathcal N(z;\mu,\Sigma),5 is qθ(z)=N(z;μ,Σ),q_\theta(z)=\mathcal N(z;\mu,\Sigma),6, the mean qθ(z)=N(z;μ,Σ),q_\theta(z)=\mathcal N(z;\mu,\Sigma),7 difference is qθ(z)=N(z;μ,Σ),q_\theta(z)=\mathcal N(z;\mu,\Sigma),8, and Pearson’s qθ(z)=N(z;μ,Σ),q_\theta(z)=\mathcal N(z;\mu,\Sigma),9. A crossing example is England & Wales versus Iceland in 1849: the θk+1=θkηθL(θk),\theta_{k+1}=\theta_k-\eta\,\nabla_\theta \mathcal L(\theta_k),0 values are both about θk+1=θkηθL(θk),\theta_{k+1}=\theta_k-\eta\,\nabla_\theta \mathcal L(\theta_k),1 years, yet θk+1=θkηθL(θk),\theta_{k+1}=\theta_k-\eta\,\nabla_\theta \mathcal L(\theta_k),2 is large because Iceland has higher infant mortality and lower mortality at older ages. For women and men in period life tables from 1990 to 2020, the relationship is even tighter: the average absolute difference between θk+1=θkηθL(θk),\theta_{k+1}=\theta_k-\eta\,\nabla_\theta \mathcal L(\theta_k),3 and θk+1=θkηθL(θk),\theta_{k+1}=\theta_k-\eta\,\nabla_\theta \mathcal L(\theta_k),4 differences is below θk+1=θkηθL(θk),\theta_{k+1}=\theta_k-\eta\,\nabla_\theta \mathcal L(\theta_k),5 in every year, and for women versus men specifically it is below θk+1=θkηθL(θk),\theta_{k+1}=\theta_k-\eta\,\nabla_\theta \mathcal L(\theta_k),6 each year, with maximum difference about θk+1=θkηθL(θk),\theta_{k+1}=\theta_k-\eta\,\nabla_\theta \mathcal L(\theta_k),7 years in 2008 (Sauerberg, 24 Aug 2025).

5. Sliced Wasserstein deficits as multidimensional excess

In “Rigidity and Quantitative Stability of the Sliced Wasserstein Deficit” (Han, 25 May 2026), the excess Wasserstein gap is formalized as a deficit: θk+1=θkηθL(θk),\theta_{k+1}=\theta_k-\eta\,\nabla_\theta \mathcal L(\theta_k),8 Here θk+1=θkηθL(θk),\theta_{k+1}=\theta_k-\eta\,\nabla_\theta \mathcal L(\theta_k),9 averages one-dimensional quadratic Wasserstein distances over linear projections. The elementary comparison

θk+1=θk+ηθELBO(θk)^.\theta_{k+1}=\theta_k+\eta\,\widehat{\nabla_\theta \mathrm{ELBO}(\theta_k)}.0

follows by projecting couplings and averaging

θk+1=θk+ηθELBO(θk)^.\theta_{k+1}=\theta_k+\eta\,\widehat{\nabla_\theta \mathrm{ELBO}(\theta_k)}.1

The deficit is therefore the excess left after accounting for all one-dimensional projections.

A central identity shows that, when θk+1=θk+ηθELBO(θk)^.\theta_{k+1}=\theta_k+\eta\,\widehat{\nabla_\theta \mathrm{ELBO}(\theta_k)}.2 and θk+1=θk+ηθELBO(θk)^.\theta_{k+1}=\theta_k+\eta\,\widehat{\nabla_\theta \mathrm{ELBO}(\theta_k)}.3 is the Brenier map from θk+1=θk+ηθELBO(θk)^.\theta_{k+1}=\theta_k+\eta\,\widehat{\nabla_\theta \mathrm{ELBO}(\theta_k)}.4 to θk+1=θk+ηθELBO(θk)^.\theta_{k+1}=\theta_k+\eta\,\widehat{\nabla_\theta \mathrm{ELBO}(\theta_k)}.5,

θk+1=θk+ηθELBO(θk)^.\theta_{k+1}=\theta_k+\eta\,\widehat{\nabla_\theta \mathrm{ELBO}(\theta_k)}.6

This expresses the deficit as an average of nonnegative one-dimensional optimality gaps. In that form, the deficit becomes a rigidity functional rather than merely a difference of distances.

The rigidity theorem states that

θk+1=θk+ηθELBO(θk)^.\theta_{k+1}=\theta_k+\eta\,\widehat{\nabla_\theta \mathrm{ELBO}(\theta_k)}.7

for some θk+1=θk+ηθELBO(θk)^.\theta_{k+1}=\theta_k+\eta\,\widehat{\nabla_\theta \mathrm{ELBO}(\theta_k)}.8 and θk+1=θk+ηθELBO(θk)^.\theta_{k+1}=\theta_k+\eta\,\widehat{\nabla_\theta \mathrm{ELBO}(\theta_k)}.9. Equality in the sliced inequality thus forces the Brenier map to be homothetic affine, not merely affine. The paper also proves a higher-dimensional Grassmannian analogue for z=gθ(ε),εp(ε),z=g_\theta(\varepsilon), \qquad \varepsilon\sim p(\varepsilon),0-dimensional projections.

For quantitative stability, the paper introduces the ridge defect

z=gθ(ε),εp(ε),z=g_\theta(\varepsilon), \qquad \varepsilon\sim p(\varepsilon),1

and the sliced Poincaré–Korn constant z=gθ(ε),εp(ε),z=g_\theta(\varepsilon), \qquad \varepsilon\sim p(\varepsilon),2, whose null space is

z=gθ(ε),εp(ε),z=g_\theta(\varepsilon), \qquad \varepsilon\sim p(\varepsilon),3

If z=gθ(ε),εp(ε),z=g_\theta(\varepsilon), \qquad \varepsilon\sim p(\varepsilon),4 and the projected monotone transports have uniformly bounded one-dimensional Lipschitz scale

z=gθ(ε),εp(ε),z=g_\theta(\varepsilon), \qquad \varepsilon\sim p(\varepsilon),5

then

z=gθ(ε),εp(ε),z=g_\theta(\varepsilon), \qquad \varepsilon\sim p(\varepsilon),6

The paper emphasizes that z=gθ(ε),εp(ε),z=g_\theta(\varepsilon), \qquad \varepsilon\sim p(\varepsilon),7 is a one-dimensional regularity scale; for Gaussian source measures and strongly log-concave targets, Caffarelli’s contraction theorem yields the uniform bound z=gθ(ε),εp(ε),z=g_\theta(\varepsilon), \qquad \varepsilon\sim p(\varepsilon),8.

The Gaussian model gives the sharp benchmark: z=gθ(ε),εp(ε),z=g_\theta(\varepsilon), \qquad \varepsilon\sim p(\varepsilon),9 for the standard Gaussian θEε[f(gθ(ε))]=Eε[zf(gθ(ε))θgθ(ε)].\nabla_\theta \mathbb E_{\varepsilon}\big[f(g_\theta(\varepsilon))\big] = \mathbb E_{\varepsilon}\big[\nabla_z f(g_\theta(\varepsilon))\,\nabla_\theta g_\theta(\varepsilon)\big].0, and the same sharp constant holds for every isotropic Gaussian θEε[f(gθ(ε))]=Eε[zf(gθ(ε))θgθ(ε)].\nabla_\theta \mathbb E_{\varepsilon}\big[f(g_\theta(\varepsilon))\big] = \mathbb E_{\varepsilon}\big[\nabla_z f(g_\theta(\varepsilon))\,\nabla_\theta g_\theta(\varepsilon)\big].1. Positive SPK bounds also persist under bounded θEε[f(gθ(ε))]=Eε[zf(gθ(ε))θgθ(ε)].\nabla_\theta \mathbb E_{\varepsilon}\big[f(g_\theta(\varepsilon))\big] = \mathbb E_{\varepsilon}\big[\nabla_z f(g_\theta(\varepsilon))\,\nabla_\theta g_\theta(\varepsilon)\big].2 perturbations of the Gaussian: θEε[f(gθ(ε))]=Eε[zf(gθ(ε))θgθ(ε)].\nabla_\theta \mathbb E_{\varepsilon}\big[f(g_\theta(\varepsilon))\big] = \mathbb E_{\varepsilon}\big[\nabla_z f(g_\theta(\varepsilon))\,\nabla_\theta g_\theta(\varepsilon)\big].3

The main obstruction is anisotropy. The anisotropic Gaussians

θEε[f(gθ(ε))]=Eε[zf(gθ(ε))θgθ(ε)].\nabla_\theta \mathbb E_{\varepsilon}\big[f(g_\theta(\varepsilon))\big] = \mathbb E_{\varepsilon}\big[\nabla_z f(g_\theta(\varepsilon))\,\nabla_\theta g_\theta(\varepsilon)\big].4

satisfy a uniform Bakry–Émery lower curvature bound and a uniform Poincaré constant, yet θEε[f(gθ(ε))]=Eε[zf(gθ(ε))θgθ(ε)].\nabla_\theta \mathbb E_{\varepsilon}\big[f(g_\theta(\varepsilon))\big] = \mathbb E_{\varepsilon}\big[\nabla_z f(g_\theta(\varepsilon))\,\nabla_\theta g_\theta(\varepsilon)\big].5 as θEε[f(gθ(ε))]=Eε[zf(gθ(ε))θgθ(ε)].\nabla_\theta \mathbb E_{\varepsilon}\big[f(g_\theta(\varepsilon))\big] = \mathbb E_{\varepsilon}\big[\nabla_z f(g_\theta(\varepsilon))\,\nabla_\theta g_\theta(\varepsilon)\big].6. This shows that neither a Bakry–Émery lower curvature bound nor a usual Poincaré inequality alone can imply a global sliced Poincaré–Korn inequality (Han, 25 May 2026).

6. Objective mismatches in conditional Wasserstein GANs

In “About exchanging expectation and supremum for conditional Wasserstein GANs” (Martin, 2021), the excess Wasserstein gap is the potential discrepancy between the theoretically correct conditional Wasserstein objective and the implementable discriminator objective. The paper distinguishes two formulations.

The joint Wasserstein objective on θEε[f(gθ(ε))]=Eε[zf(gθ(ε))θgθ(ε)].\nabla_\theta \mathbb E_{\varepsilon}\big[f(g_\theta(\varepsilon))\big] = \mathbb E_{\varepsilon}\big[\nabla_z f(g_\theta(\varepsilon))\,\nabla_\theta g_\theta(\varepsilon)\big].7 is

θEε[f(gθ(ε))]=Eε[zf(gθ(ε))θgθ(ε)].\nabla_\theta \mathbb E_{\varepsilon}\big[f(g_\theta(\varepsilon))\big] = \mathbb E_{\varepsilon}\big[\nabla_z f(g_\theta(\varepsilon))\,\nabla_\theta g_\theta(\varepsilon)\big].8

and requires discriminators that are θEε[f(gθ(ε))]=Eε[zf(gθ(ε))θgθ(ε)].\nabla_\theta \mathbb E_{\varepsilon}\big[f(g_\theta(\varepsilon))\big] = \mathbb E_{\varepsilon}\big[\nabla_z f(g_\theta(\varepsilon))\,\nabla_\theta g_\theta(\varepsilon)\big].9-Lipschitz on W1(PA,PB)e0,Ae0,BW_1(P_A,P_B)-|e_{0,A}-e_{0,B}|00, hence in both arguments.

The averaged conditional objective is

W1(PA,PB)e0,Ae0,BW_1(P_A,P_B)-|e_{0,A}-e_{0,B}|01

Formally applying Kantorovich–Rubinstein duality pointwise suggests the practical objective

W1(PA,PB)e0,Ae0,BW_1(P_A,P_B)-|e_{0,A}-e_{0,B}|02

where W1(PA,PB)e0,Ae0,BW_1(P_A,P_B)-|e_{0,A}-e_{0,B}|03 consists of measurable discriminators that are W1(PA,PB)e0,Ae0,BW_1(P_A,P_B)-|e_{0,A}-e_{0,B}|04-Lipschitz in W1(PA,PB)e0,Ae0,BW_1(P_A,P_B)-|e_{0,A}-e_{0,B}|05 only. The delicate step is the exchange of W1(PA,PB)e0,Ae0,BW_1(P_A,P_B)-|e_{0,A}-e_{0,B}|06 and W1(PA,PB)e0,Ae0,BW_1(P_A,P_B)-|e_{0,A}-e_{0,B}|07. The obstacle is measurability: the pointwise optimizer W1(PA,PB)e0,Ae0,BW_1(P_A,P_B)-|e_{0,A}-e_{0,B}|08 from one-dimensional Kantorovich–Rubinstein duality need not depend measurably on W1(PA,PB)e0,Ae0,BW_1(P_A,P_B)-|e_{0,A}-e_{0,B}|09. Earlier work such as Adler et al. used the objective, but the exchange was not fully justified.

The paper proves the equality under three assumptions: compact support of the condition W1(PA,PB)e0,Ae0,BW_1(P_A,P_B)-|e_{0,A}-e_{0,B}|10, Wasserstein continuity of the conditional law W1(PA,PB)e0,Ae0,BW_1(P_A,P_B)-|e_{0,A}-e_{0,B}|11, and continuity of the generator in expectation,

W1(PA,PB)e0,Ae0,BW_1(P_A,P_B)-|e_{0,A}-e_{0,B}|12

uniformly on the compact support. A lemma first shows that

W1(PA,PB)e0,Ae0,BW_1(P_A,P_B)-|e_{0,A}-e_{0,B}|13

is uniformly continuous. The proof of the main theorem then partitions the compact W1(PA,PB)e0,Ae0,BW_1(P_A,P_B)-|e_{0,A}-e_{0,B}|14-space into finitely many boxes, selects representative points, chooses approximately optimal Kantorovich potentials on those representatives, and assembles a measurable piecewise discriminator

W1(PA,PB)e0,Ae0,BW_1(P_A,P_B)-|e_{0,A}-e_{0,B}|15

Under these assumptions,

W1(PA,PB)e0,Ae0,BW_1(P_A,P_B)-|e_{0,A}-e_{0,B}|16

Hence the excess gap disappears: the practical adversarial objective is exactly the intended averaged conditional Wasserstein distance. The associated misconception is precise. If one optimizes the joint Wasserstein distance on W1(PA,PB)e0,Ae0,BW_1(P_A,P_B)-|e_{0,A}-e_{0,B}|17, the discriminator must be Lipschitz in both W1(PA,PB)e0,Ae0,BW_1(P_A,P_B)-|e_{0,A}-e_{0,B}|18 and W1(PA,PB)e0,Ae0,BW_1(P_A,P_B)-|e_{0,A}-e_{0,B}|19. If one optimizes the averaged conditional Wasserstein objective, then under the theorem’s assumptions it is enough to enforce Lipschitzness only in W1(PA,PB)e0,Ae0,BW_1(P_A,P_B)-|e_{0,A}-e_{0,B}|20 (Martin, 2021).

Taken together, these works show that an excess Wasserstein gap can arise from very different mechanisms—time discretization, surrogate transport costs, cancellation in signed summaries, projection onto lower-dimensional marginals, or dual-formulation subtleties. The unifying theme is that Wasserstein geometry supplies an exact reference object, while the gap records what remains after replacing that object by a tractable approximation, projection, or reduced statistic.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Excess Wasserstein Gap.