---
title: Generalized Local Squared Wasserstein-2 Loss
url: https://www.emergentmind.com/topics/generalized-local-squared-wasserstein-2-loss
type: topic
---

# Generalized Local Squared Wasserstein-2 Loss

Generalized Local Squared Wasserstein-2 Loss denotes a family of objectives that replace a single global transport discrepancy by a sum, average, or asymptotic expansion of squared \(W_2\) terms computed in localized contexts. In the literature, the localization mechanism is not uniform: it may be temporal and barycentric along a discrete curve of measures, infinitesimal in the Riemannian geometry induced by \(W_2\), conditional on neighborhoods in input or initial-condition space, or local in parameter space through second-order expansions of \(W_2^2\)-induced risks. The shared principle is the same: penalize deviations from a local Wasserstein reference rather than only a global mismatch [2302.10682] [1909.06860] [2406.06825] [2510.05645].

## 1. Multiple meanings of locality

The expression is used for several related constructions rather than a single standardized formula. This suggests that it is best understood as a family name for squared \(W_2\)-based local objectives whose precise form depends on the ambient problem: spline interpolation in Wasserstein space, smoothness regularization on image manifolds, conditional distribution matching in uncertainty quantification, or local quadratic approximation in Bayesian asymptotics [2302.10682] [1909.06860] [2503.05068] [2507.05143].

| Setting | Locality mechanism | Representative form |
|---|---|---|
| Wasserstein splines | interior time index \(k\), anchored barycentric average of \(\mu_{k-1},\mu_{k+1}\) around \(\mu_k\) | \(4K^3\sum_{k=1}^{K-1}W_2^2\!\big(\mu_k,\operatorname{Bar}_{\mu_k}(\mu_{k-1},\mu_{k+1})\big)+\delta K\sum_{k=0}^{K-1}W_2^2(\mu_k,\mu_{k+1})\) |
| W2 diffusion regularization | small \(W_2\)-geodesic displacements in input space | \(\mathcal L_{\mathrm{GLSW2}}(f)\approx \frac{\eta^2}{2}\,\mathbb E_x\,\operatorname{Tr}\!\big(J_f(x)^\top L(x)J_f(x)\big)\) |
| Conditional UQ / inverse problems | neighborhoods \(B(x,\delta)\) in input space | \(\int_D W_2^2(\mu_{x,\delta}^e,\hat\mu_{x,\delta}^e)\,\nu^e(dx)\) |
| Time-decoupled dynamics | neighborhoods in initial conditions and separate time slices | \(\int_0^T W_{2,\delta}^{2,e}(X(t),\hat X(t))\,dt\) |
| Parametric Bayes asymptotics | local expansion around \((\vartheta_0,\vartheta_0)\) | \(\ell_{W_2}(t,\vartheta)=\langle t-\vartheta,A(\vartheta)(t-\vartheta)\rangle+\xi(t,\vartheta)\) |

A recurrent misconception is that “local” must mean spatial locality in the underlying state space. In these works it can instead mean temporal locality, tangent-space locality, conditional locality with respect to covariates, or local asymptotic structure in parameter space. Likewise, “generalized” may refer either to anchored barycenters, generalized costs for mixed continuous–categorical outputs, time-decoupled formulations, or unbalanced transport with source terms.

## 2. Optimal-transport foundations and generalized variants

For probability measures \(\mu,\nu\in P(\Omega)\) with finite second moments, the squared Wasserstein-2 distance is defined by
\[
W_2^2(\mu,\nu)=\inf_{\Pi\in U(\mu,\nu)}\int_{\Omega\times\Omega} d^2(x,y)\,d\Pi(x,y),
\]
and in the Benamou–Brenier dynamic formulation,
\[
W_2^2(\mu_0,\mu_1)=\inf_{(\mu,v)\in CE(\mu_0,\mu_1)}\int_0^1\int_{\mathbb R^d}\|v_t(x)\|^2\,d\mu_t(x)\,dt.
\]
In the spline framework, this Riemannian viewpoint is combined with local barycentric constructions. For \(\mu,\nu\in P(\Omega)\) and \(t\in[0,1]\), the \(t\)-barycenter is
\[
\operatorname{Bar}^t(\mu,\nu)\in\arg\min_{\rho\in P(\Omega)}(1-t)W_2^2(\rho,\mu)+tW_2^2(\rho,\nu),
\]
while the anchored generalized barycenter with base \(\mu_2\) is
\[
\operatorname{Bar}_{\mu_2}^t(\mu_1,\mu_3):=((1-t)\pi^1+t\pi^3)_\#\Pi
\]
for a three-marginal optimal plan \(\Pi\) whose \((\pi^1,\pi^2)\)- and \((\pi^2,\pi^3)\)-marginals are optimal couplings [2302.10682].

A distinct generalization allows mass removal and creation. For \(a,b>0\) and \(p\ge 1\),
\[
T_{a,b}(\mu,\nu):=\inf_{\tilde\mu,\tilde\nu\in M,\ |\tilde\mu|=|\tilde\nu|} a^p\big(|\mu-\tilde\mu|+|\nu-\tilde\nu|\big)^p+b^pW_p(\tilde\mu,\tilde\nu)^p,
\]
and
\[
W_p^{a,b}(\mu,\nu):=[T_{a,b}(\mu,\nu)]^{1/p}.
\]
For \(p=2\), the generalized Benamou–Brenier formula states
\[
W_2^{a,b}(\mu_0,\mu_1)^2=\inf\{B_{a,b}[\mu,v,s]:(\mu,v,s)\in V(\mu_0,\mu_1)\},
\]
with
\[
B_{a,b}[\mu,v,s]=a^2\int_0^1\int_{\mathbb R^d}d|s_t|(x)\,dt+b^2\int_0^1\int_{\mathbb R^d}|v_t(x)|^2\,d\mu_t(x)\,dt.
\]
On this basis, a “Generalized Local Squared Wasserstein-2 Loss” can be formed patchwise,
\[
L_{\mathrm{local}}(\mu,\nu)=\sum_{k=1}^K\kappa_k\,[W_2^{a,b}(\mu|_{U_k},\nu|_{U_k})]^2,
\]
thereby localizing both transport and source penalties [1304.7014].

## 3. Anchored barycentric spline loss in Wasserstein space

The most explicit practical identification of the generalized local squared \(W_2\) loss appears in the discrete spline model for measure-valued interpolation. The continuous spline energy is defined as
\[
\mathfrak F((\mu_t)_t)=\inf_{v:(\mu,v)\in CE(\mu_0,\mu_1),\ v_t\in T_{\mu_t}}
\int_0^1\int_{\mathbb R^d}\left\|\partial_t v_t+\frac12\nabla\|v_t\|^2\right\|^2\,d\mu_t\,dt,
\]
with action regularization
\[
\mathfrak E((\mu_t)_t)=\inf_{v:(\mu,v)\in CE(\mu_0,\mu_1)}\int_0^1\int_{\mathbb R^d}\|v_t\|^2\,d\mu_t\,dt,
\]
and regularized energy \(\mathfrak F^\delta:=\mathfrak F+\delta\mathfrak E\). After time discretization with \(t_k=k/K\), the discrete action is
\[
\mathfrak E^K(\mu^K):=K\sum_{k=0}^{K-1}W_2^2(\mu_k^K,\mu_{k+1}^K).
\]

The non-anchored squared-acceleration term is
\[
\mathfrak F^K(\mu^K):=\inf_{\{\tilde\mu_k^K\}}4K^3\sum_{k=1}^{K-1}W_2^2(\mu_k^K,\tilde\mu_k^K),
\quad \tilde\mu_k^K\in\operatorname{Bar}(\mu_{k-1}^K,\mu_{k+1}^K),
\]
whereas the anchored or generalized local version is
\[
\mathfrak F_G^K(\mu^K):=\inf_{\{\tilde\mu_k^K\}}4K^3\sum_{k=1}^{K-1}W_2^2(\mu_k^K,\tilde\mu_k^K),
\quad \tilde\mu_k^K\in\operatorname{Bar}_{\mu_k^K}(\mu_{k-1}^K,\mu_{k+1}^K).
\]
The corresponding regularized energies are \(\mathfrak F^{\delta,K}:=\mathfrak F^K+\delta\mathfrak E^K\) and \(\mathfrak F_G^{\delta,K}:=\mathfrak F_G^K+\delta\mathfrak E^K\). The practical loss is then
\[
L_{\mathrm{GLSW2}}(\mu^K)=4K^3\sum_{k=1}^{K-1}W_2^2\!\big(\mu_k^K,\operatorname{Bar}_{\mu_k^K}(\mu_{k-1}^K,\mu_{k+1}^K)\big)
+\delta K\sum_{k=0}^{K-1}W_2^2(\mu_k^K,\mu_{k+1}^K),
\]
subject to interpolation constraints \(\mu_{K\bar t_i}^K=\bar\mu_i\) and one discrete boundary condition among natural, Hermite, or periodic [2302.10682].

For \(\delta>0\), existence of discrete regularized spline interpolations follows from tightness and lower semicontinuity. On the Gaussian subspace \(\mu_i=\mathcal N(m_i,\Sigma_i)\), \(W_2^2\) is explicit,
\[
W_2^2(\mathcal N(m_1,\Sigma_1),\mathcal N(m_2,\Sigma_2))
=\|m_1-m_2\|^2+\operatorname{Tr}\!\Big(\Sigma_1+\Sigma_2-2(\Sigma_2^{1/2}\Sigma_1\Sigma_2^{1/2})^{1/2}\Big),
\]
the anchored barycenter covariance is explicit, and the discrete energies satisfy the consistency estimates
\[
\mathfrak E((\mu_t)_t)=\mathfrak E^K((\mu_0^K,\dots,\mu_K^K))+O(K^{-1}),\qquad
\mathfrak F((\mu_t)_t)=\mathfrak F_G^K((\mu_0^K,\dots,\mu_K^K))+O(K^{-1}).
\]

Numerically, the fully discrete functional is optimized with entropy-regularized OT via Sinkhorn and a variant of Nesterov’s accelerated gradient descent on simplex-constrained weights. The update uses extrapolation, automatic differentiation through Sinkhorn, multiplicative updates
\[
\tilde\omega^k\leftarrow \tilde\omega^k\odot \exp(-t\beta\,\mathrm{grad}_k),
\]
followed by renormalization. The paper explicitly states that it does not use iPALM; rather, it uses an accelerated first-order scheme tailored to the simplex constraints. The method is applied to numerical spline interpolation and to synthesized textures, where local feature distributions are spline-interpolated and then matched by a generator in feature space.

## 4. Infinitesimal geometry, diffusion regularization, and local quadratic asymptotics

In discriminative learning, a generalized local squared \(W_2\) loss arises from the Riemannian geometry induced by Wasserstein-2 on images or histograms. If \(x\in\mathbb R^n\) denotes an image and \(G_W(x)=L(x)^{-1}\) is the inverse of a weighted graph Laplacian \(L(x)\), then locally
\[
d_W^2(x,x+\delta)=\langle \delta,G_W(x)\delta\rangle+o(\|\delta\|^2).
\]
Using an input-dependent additive noise model
\[
x'=x+\delta,\qquad \delta\sim \mathcal N(0,\Sigma_W(x)),\qquad \Sigma_W(x)=\eta^2L(x),
\]
the second-order expansion of the augmented risk yields
\[
\mathcal L_{\mathrm{aug}}(f)
=
\mathcal L(f)
+
\frac{\eta^2}{2}\,
\mathbb E_{(x,y)}
\Big[
\ell''(f(x),y)\,\nabla f(x)^\top L(x)\nabla f(x)
+
\ell'(f(x),y)\,\Delta_W f(x)
\Big]
+o(\eta^2).
\]
The stand-alone smoothness version is
\[
\mathcal L_{\mathrm{GLSW2}}(f)
:=
\frac12\,
\mathbb E_x\,
\mathbb E_{\delta\sim \mathcal N(0,\eta^2L(x))}
\big[\|f(x+\delta)-f(x)\|^2\big]
\approx
\frac{\eta^2}{2}\,
\mathbb E_x\,\operatorname{Tr}\!\big(J_f(x)^\top L(x)J_f(x)\big),
\]
with scalar specialization \(\epsilon^2\,\mathbb E_x[\nabla f(x)^\top L(x)\nabla f(x)]\) in the small-displacement expansion [1909.06860].

This formulation makes locality metric-dependent rather than patch-dependent. The operator \(L(x)\) privileges mass-transport directions on the pixel graph, so the resulting regularizer is anisotropic and input-dependent, unlike isotropic Jacobian penalties. The same paper reports empirical robustness improvements on CIFAR-10, including natural test error \(16.29\%\to 15.35\%\) and FGSM \(\epsilon=8/255\) robust test error \(82.22\%\to 30.20\%\) for W2 grad regularization, with convolutional implementations whose overhead is negligible relative to standard backpropagation [1909.06860].

A different local interpretation appears in Bayesian asymptotics. For a parametric family \(\{P_\vartheta:\vartheta\in\Theta\}\), the squared Wasserstein loss
\[
\ell_{W_2}(t,\vartheta):=W_2^2(P_t,P_\vartheta)
\]
is shown, near \(\vartheta_0\), to admit a local quadratic expansion
\[
\ell_{W_2}(t,\vartheta)
=
\langle t-\vartheta, A(\vartheta)(t-\vartheta)\rangle+\xi(t,\vartheta),
\]
with \(A(\vartheta_0)\) positive definite and \(\xi(t,\vartheta)=o(\|t-\vartheta\|^2)\) locally. In this setting, “generalized local squared \(W_2\) loss” refers to the locally quadratic structure of \(W_2^2\) as a function of parameters, which, combined with Bernstein–von Mises theory, yields asymptotic normality and efficiency of Bayes estimators under broad regularity conditions [2510.05645]. In one-dimensional models the paper also provides an explicit second-derivative formula involving the optimal map \(T_t^\vartheta(x)=F_\vartheta^{-1}(F_t(x))\).

A closely related exact reduction occurs for predictive density estimation in location and location-scale families. There, plug-in predictive densities form a complete class under \(W_2^2\), and the Bayes predictive density is the plug-in density evaluated at posterior means of location and scale parameters. In location-scale families with spherical scale, predictive risk reduces to
\[
\|\mu-\hat\mu\|^2+d(\sigma-\hat\sigma)^2,
\]
so the Wasserstein loss becomes a quadratic parameter loss [1904.02880].

## 5. Empirical neighborhood losses, time decoupling, and mixed-type outputs

In uncertainty quantification and inverse problems with latent randomness, the local squared Wasserstein-2 method matches conditional output laws in neighborhoods of the input. For \(y(x;\omega)=f(x,\omega)\) and an approximate model \(\hat y(x;\hat\omega)=\hat f(x,\hat\omega)\), the ideal objective is
\[
\tilde W_2^2(y,\hat y):=\int_D W_2^2(\mu_x,\hat\mu_x)\,\nu(dx),
\]
where \(\mu_x=P_{\mathrm{data}}(y|x)\) and \(\hat\mu_x=P_{\mathrm{model}}(y|x;\theta)\). Because \(\mu_x\) is unavailable, the empirical local objective is
\[
\tilde W_{2,\delta}^{2,e}(y,\hat y):=\int_D W_2^2(\mu_{x,\delta}^e,\hat\mu_{x,\delta}^e)\,\nu^e(dx)
\approx
\frac1N\sum_{i=1}^N W_2^2(\mu_{x_i,\delta}^e,\hat\mu_{x_i,\delta}^e),
\]
where neighborhoods are defined by \(|\tilde x-x|_x\le \delta\). The paper solves the discrete OT exactly with Python POT using `ot.emd2`, without entropic regularization in the main experiments, and proves
\[
\mathbb E\big[|\tilde W_2^2-\tilde W_{2,\delta}^{2,e}|\big]
\le
\frac{4M}{\sqrt N}
+
8CM\,\mathbb E[h(N(x,\delta),n)]
+
8\sqrt M\,L\delta,
\]
exhibiting the bias–variance trade-off in \(\delta\) [2406.06825].

For stochastic dynamical systems, locality is combined with time decoupling. Given local empirical conditional distributions \(\nu_{X_0,\delta}^e(t)\) and \(\hat\nu_{X_0,\delta}^e(t)\) at time \(t\), the local discrepancy is
\[
W_{2,\delta}^{2,e}(X(t),\hat X(t))
=
\int W_2^2\big(\nu_{X_0,\delta}^e(t),\hat\nu_{X_0,\delta}^e(t)\big)\,\nu_0^e(dX_0),
\]
and the time-decoupled loss is
\[
\tilde W_{2,\delta}^{2,e}(X,\hat X)
=
\int_0^T W_{2,\delta}^{2,e}(X(t),\hat X(t))\,dt
\approx
\sum_{i=0}^{n-1}W_{2,\delta}^{2,e}(X(t_i),\hat X(t_i))(t_{i+1}-t_i).
\]
This avoids OT on path space and is proved to be necessary for recovering the parameter distribution: under the stated Lipschitz and moment assumptions,
\[
\mathbb E[\tilde W_{2,\delta}^{2,e}(X,\hat X)]
\le
8C_0T\delta^2e^{C_0T}
+
\frac{6C_1}{C_0}Te^{C_0T}
\Big[
W_2^2(\mu,\hat\mu)
+
2C_3\mathbb E\big[h(N^\#(X_0;\delta),\ell)(\Theta_6^{1/3}+\hat\Theta_6^{1/3})\big]
\Big].
\]
The formulation is developed for ODEs, jump–diffusions, and SPDEs, with POT-based EMD in experiments and Adam or AdamW for training stochastic neural networks [2503.05068].

A further generalization addresses mixed outputs with both continuous and categorical components. On \(\mathcal Y=\mathbb R^{d_1}\times S_{d-d_1}\), the paper defines the squared mixed-type cost
\[
\|y\|^2
=
\lambda\sum_{i=1}^{d_1}y_i^2
+
\sum_{j=d_1+1}^d \hat\delta_{y_j,0},
\qquad
\hat\delta_{z,0}
=
\begin{cases}
4z^2,& |z|\le \frac12,\\
1,& |z|>\frac12,
\end{cases}
\]
and the generalized Wasserstein-2 distance \(\hat W_2\) by minimizing the expected value of this cost over couplings. The generalized local squared loss is
\[
\mathcal L_{\mathrm{GLSW2}}
\equiv
\hat W_{2,\delta}^{2,e}(y_x,\hat y_x)
:=
\int_D \hat W_2^2(f_{x,\delta}^e,\hat f_{x,\delta}^e)\,\nu^e(dx),
\]
with a time-averaged extension for spatiotemporal systems. For differentiable training, the categorical part is replaced by a surrogate \(\tilde\delta_{y_j,\hat y_j}\) together with
\[
\mathrm{round}_1(\hat y_j)=\hat y_j-(\hat y_j-\mathrm{round}(\hat y_j))\mathrm{.detach()},
\]
so gradients do not flow through the rounding gap. The paper proves a universal approximation theorem for stochastic neural networks under this generalized \(W_2\) metric and a finite-sample bound analogous in structure to the purely continuous case [2507.05143].

## 6. Related design patterns, applications, and limitations

Several adjacent research lines instantiate local squared \(W_2\) ideas in more specialized forms. In generative modeling via restricted convex potentials, the paper on restricted \(W_2\) approximation does not explicitly propose a local \(W_2^2\) objective, but it states that a local version can be built as
\[
W_{2,F}^{\mathrm{loc}}(P,Q)
=
\sum_{k=1}^K
W_{2,F_k}^2\big((\Phi_k\#P),(\Phi_k\#Q)\big),
\]
with local feature maps \(\Phi_k\) and local convex classes \(F_k\). The same work emphasizes restricted moment matching on tangent spaces and projected SGD with ICNN-based convex potentials [1902.07197]. In end-to-end \(W_2\) generative networks, the global quadratic-cost OT objective is trained with ICNN potentials and cycle-consistency regularization, and the provided adaptation to a local squared \(W_2\) loss weights correlations and cycle penalties over neighborhoods or kernels while preserving convexity in the transported variable [1909.13082].

In vision, locality can be imposed at the class, patch, or region level. For ordered single-label classification, the squared EMD loss
\[
W_2^2(p,q)=\sum_{k=1}^K\left(\sum_{i=1}^k(p_i-q_i)\right)^2
\]
penalizes errors according to a ground distance matrix \(D\), and the paper also proposes learning \(D\) from CNN features during training [1611.05916]. In superpixel-based segmentation, adjacent region histograms \(v_{\Omega_i},v_{\Omega_j}\) are compared by
\[
E_{ij}=d^2(v_{\Omega_i},v_{\Omega_j}),
\qquad
\kappa_{ij}=E_{ij}-h_i-h_j,
\]
and greedy adjacency-constrained merging produces a local graph-based squared \(W_2\) criterion over superpixels [2601.17071]. In texture interpolation, the spline formulation uses local feature maps \(F_m\) on image patches, constructs feature-space empirical measures \(\nu=(1/M)\sum\delta_{F_m[u]}\), and reports visually superior interpolation relative to metamorphosis splines that blend intensities rather than transporting mass in feature space [2302.10682].

The main limitations are likewise context-dependent. Entropic regularization introduces bias and smoothing, although it improves differentiability and numerical stability; this is emphasized both in spline computation and in OT-based learning more broadly [2302.10682] [1909.13082]. Higher-order spline terms can overshoot near key frames, with the anchored generalized barycenter and \(\delta>0\) recommended as stabilizers [2302.10682]. Neighborhood-based losses depend sensitively on the radius \(\delta\): too small a neighborhood leads to poor empirical estimation, while too large a neighborhood induces bias by mixing distinct conditional laws [2406.06825] [2503.05068] [2507.05143]. Time decoupling reduces computation but may discard temporal-correlation information when parameter effects are expressed primarily through trajectory correlations rather than time marginals [2503.05068]. In mixed-output settings, the choice of categorical ground penalty matters, and the paper notes sensitivity to the constant \(4\) in \(\hat\delta\) [2507.05143].

Taken together, these formulations establish a coherent but plural notion of generalized local squared Wasserstein-2 loss. The core object is always a squared \(W_2\)-type comparison, but the operative notion of locality changes with the modeling task: barycentric averaging in time, infinitesimal diffusion in input geometry, empirical conditioning in covariate neighborhoods, time-slice aggregation for stochastic dynamics, patchwise unbalanced transport, or local quadratic structure in parameter space. This diversity is not a defect of the term; it is a direct reflection of the fact that \(W_2\) supplies geometry, while “local” specifies where that geometry is sampled, linearized, or constrained.

Source: https://www.emergentmind.com/topics/generalized-local-squared-wasserstein-2-loss