Papers
Topics
Authors
Recent
Search
2000 character limit reached

Divergence-Regularized Optimal Transport

Updated 25 December 2025
  • Divergence-Regularized Optimal Transport is a framework that incorporates convex divergence penalties (e.g., KL, Tsallis) to smooth the classical OT problem and facilitate convex relaxations.
  • It leverages dual formulations and specialized numerical algorithms such as Sinkhorn, mirror descent, and Bregman projections to achieve efficient, robust, and scalable optimization.
  • This approach improves sample complexity, statistical convergence, and robustness, with applications in robust learning, statistical inference, and generative modeling.

Divergence-Regularized Optimal Transport is a broad framework for convex relaxations of the classical optimal transport (OT) problem, in which the original linear programming problem is smoothed or "regularized" by a convex divergence term acting on the space of couplings. This methodology encompasses entropic regularization (based on Kullback-Leibler (KL) divergence), Tsallis, Rényi, β\beta-divergence, Bregman, and other ff-divergence regularizations, as well as kernel methods such as MMD-regularized OT. It enables scalable numerical algorithms, admits rigorous statistical theory, improves robustness and sample complexity, and provides a unifying interface between geometry, information, convex analysis, and applications, especially in high-dimensional data science, statistical inference, and generative modeling.

1. Mathematical Formulation of Divergence-Regularized OT

Let (X1,B1,μ)(X_1,\mathcal{B}_1,\mu) and (X2,B2,ν)(X_2,\mathcal{B}_2,\nu) be Polish probability spaces and c:X1×X2[0,+)c : X_1 \times X_2 \to [0, +\infty) a lower-semicontinuous cost. The set of couplings is

Π(μ,ν)={πP(X1×X2):πX1=μ,πX2=ν}.\Pi(\mu, \nu) = \{ \pi \in \mathcal P(X_1 \times X_2) : \pi_{X_1} = \mu,\, \pi_{X_2} = \nu \}.

The unregularized Monge–Kantorovich problem is

OT(μ,ν)=infπΠ(μ,ν)X1×X2c(x,y)dπ(x,y).\mathrm{OT}(\mu, \nu) = \inf_{\pi \in \Pi(\mu, \nu)} \int_{X_1 \times X_2} c(x, y)\,d\pi(x, y).

Divergence regularization introduces an additive penalty on π\pi with respect to a reference, usually μν\mu\otimes\nu, via a convex divergence DfD_f:

ff0

with

ff1

for convex ff2 satisfying ff3. Notable cases:

  • KL (entropic): ff4
  • Tsallis: ff5 for ff6
  • ff7-divergence: ff8 as in (Nakamura et al., 2022)
  • Bregman, Rényi, MMD, and others

Dual formulations are given via Fenchel–Rockafellar conjugacy using the convex dual ff9. For Tsallis-regularized OT (Suguro et al., 2023), e.g.,

(X1,B1,μ)(X_1,\mathcal{B}_1,\mu)0

2. Classes of Divergences, Interpolation, and Limiting Behavior

Divergence regularizations span a wide spectrum. KL, Tsallis, (X1,B1,μ)(X_1,\mathcal{B}_1,\mu)1-divergence, Bregman divergences, and Rényi divergences (for (X1,B1,μ)(X_1,\mathcal{B}_1,\mu)2) are prominent:

  • KL ((X1,B1,μ)(X_1,\mathcal{B}_1,\mu)3, (X1,B1,μ)(X_1,\mathcal{B}_1,\mu)4): recovers the entropic regularizer and Sinkhorn algorithm.
  • Tsallis ((X1,B1,μ)(X_1,\mathcal{B}_1,\mu)5): allows polynomial penalty structure, fusing sparsity with smoothness, converges to KL as (X1,B1,μ)(X_1,\mathcal{B}_1,\mu)6; as (X1,B1,μ)(X_1,\mathcal{B}_1,\mu)7, recovers classic OT (Suguro et al., 2023).
  • (X1,B1,μ)(X_1,\mathcal{B}_1,\mu)8-divergence (Nakamura et al., 2022): interpolates entropic (KL) and robust hard-thresholding regimes. Robust to outliers for (X1,B1,μ)(X_1,\mathcal{B}_1,\mu)9.
  • Rényi divergence (Bresch et al., 2024): for (X2,B2,ν)(X_2,\mathcal{B}_2,\nu)0 recovers KL; for (X2,B2,ν)(X_2,\mathcal{B}_2,\nu)1 recovers unregularized OT. Not an (X2,B2,ν)(X_2,\mathcal{B}_2,\nu)2-divergence/Bregman distance, but admits strict convexity, metrization, and symmetry.

This interpolation property allows one to tune regularization parameters (e.g. (X2,B2,ν)(X_2,\mathcal{B}_2,\nu)3 in Rényi or (X2,B2,ν)(X_2,\mathcal{B}_2,\nu)4 in Tsallis) to squeeze the regularized solution toward the true OT plan or, inversely, to maximize smoothness for tractable Sinkhorn-like computation.

3. Convergence Rates and (X2,B2,ν)(X_2,\mathcal{B}_2,\nu)5-Convergence

Central theoretical results concern the vanishing regularization limit (X2,B2,ν)(X_2,\mathcal{B}_2,\nu)6:

  • The functional (X2,B2,ν)(X_2,\mathcal{B}_2,\nu)7 (X2,B2,ν)(X_2,\mathcal{B}_2,\nu)8-converges to the unregularized (X2,B2,ν)(X_2,\mathcal{B}_2,\nu)9 (Suguro et al., 2023, Eckstein et al., 2022).
  • Quantization and shadow coupling arguments yield explicit convergence rates: in the entropic case (KL), the gap c:X1×X2[0,+)c : X_1 \times X_2 \to [0, +\infty)0 for typical costs and dimension (Suguro et al., 2023, Eckstein et al., 2022).
  • For Tsallis regularization c:X1×X2[0,+)c : X_1 \times X_2 \to [0, +\infty)1, rates slow to polynomial decay in c:X1×X2[0,+)c : X_1 \times X_2 \to [0, +\infty)2: c:X1×X2[0,+)c : X_1 \times X_2 \to [0, +\infty)3, and the KL case is provably optimal among all c:X1×X2[0,+)c : X_1 \times X_2 \to [0, +\infty)4 (Suguro et al., 2023).

The limits in other regularization families parallel this pattern—c:X1×X2[0,+)c : X_1 \times X_2 \to [0, +\infty)5-Rényi divergence regularization recovers the hard OT plan for c:X1×X2[0,+)c : X_1 \times X_2 \to [0, +\infty)6 without the numerical instability attendant to c:X1×X2[0,+)c : X_1 \times X_2 \to [0, +\infty)7 in classical entropic regularization (Bresch et al., 2024).

4. Numerical Algorithms and Computation

Divergence regularization transforms the OT problem into a strictly convex optimization, admitting scalable algorithms:

Strict convexity and superlinear growth of divergence ensure unique minimizers, strong duality, and global convergence, subject to compactness and regularity.

5. Regularized OT: Statistical Theory and Sample Complexity

Divergence regularization fundamentally shapes the statistical behavior of empirical OT estimators:

  • Parametric rates (c:X1×X2[0,+)c : X_1 \times X_2 \to [0, +\infty)9) for all Π(μ,ν)={πP(X1×X2):πX1=μ,πX2=ν}.\Pi(\mu, \nu) = \{ \pi \in \mathcal P(X_1 \times X_2) : \pi_{X_1} = \mu,\, \pi_{X_2} = \nu \}.0-divergences: Provided the cost is bounded and Π(μ,ν)={πP(X1×X2):πX1=μ,πX2=ν}.\Pi(\mu, \nu) = \{ \pi \in \mathcal P(X_1 \times X_2) : \pi_{X_1} = \mu,\, \pi_{X_2} = \nu \}.1 is Π(μ,ν)={πP(X1×X2):πX1=μ,πX2=ν}.\Pi(\mu, \nu) = \{ \pi \in \mathcal P(X_1 \times X_2) : \pi_{X_1} = \mu,\, \pi_{X_2} = \nu \}.2, empirical regularized OT achieves the parametric rate for sample complexity, in sharp contrast with the curse of dimensionality intrinsic to unregularized OT (Yang et al., 2 Oct 2025, González-Sanz et al., 7 May 2025, Bayraktar et al., 2022).
  • Central Limit Theorems: Limiting Gaussian distributions for the regularized OT cost, plan, and dual potentials, in both one- and two-sample regimes, with explicit covariance formulas (Klatt et al., 2018, Bigot et al., 2017, Yang et al., 2 Oct 2025, González-Sanz et al., 7 May 2025).
  • Bootstrap Consistency: Ordinary Π(μ,ν)={πP(X1×X2):πX1=μ,πX2=ν}.\Pi(\mu, \nu) = \{ \pi \in \mathcal P(X_1 \times X_2) : \pi_{X_1} = \mu,\, \pi_{X_2} = \nu \}.3-out-of-Π(μ,ν)={πP(X1×X2):πX1=μ,πX2=ν}.\Pi(\mu, \nu) = \{ \pi \in \mathcal P(X_1 \times X_2) : \pi_{X_1} = \mu,\, \pi_{X_2} = \nu \}.4 bootstrap is valid for empirical divergence-regularized OT, enabling statistical inference and confidence bands (Bigot et al., 2017, Klatt et al., 2018).
  • Stability and Regularity: Quantitative bounds on the change in optimizers under marginal perturbations, strengthening robustness.
  • Intrinsic Dimension and Smoothness: Fast rates are achievable depending on the regularity of cost/divergence and the “intrinsic dimension” of the data (Bayraktar et al., 2022).

6. Applications: Robustness, Inference, and Geometry

Divergence-regularized OT is used extensively across disciplines:

  • Robust Learning: Π(μ,ν)={πP(X1×X2):πX1=μ,πX2=ν}.\Pi(\mu, \nu) = \{ \pi \in \mathcal P(X_1 \times X_2) : \pi_{X_1} = \mu,\, \pi_{X_2} = \nu \}.5-potential regularization prevents mass transport to outliers, achieving statistical performance superior to entropic OT in contaminated data (Nakamura et al., 2022).
  • Ecological inference: Tsallis-regularized OT offers state-of-the-art marginal reconstruction for political science datasets (Muzellec et al., 2016).
  • Kernel and RKHS-based OT: MMD regularization interpolates classical OT and kernel distances, combining sample efficiency and ground-metric geometry (Manupriya et al., 2020).
  • Generative modeling: Proximal OT divergences provide tractable interpolants between GAN-style Π(μ,ν)={πP(X1×X2):πX1=μ,πX2=ν}.\Pi(\mu, \nu) = \{ \pi \in \mathcal P(X_1 \times X_2) : \pi_{X_1} = \mu,\, \pi_{X_2} = \nu \}.6-divergences and OT distances, governing flows in probability space (Baptista et al., 17 May 2025, Birrell et al., 2023).
  • Distributionally Robust Optimization: Infimal-convolution (OT-regularized divergence) sets define ambiguity sets that generalize Wasserstein and Π(μ,ν)={πP(X1×X2):πX1=μ,πX2=ν}.\Pi(\mu, \nu) = \{ \pi \in \mathcal P(X_1 \times X_2) : \pi_{X_1} = \mu,\, \pi_{X_2} = \nu \}.7-divergence DRO schemes (Birrell et al., 2023, Baptista et al., 17 May 2025).

Empirical results show that, for moderate regularization, sparser couplings (Tsallis, Rényi, Π(μ,ν)={πP(X1×X2):πX1=μ,πX2=ν}.\Pi(\mu, \nu) = \{ \pi \in \mathcal P(X_1 \times X_2) : \pi_{X_1} = \mu,\, \pi_{X_2} = \nu \}.8) yield OT plans closer to the ground truth than KL-regularized schemes, particularly in real-world inference tasks (Bresch et al., 2024).

7. Comparisons, Limitations, and Generalizations

  • KL regularization: Fast convergence, maximal smoothness, and computational simplicity via the Sinkhorn algorithm; cost bias vanishes logarithmically as Π(μ,ν)={πP(X1×X2):πX1=μ,πX2=ν}.\Pi(\mu, \nu) = \{ \pi \in \mathcal P(X_1 \times X_2) : \pi_{X_1} = \mu,\, \pi_{X_2} = \nu \}.9.
  • Tsallis/OT(μ,ν)=infπΠ(μ,ν)X1×X2c(x,y)dπ(x,y).\mathrm{OT}(\mu, \nu) = \inf_{\pi \in \Pi(\mu, \nu)} \int_{X_1 \times X_2} c(x, y)\,d\pi(x, y).0 regularization: Allows sparse couplings and variable convergence rates polynomial in OT(μ,ν)=infπΠ(μ,ν)X1×X2c(x,y)dπ(x,y).\mathrm{OT}(\mu, \nu) = \inf_{\pi \in \Pi(\mu, \nu)} \int_{X_1 \times X_2} c(x, y)\,d\pi(x, y).1; bias decays slower than KL (Suguro et al., 2023, Muzellec et al., 2016, González-Sanz et al., 7 May 2025).
  • Rényi regularization: Enables interpolation from OT to KL without numerical instability, yielding plans that empirically outperform both (Bresch et al., 2024).
  • General OT(μ,ν)=infπΠ(μ,ν)X1×X2c(x,y)dπ(x,y).\mathrm{OT}(\mu, \nu) = \inf_{\pi \in \Pi(\mu, \nu)} \int_{X_1 \times X_2} c(x, y)\,d\pi(x, y).2-divergence/Bregman regularization: All known convex regularizers with suitable smoothness yield strict convexity, well-behaved limits, statistical efficiency, and strong convergence theory.
  • Extensions: Multi-marginal divergence-regularized OT, barycenters, unbalanced OT, and models with spatially varying divergences (e.g., homogeneous UROT/OT with boundary (Lacombe, 2022)), and optimal transport-regularized divergences (infimal convolution) (Baptista et al., 17 May 2025).

Sharpness of rates and uniform statistical bounds depend critically on the specific divergence, cost regularity, and data geometry. KL regularization remains optimal in terms of fastest vanishing bias, but non-entropy divergences enable practical gains in robustness and inference accuracy.


References

Definition Search Book Streamline Icon: https://streamlinehq.com
References (19)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Divergence-Regularized Optimal Transport.