Papers
Topics
Authors
Recent
Search
2000 character limit reached

Almost-sharp O(k1logk)O(k^{-1} \log k) convergence rate for the Sinkhorn algorithm in the asymptotically scalable case

Published 29 Apr 2026 in math.OC | (2604.26265v1)

Abstract: We prove that the Sinkhorn algorithm converges at a rate of O(k<sup>1</sup>logk)O(k<sup>{-1}</sup> \log k) in 1\ell_1-norm marginal error, in the asymptotically scalable case. This almost closes the gap between the lower bound Ω(k<sup>1)Ω(k<sup>{-1}) (Qu et al., 2025) and the previously best known upper bound O(k<sup>1/2)O(k<sup>{-1/2}) (Léger, 2021), and generalizes the analysis for the positive case by Dvurechensky et al. (2018).

Authors (1)

Summary

  • The paper introduces an O(log k/k) convergence rate for the Sinkhorn algorithm in the asymptotically scalable regime, nearly matching the known lower bound of Ω(1/k).
  • It utilizes convex duality and a rate-function framework, leveraging the Dulmage–Mendelsohn decomposition to analyze non-attainable optima in degenerate cases.
  • The explicit rate guarantees provided are critical for real-world applications in sparse or ill-posed entropy-regularized optimal transport problems.

O(k1logk)O(k^{-1} \log k) Convergence Rate for Sinkhorn in the Asymptotically Scalable Case

Introduction and Context

This work rigorously analyzes the convergence rate of the classical Sinkhorn algorithm for entropy-regularized optimal transport (EOT) in the general asymptotically scalable regime. This setting strictly generalizes the standard exactly scalable (matrix scaling) case, where strong regularity properties are not assumed and feasible EOT couplings may be singular or supported only on subsets defined by the cost matrix structure. The asymptotically scalable scenario is particularly relevant for sparse problems or instances with structural zeros (Cij=C_{ij} = \infty).

While earlier results establish exponential convergence when the input matrix AA is positive and O(1/k)O(1/k) convergence in the exactly scalable case, the only known upper bound in the asymptotically scalable case prior to this work was O(1/k)O(1/\sqrt{k}) in 1\ell_1 marginal error. A matching non-improvable lower bound of Ω(1/k)\Omega(1/k) was known, leaving a significant gap in the non-exact setting.

Problem Formulation

Given marginals μΔm\mu \in \Delta_m, νΔn\nu \in \Delta_n, and a cost matrix C(R{})m×nC \in (\mathbb{R} \cup \{\infty\})^{m \times n}, the entropy-regularized OT problem is solved via the Sinkhorn iteration—alternating Bregman projections onto the two sets of prescribed marginals in the space of positive matrices weighted by Cij=C_{ij} = \infty0. For the EOT dual, the natural performance metric is the Cij=C_{ij} = \infty1-norm error in matching marginals: Cij=C_{ij} = \infty2.

The matrix Cij=C_{ij} = \infty3 is asymptotically Cij=C_{ij} = \infty4-scalable if the dual problem is bounded below, i.e., Cij=C_{ij} = \infty5, which encompasses all cases where the EOT is feasible, even if the optimizer is only approached in the limit and possibly singular.

Theoretical Contributions

The main result is an upper bound on the marginal error of the Sinkhorn algorithm when Cij=C_{ij} = \infty6 is only asymptotically scalable:

Cij=C_{ij} = \infty7

This result matches the known lower bound Cij=C_{ij} = \infty8 up to a logarithmic factor, resolving a long-standing question by almost closing the gap between lower and upper convergence rates in this regime.

The explicit statement (Theorem 1) provides: Cij=C_{ij} = \infty9 where AA0 is the diameter of a DAG associated with the Dulmage–Mendelsohn decomposition, AA1 and AA2 are explicit constants determined by the cost and marginals, and AA3 is the minimal entry in AA4.

Technical Approach

The analysis innovates over prior work in several ways:

  • Extends convergence analysis based on convex duality and Bregman proximal point geometry to settings without exact attainability.
  • Introduces a rate-function perspective: convergence is reduced to bounding a specific rate function AA5 for vanishing regularization AA6.
  • Constructs near-minimizers of the dual with controlled oscillation, leveraging a decompositional structure (generalized Dulmage–Mendelsohn) on the induced bipartite graph of admissible couplings.
  • Connects the logarithmic penalty to the structure of the interaction DAG between DM components.
  • Provides fully explicit constants and uniformity in the model parameters, robust to degenerate and sparse cases.

Comparative Results

The previously tightest upper bound was AA7, attainable via classical convex optimization techniques. The polynomial time bound presented here improves this to AA8, which is only a logarithmic factor above the minimax rate. Moreover, the analysis is shown to be tight (up to the AA9 factor) via explicit examples, e.g., in O(1/k)O(1/k)0 block-matrix settings with marginally feasible but non-exact scaling, as illustrated by Soules (1991).

Implications

Practical:

This result provides explicit rate guarantees for the Sinkhorn algorithm in the broad class of real-world EOT and matrix scaling problems with singularity or sparsity. As many applications in data science, generative models, and domain adaptation involve such ill-posed or semi-feasible instances, these guarantees are essential for algorithmic reliability.

Theoretical:

The result demonstrates that asymptotic polynomial O(1/k)O(1/k)1 convergence is almost always achieved in practice, barring the logarithmic term, even when exact dual minimizers do not exist. The connection to the DM decomposition and the explicit role of the graph structure points toward a precise geometric understanding of non-attainability in OT regularization. The rate-function framework developed may have further applications in the analysis of Bregman proximal algorithms with non-attainable minimizers.

Limits and Future Work:

The O(1/k)O(1/k)2 gap between lower and upper rates is shown to be inherent to the analysis method, even if not generally tight for every instance. Removing this gap—possibly using a different analytic strategy than the rate-function technique—remains an open question. Additionally, the result suggests exploring more refined structural decompositions of the cost support graph to directly obtain tighter, possibly instance-optimal, convergence rates.

Conclusion

This work advances the theory of entropy-regularized OT by almost closing the convergence rate gap for Sinkhorn in the asymptotically scalable case. The O(1/k)O(1/k)3 upper bound is the sharpest known in this general regime, and the analysis links the algorithmic rate to precise combinatorial properties of the problem instance. These results solidify the understanding of Sinkhorn's behavior in degenerate and sparse problems, which are pervasive in modern applications of optimal transport and matrix scaling (2604.26265).

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.