Almost-sharp O(k−1logk) convergence rate for the Sinkhorn algorithm in the asymptotically scalable case
Published 29 Apr 2026 in math.OC | (2604.26265v1)
Abstract: We prove that the Sinkhorn algorithm converges at a rate of O(k<sup>−1</sup>logk) in ℓ1-norm marginal error, in the asymptotically scalable case. This almost closes the gap between the lower bound Ω(k<sup>−1) (Qu et al., 2025) and the previously best known upper bound O(k<sup>−1/2) (Léger, 2021), and generalizes the analysis for the positive case by Dvurechensky et al. (2018).
The paper introduces an O(log k/k) convergence rate for the Sinkhorn algorithm in the asymptotically scalable regime, nearly matching the known lower bound of Ω(1/k).
It utilizes convex duality and a rate-function framework, leveraging the Dulmage–Mendelsohn decomposition to analyze non-attainable optima in degenerate cases.
The explicit rate guarantees provided are critical for real-world applications in sparse or ill-posed entropy-regularized optimal transport problems.
O(k−1logk) Convergence Rate for Sinkhorn in the Asymptotically Scalable Case
Introduction and Context
This work rigorously analyzes the convergence rate of the classical Sinkhorn algorithm for entropy-regularized optimal transport (EOT) in the general asymptotically scalable regime. This setting strictly generalizes the standard exactly scalable (matrix scaling) case, where strong regularity properties are not assumed and feasible EOT couplings may be singular or supported only on subsets defined by the cost matrix structure. The asymptotically scalable scenario is particularly relevant for sparse problems or instances with structural zeros (Cij=∞).
While earlier results establish exponential convergence when the input matrix A is positive and O(1/k) convergence in the exactly scalable case, the only known upper bound in the asymptotically scalable case prior to this work was O(1/k) in ℓ1 marginal error. A matching non-improvable lower bound of Ω(1/k) was known, leaving a significant gap in the non-exact setting.
Problem Formulation
Given marginals μ∈Δm, ν∈Δn, and a cost matrix C∈(R∪{∞})m×n, the entropy-regularized OT problem is solved via the Sinkhorn iteration—alternating Bregman projections onto the two sets of prescribed marginals in the space of positive matrices weighted by Cij=∞0. For the EOT dual, the natural performance metric is the Cij=∞1-norm error in matching marginals: Cij=∞2.
The matrix Cij=∞3 is asymptotically Cij=∞4-scalable if the dual problem is bounded below, i.e., Cij=∞5, which encompasses all cases where the EOT is feasible, even if the optimizer is only approached in the limit and possibly singular.
Theoretical Contributions
The main result is an upper bound on the marginal error of the Sinkhorn algorithm when Cij=∞6 is only asymptotically scalable:
Cij=∞7
This result matches the known lower bound Cij=∞8 up to a logarithmic factor, resolving a long-standing question by almost closing the gap between lower and upper convergence rates in this regime.
The explicit statement (Theorem 1) provides: Cij=∞9
where A0 is the diameter of a DAG associated with the Dulmage–Mendelsohn decomposition, A1 and A2 are explicit constants determined by the cost and marginals, and A3 is the minimal entry in A4.
Technical Approach
The analysis innovates over prior work in several ways:
Extends convergence analysis based on convex duality and Bregman proximal point geometry to settings without exact attainability.
Introduces a rate-function perspective: convergence is reduced to bounding a specific rate function A5 for vanishing regularization A6.
Constructs near-minimizers of the dual with controlled oscillation, leveraging a decompositional structure (generalized Dulmage–Mendelsohn) on the induced bipartite graph of admissible couplings.
Connects the logarithmic penalty to the structure of the interaction DAG between DM components.
Provides fully explicit constants and uniformity in the model parameters, robust to degenerate and sparse cases.
Comparative Results
The previously tightest upper bound was A7, attainable via classical convex optimization techniques. The polynomial time bound presented here improves this to A8, which is only a logarithmic factor above the minimax rate. Moreover, the analysis is shown to be tight (up to the A9 factor) via explicit examples, e.g., in O(1/k)0 block-matrix settings with marginally feasible but non-exact scaling, as illustrated by Soules (1991).
Implications
Practical:
This result provides explicit rate guarantees for the Sinkhorn algorithm in the broad class of real-world EOT and matrix scaling problems with singularity or sparsity. As many applications in data science, generative models, and domain adaptation involve such ill-posed or semi-feasible instances, these guarantees are essential for algorithmic reliability.
Theoretical:
The result demonstrates that asymptotic polynomial O(1/k)1 convergence is almost always achieved in practice, barring the logarithmic term, even when exact dual minimizers do not exist. The connection to the DM decomposition and the explicit role of the graph structure points toward a precise geometric understanding of non-attainability in OT regularization. The rate-function framework developed may have further applications in the analysis of Bregman proximal algorithms with non-attainable minimizers.
Limits and Future Work:
The O(1/k)2 gap between lower and upper rates is shown to be inherent to the analysis method, even if not generally tight for every instance. Removing this gap—possibly using a different analytic strategy than the rate-function technique—remains an open question. Additionally, the result suggests exploring more refined structural decompositions of the cost support graph to directly obtain tighter, possibly instance-optimal, convergence rates.
Conclusion
This work advances the theory of entropy-regularized OT by almost closing the convergence rate gap for Sinkhorn in the asymptotically scalable case. The O(1/k)3 upper bound is the sharpest known in this general regime, and the analysis links the algorithmic rate to precise combinatorial properties of the problem instance. These results solidify the understanding of Sinkhorn's behavior in degenerate and sparse problems, which are pervasive in modern applications of optimal transport and matrix scaling (2604.26265).
“Emergent Mind helps me see which AI papers have caught fire online.”
Philip
Creator, AI Explained on YouTube
Sign up for free to explore the frontiers of research
Discover trending papers, chat with arXiv, and track the latest research shaping the future of science and technology.Discover trending papers, chat with arXiv, and more.