- The paper establishes a sharp O(1/k) convergence rate for Sinkhorn algorithm iterates in entropy-regularized optimal transport problems.
- It employs a two-step methodology by first performing local analysis on the optimal support and then bootstrapping these results to achieve global convergence with explicit constants.
- The analysis refines previous bounds and provides actionable insights for parameter tuning and understanding support structures in scaling applications.
Sharp O(1/k) Convergence Rate for the Sinkhorn Algorithm via Local Analysis
Overview and Problem Statement
The paper "Sharp O(1/k) convergence rate for the Sinkhorn algorithm via a local analysis" (2606.28973) addresses the rate of convergence of the Sinkhorn algorithm, a fundamental method for solving entropy-regularized optimal transport (EOT) and matrix scaling problems. The main contribution is a non-asymptotic, sharp O(1/k) convergence guarantee (in ℓ1 marginal error and relative entropy) for the Sinkhorn iterates in the general, asymptotically scalable case. The analysis separates the behavior on the optimal support of the transport plan from the rest of the feasible support, yielding a local convergence bound that is bootstrapped to global rates using previous nearly-sharp global results.
Let μ∈Δm and ν∈Δn be strictly positive probability vectors, and C∈(R∪{∞})m×n a cost matrix with admissible support E⊆{1,…,m}×{1,…,n}. The EOT problem is: π∈ΔEmin(i,j)∈E∑Cijπij+τKL(π∥μ⊗ν)subject toX#π=μ, Y#π=ν
The Sinkhorn algorithm alternately rescales rows and columns of the coupling to achieve the desired marginals, producing iterates πk. The main convergence metric is the O(1/k)0-marginal error: O(1/k)1
Review of Prior Results and Sharpness
Classical results establish that Sinkhorn iterates converge (O(1/k)2) whenever a feasible solution exists. However, explicit non-asymptotic rates in general support patterns were incomplete. Previous work (\citet{leger2021gradient}) provided an O(1/k)3 bound, while more recent work (\citet{wang2026almost}) improved this to O(1/k)4. Soules constructed an example instance achieving only O(1/k)5, which puts a lower bound on possible convergence rates. Recently, Qu et al.\ demonstrated that O(1/k)6 necessarily in asymptotically scalable, but not exactly scalable, instances. This pinned down the sharpness question up to a log factor.
The contribution of this paper is the closure of this gap: it proves that O(1/k)7 in full generality, matching the known lower bound. Moreover, the results subsume and refine all previous nonasymptotic convergence bounds for the Sinkhorn algorithm in this setting.
Technical Approach and Main Results
The core technical novelty lies in a two-step approach leveraging the structure of the EOT problem’s support graph:
- Local Analysis: The proof conducts a fine-grained analysis on the bipartite graph associated with the EOT. It distinguishes between edges in the optimal support O(1/k)8 (edges supporting strictly positive mass in the unique optimal plan O(1/k)9) and those outside O(1/k)0. The local convergence result proves that, once the Sinkhorn iterates are sufficiently close to O(1/k)1 in the joint KL divergence, convergence proceeds at the sharp O(1/k)2 rate. The argument exploits:
- The decomposition of the support via the Dulmage-Mendelsohn (DM) partition and interaction DAG;
- Careful propagation of errors across blocks in the graph, with explicit dependency on the minimal positive entry of O(1/k)3, block diameters, and other structural graph parameters.
- Global Bootstrapping: To lift the local rate to a global guarantee, the proof uses the author's previous O(1/k)4 bound to quantify the number of iterates needed to reach the local region where the O(1/k)5 convergence kicks in. Explicit constants are given, depending on the cost matrix’s oscillation, regularization parameter, graph parameters, and minimal marginals.
The main theorem (see Theorem 1 and 3 in the paper) states that there exist explicit constants (dependent only on problem sparsity, minimal marginals, cost oscillation, and the regularization parameter O(1/k)6) such that for all large enough O(1/k)7: O(1/k)8
This also holds for the discrepancy in joint relative entropy, demonstrating that the O(1/k)9 rate is attained for the marginal error and suboptimality simultaneously.
Numerical and Theoretical Strength
The guarantees include:
- Sharpness: The ℓ10 rate is nonimprovable in the general scalable setting, as demonstrated by explicit problem instances.
- Explicit Constants: Rates depend explicitly yet succinctly on ℓ11, cost oscillation, and minimal marginals, as well as combinatorial properties of the support graph.
- General Support Patterns: The analysis covers all admissible ℓ12, not just full or dense supports, subsuming previous results.
- Separation of Local vs Global Behavior: The methodology shows that the log factor gap in prior upper bounds was a transient artifact; after a modest burn-in, sharp rates are achieved globally.
Implications and Future Directions
Practical Impact
This result clarifies and tightens the theoretical understanding of the computational complexity of Sinkhorn and related algorithms for regularized optimal transport and matrix scaling. It provides practitioners with sharper bounds for algorithmic performance, especially in large-scale or structured OT problems where sparse supports are common (e.g., in graphics, structured matching, or subset transport). The explicit dependency on ℓ13 and structural graph parameters also informs the selection and tuning of entropic regularization for optimal computational tradeoffs.
Theoretical Developments
The methodology leverages structural graph-theoretic decompositions (Dulmage-Mendelsohn) and bootstrapping between global and local convergence—a framework that can potentially be generalized to other multiterminal OT, multi-marginal, or regularized variants.
Open problems articulated by the author include:
- Generalizing the sharp ℓ14 rate to multi-marginal or unbalanced Sinkhorn settings, where the decomposability and optimal support characterization are less direct.
- Investigating whether the same local-to-global bootstrapping can apply when additional constraints (such as unbalance or dynamical regularization) are present.
- Further tightening constants and understanding pre-asymptotic behavior in practical regimes, especially for very sparse or degenerate cases.
Conclusion
This work establishes the sharp ℓ15 nonasymptotic convergence rate for the Sinkhorn algorithm in the most general scalable setting, reconciling upper and lower bounds and removing previous logarithmic gaps. The analysis highlights the value of local, graph-structural decomposition in proving nonasymptotic complexity results for iterative projection and scaling algorithms. The techniques and guarantees presented set a new theoretical benchmark for entropic OT and scaling algorithms, and open the way for similar results in more general transport and regularized matching domains (2606.28973).