Papers
Topics
Authors
Recent
Search
2000 character limit reached

Entropy-Regularized Linear Programming

Updated 20 November 2025
  • The paper introduces an entropy-regularized LP that adds a negative Shannon penalty to smooth the feasible region and enforce strict convexity.
  • It demonstrates exponential convergence to LP optima with non-asymptotic error bounds achieved via properties like weak convexity and the Sinkhorn iteration.
  • The approach underpins scalable algorithms for optimal transport and machine learning, balancing computational efficiency with trade-offs between accuracy and runtime.

An entropy-regularized linear programming approach augments a standard linear program (LP) with a negative Shannon entropy penalty. This method smooths the polyhedral feasible region, leads to strictly convex objectives, and critically underpins the scalability of algorithms for optimal transport and related large-scale optimization in machine learning. At its core, an entropy penalty enables exponentially fast, quantifiable convergence to LP optima while admitting algorithmic strategies (e.g., Sinkhorn iteration) with favorable computational and parallelization properties. This framework also provides non-asymptotic explicit error bounds, elucidates sharp trade-offs between accuracy and computational effort, and demonstrates fundamental limits on the achievable complexity for certain combinatorial LPs such as the assignment problem (Weed, 2018).

1. Classical Linear Programs and Entropic Penalties

A standard LP in minimization form is

(LP)min⁡x≥0  c⊤xsubject to  Ax=b,\text{(LP)}\quad \min_{x \geq 0} \; c^\top x \quad \text{subject to} \; A x = b,

where P:={x≥0:Ax=b}P := \{x \geq 0 : A x = b\} is assumed bounded and nonempty, and c⊤xc^\top x is not constant on PP. The entropy-regularized variant introduces a negative Shannon entropy penalty,

ent(x):=∑ixilog⁡(1/xi),\mathrm{ent}(x) := \sum_i x_i \log(1/x_i),

with regularization parameter η>0\eta > 0, transforming the objective to

(Pen)min⁡x  Fη(x):=c⊤x−η−1ent(x)subject to  Ax=b.\text{(Pen)}\quad \min_x \; F_\eta(x) := c^\top x - \eta^{-1} \mathrm{ent}(x) \quad \text{subject to} \; A x = b.

As η→∞\eta \rightarrow \infty, the penalty vanishes, recovering the original LP. For moderate η\eta, strong convexity of the entropic term facilitates efficient algorithms, notably the Sinkhorn method.

2. Quantitative Error Bounds and Exponential Convergence

Let f∗=min⁡x∈Pc⊤xf^* = \min_{x \in P} c^\top x and P:={x≥0:Ax=b}P := \{x \geq 0 : A x = b\}0. To relate the entropic and original optima, define the suboptimality gap P:={x≥0:Ax=b}P := \{x \geq 0 : A x = b\}1, the P:={x≥0:Ax=b}P := \{x \geq 0 : A x = b\}2-radius P:={x≥0:Ax=b}P := \{x \geq 0 : A x = b\}3, and the entropic radius P:={x≥0:Ax=b}P := \{x \geq 0 : A x = b\}4.

A non-asymptotic convergence theorem establishes

P:={x≥0:Ax=b}P := \{x \geq 0 : A x = b\}5

valid for any LP. Explicitly, if P:={x≥0:Ax=b}P := \{x \geq 0 : A x = b\}6,

P:={x≥0:Ax=b}P := \{x \geq 0 : A x = b\}7

The proof decomposes P:={x≥0:Ax=b}P := \{x \geq 0 : A x = b\}8 as a convex combination of optimal and suboptimal vertices and uses weak convexity properties of entropy. Notably, the exponential rate is optimal and matching lower bounds exist: for a rescaled simplex, with P:={x≥0:Ax=b}P := \{x \geq 0 : A x = b\}9, the rate c⊤xc^\top x0 is tight up to constants. No improvement is possible in the dependencies on c⊤xc^\top x1 (Weed, 2018).

3. Limitation: Assignment Problem and Complexity Barriers

Consider the c⊤xc^\top x2 assignment (minimum-cost perfect matching) LP: c⊤xc^\top x3 The Birkhoff polytope then has c⊤xc^\top x4. The exponential convergence theorem implies that to reach c⊤xc^\top x5-objective accuracy, one must set

c⊤xc^\top x6

Sinkhorn-type algorithms require c⊤xc^\top x7 time per run, resulting in a total complexity of c⊤xc^\top x8, which precludes the existence of a near-linear time (c⊤xc^\top x9) approximation scheme for the assignment problem by entropy-regularized means alone. Furthermore, if PP0, recovery of even a constant-factor approximate assignment is impossible (Weed, 2018).

4. Methodological Components and Key Lemmas

The analysis leverages several fundamental properties of the entropy function. Key results include:

  • Weak convexity: For any nonnegative PP1, and PP2,

PP3

with PP4.

  • Monotonicity and scalar bounds on the binary entropy facilitate the derivation of sharp fixed-point bounds for the convex combination weights in the optimality analysis.
  • The analysis holds uniformly for arbitrary LPs, not just for specific instances such as transport polytopes.

These structural insights underpin both the exponential convergence rate and the necessity for large regularization in combinatorially complex LPs.

5. Practical Implications for Machine Learning and Optimal Transport

In large-scale optimal transport (OT) and machine learning, entropy regularization via the Sinkhorn algorithm is widely adopted for computational expedience on GPUs and other parallel hardware. However, exact objective accuracy requires

PP5

For OT problems on size PP6, typically PP7, whence PP8 is required. Small PP9 yields fast but biased solutions, while increasing ent(x):=∑ixilog⁡(1/xi),\mathrm{ent}(x) := \sum_i x_i \log(1/x_i),0 reduces bias exponentially slowly. There is a trade-off: computationally efficient but approximate solutions when ent(x):=∑ixilog⁡(1/xi),\mathrm{ent}(x) := \sum_i x_i \log(1/x_i),1 is modest, and high-precision solutions only at high computational cost. In coarse ML applications where approximate distances suffice, moderate entropic regularization is typically acceptable. For exact recovery or fine-grained OT, the exponential rate in ent(x):=∑ixilog⁡(1/xi),\mathrm{ent}(x) := \sum_i x_i \log(1/x_i),2 governs achievable bias (Weed, 2018).

A critical lesson is that the entropic diameter (ent(x):=∑ixilog⁡(1/xi),\mathrm{ent}(x) := \sum_i x_i \log(1/x_i),3) and the data-dependent condition number (ent(x):=∑ixilog⁡(1/xi),\mathrm{ent}(x) := \sum_i x_i \log(1/x_i),4) jointly determine the optimal choice of ent(x):=∑ixilog⁡(1/xi),\mathrm{ent}(x) := \sum_i x_i \log(1/x_i),5. Future work aims to precisely estimate the energy spectrum of near-optima (the distribution of ent(x):=∑ixilog⁡(1/xi),\mathrm{ent}(x) := \sum_i x_i \log(1/x_i),6 and ent(x):=∑ixilog⁡(1/xi),\mathrm{ent}(x) := \sum_i x_i \log(1/x_i),7) for adaptive parameter tuning.

6. Summary and Outlook

The entropy-regularized LP approach yields exponentially fast, fully explicit convergence to LP optima across arbitrary problem instances, supported by sharp upper and lower bounds. There are foundational limitations for combinatorial LPs—for instance, the assignment problem cannot be solved in near-linear time by entropic smoothing alone. Nonetheless, the method underpins scalable algorithms for large-scale OT and machine learning, where practical trade-offs between accuracy and runtime must be balanced by tuning the regularization parameter in light of intrinsic geometric characteristics of the LP feasible region (Weed, 2018).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Entropy-Regularized Linear Programming Approach.