Papers
Topics
Authors
Recent
Search
2000 character limit reached

Difference of Convex Functions Algorithm (DCA)

Updated 29 January 2026
  • Difference-of-Convex functions Algorithm (DCA) is a method that represents a nonconvex function as the difference of two convex functions to simplify optimization.
  • The Boosted DC Algorithm (BDCA) enhances classical DCA by incorporating an extrapolation step, leading to significant convergence speedups and robust descent properties.
  • BDCA is practically effective for linearly constrained DC programming, with rigorous convergence guarantees to KKT points and demonstrated efficiency in applications like quadratic programming and copositivity detection.

A difference-of-convex (DC) function is any function that can be represented as the difference of two convex functions, i.e. ϕ(x)=g(x)h(x)\phi(x) = g(x) - h(x) with g,hg, h convex and proper. The Difference of Convex functions Algorithm (DCA) is a foundational method for DC programming, iteratively linearizing the “concave” part and minimizing a convex surrogate. The Boosted DC Algorithm (BDCA) extends classical DCA by incorporating an extrapolation step along the DCA descent direction, typically using a line-search. This simple but effective acceleration mechanism yields provable and substantial improvements in convergence properties and empirical performance, enabling broad applicability to linearly constrained DC programs and quadratic programming with box or trust-region constraints. BDCA guarantees monotonic descent, generalizes to large-scale and nonsmooth settings, and provides rigorous convergence to Karush–Kuhn–Tucker points under Slater-type conditions, with geometric rates in specific quadratic scenarios (Artacho et al., 2019).

1. Problem Formulation and DC Decomposition

The relevant problem class is linearly-constrained DC programming: minxRn  ϕ(x):=g(x)h(x)s.t.  ai,xbi,  i=1,,p,\min_{x\in\mathbb{R}^n} \; \phi(x) := g(x) - h(x) \quad\text{s.t.}\;\langle a_i, x\rangle \leq b_i, \; i=1, \dots, p, where g,h:RnR{+}g,h:\mathbb{R}^n \rightarrow \mathbb{R}\cup\{+\infty\} are proper, closed, convex; gg is differentiable, and both gg and hh are ρ\rho-strongly convex for some ρ>0\rho > 0. The feasible set is a polyhedron F={xai,xbi}F = \{x \mid \langle a_i, x \rangle \leq b_i\}, and the problem is equivalently

g,hg, h0

with g,hg, h1 the indicator of g,hg, h2. Analysis assumes:

  • (A1) g,hg, h3, g,hg, h4 are g,hg, h5-strongly convex,
  • (A2) g,hg, h6 subdifferentiable everywhere, g,hg, h7 g,hg, h8,
  • (A3) Slater: there exists g,hg, h9 strictly feasible.

2. Classical DCA: Structure and Descent

At each iteration minxRn  ϕ(x):=g(x)h(x)s.t.  ai,xbi,  i=1,,p,\min_{x\in\mathbb{R}^n} \; \phi(x) := g(x) - h(x) \quad\text{s.t.}\;\langle a_i, x\rangle \leq b_i, \; i=1, \dots, p,0:

  • Select minxRn  ϕ(x):=g(x)h(x)s.t.  ai,xbi,  i=1,,p,\min_{x\in\mathbb{R}^n} \; \phi(x) := g(x) - h(x) \quad\text{s.t.}\;\langle a_i, x\rangle \leq b_i, \; i=1, \dots, p,1,
  • Form the convex surrogate: minxRn  ϕ(x):=g(x)h(x)s.t.  ai,xbi,  i=1,,p,\min_{x\in\mathbb{R}^n} \; \phi(x) := g(x) - h(x) \quad\text{s.t.}\;\langle a_i, x\rangle \leq b_i, \; i=1, \dots, p,2,
  • Compute its unique minimizer: minxRn  ϕ(x):=g(x)h(x)s.t.  ai,xbi,  i=1,,p,\min_{x\in\mathbb{R}^n} \; \phi(x) := g(x) - h(x) \quad\text{s.t.}\;\langle a_i, x\rangle \leq b_i, \; i=1, \dots, p,3,
  • Set minxRn  ϕ(x):=g(x)h(x)s.t.  ai,xbi,  i=1,,p,\min_{x\in\mathbb{R}^n} \; \phi(x) := g(x) - h(x) \quad\text{s.t.}\;\langle a_i, x\rangle \leq b_i, \; i=1, \dots, p,4.

The key property is strong convexity descent: minxRn  ϕ(x):=g(x)h(x)s.t.  ai,xbi,  i=1,,p,\min_{x\in\mathbb{R}^n} \; \phi(x) := g(x) - h(x) \quad\text{s.t.}\;\langle a_i, x\rangle \leq b_i, \; i=1, \dots, p,5 monotonic decrease of the objective, and minxRn  ϕ(x):=g(x)h(x)s.t.  ai,xbi,  i=1,,p,\min_{x\in\mathbb{R}^n} \; \phi(x) := g(x) - h(x) \quad\text{s.t.}\;\langle a_i, x\rangle \leq b_i, \; i=1, \dots, p,6.

3. The Boosted DC Algorithm (BDCA): Algorithmic Acceleration

BDCA augments DCA’s surrogate minimization with extrapolation:

  • Compute minxRn  ϕ(x):=g(x)h(x)s.t.  ai,xbi,  i=1,,p,\min_{x\in\mathbb{R}^n} \; \phi(x) := g(x) - h(x) \quad\text{s.t.}\;\langle a_i, x\rangle \leq b_i, \; i=1, \dots, p,7,
  • If minxRn  ϕ(x):=g(x)h(x)s.t.  ai,xbi,  i=1,,p,\min_{x\in\mathbb{R}^n} \; \phi(x) := g(x) - h(x) \quad\text{s.t.}\;\langle a_i, x\rangle \leq b_i, \; i=1, \dots, p,8 (criticality), stop.
  • Check feasibility of extrapolation at minxRn  ϕ(x):=g(x)h(x)s.t.  ai,xbi,  i=1,,p,\min_{x\in\mathbb{R}^n} \; \phi(x) := g(x) - h(x) \quad\text{s.t.}\;\langle a_i, x\rangle \leq b_i, \; i=1, \dots, p,9 (active constraints).
  • Initialize g,h:RnR{+}g,h:\mathbb{R}^n \rightarrow \mathbb{R}\cup\{+\infty\}0, backtrack g,h:RnR{+}g,h:\mathbb{R}^n \rightarrow \mathbb{R}\cup\{+\infty\}1 (with g,h:RnR{+}g,h:\mathbb{R}^n \rightarrow \mathbb{R}\cup\{+\infty\}2) until

g,h:RnR{+}g,h:\mathbb{R}^n \rightarrow \mathbb{R}\cup\{+\infty\}3

for some g,h:RnR{+}g,h:\mathbb{R}^n \rightarrow \mathbb{R}\cup\{+\infty\}4. Else, g,h:RnR{+}g,h:\mathbb{R}^n \rightarrow \mathbb{R}\cup\{+\infty\}5.

  • Update g,h:RnR{+}g,h:\mathbb{R}^n \rightarrow \mathbb{R}\cup\{+\infty\}6.

BDCA pseudocode (see full details (Artacho et al., 2019)) guarantees: g,h:RnR{+}g,h:\mathbb{R}^n \rightarrow \mathbb{R}\cup\{+\infty\}7 and descent holds with finitely many backtracking steps.

4. Convergence Theory: Stationarity and Linear Rates

BDCA exhibits the following under (A1)-(A3) (strong convexity, subdifferentiability, Slater):

  • Every cluster point is a KKT point for g,h:RnR{+}g,h:\mathbb{R}^n \rightarrow \mathbb{R}\cup\{+\infty\}8.
  • The sequence g,h:RnR{+}g,h:\mathbb{R}^n \rightarrow \mathbb{R}\cup\{+\infty\}9 is nonincreasing and convergent.
  • gg0.
  • For quadratic objectives

gg1

one can split gg2, gg3 with gg4.

  • Under Slater, global gg5-linear convergence:

gg6

for some KKT point gg7.

5. Algorithmic Complexity, Practical Implementation, and Numerical Performance

Empirical tests compare DCA and BDCA on three classes:

  • Copositivity detection (gg8): BDCA is on average gg9 faster, with speedup growing in gg0.
  • gg1, gg2 trust-region subproblems: BDCA achieves gg3 (gg4) and gg5 (gg6) speedup.
  • Piecewise quadratic programs with box constraints: BDCA uniformly outperforms DCA; speedup improves with the number of pieces.

In all cases, BDCA yields greater per-iteration descent, and overhead for feasibility and objective evaluations is offset by significant improvements in convergence. BDCA always produces solutions with at least as small or better objective value compared to DCA. Boosting is activated in gg7 of iterations depending on the problem type.

Problem Class BDCA Speedup over DCA Fraction of Boosted Steps Scaling with Problem Size
Copositivity gg8 gg9 Increases with hh0
hh1 trust-region hh2 hh3 Stable across hh4
hh5 trust-region hh6 hh7 Stable across hh8
Piecewise Quadratic varies (increases with hh9) varies Grows with number of pieces

6. Theoretical Significance and Practical Guidelines

BDCA generalizes the classical DCA with a robust extrapolation mechanism, resulting in provably stronger descent steps, improved convergence rates, and practical scalability. The global KKT property, rigorous monotonicity, and linear rates in the quadratic and box-constrained cases match those found in best-in-class convex optimization algorithms (Artacho et al., 2019). Whenever the inner DCA subproblem is tractable (e.g. projection onto feasible set), it is highly recommended to adopt BDCA: line-search/extrapolation yields substantial acceleration with no loss in theoretical guarantees. The overhead of feasibility and function checking is negligible compared to gains from larger steps.

7. Extensions, Limitations, and Comparative Perspective

Though BDCA is formulated and proven for linearly constrained, strong-convexity DC programs with smooth surrogates, practical variants exist for nonsmooth and block-structured problems—see the referenced works for extensions. Limitations include reliance on strong convexity for global guarantees, and the need for efficiently computable projections or inner subproblems. In practice, BDCA’s boosting is automatically shut off at criticality, guaranteeing justification of acceleration only when empirically effective. Comparison with classical DCA in large-scale numerical tests validates the advantage of BDCA across a range of domains.

Key Reference:

"The Boosted DC Algorithm for linearly constrained DC programming" (Artacho et al., 2019)

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Difference of Convex functions Algorithm (DCA).