---
title: 'DCA: Difference of Convex Functions Algorithm'
url: https://www.emergentmind.com/topics/difference-of-convex-functions-algorithm-dca-dca1fcfa-7051-4b04-8ac7-89ba012e6f1f
type: topic
---

# DCA: Difference of Convex Functions Algorithm

The Difference of Convex Functions Algorithm (DCA) is a foundational framework for the local minimization of functions expressible as the difference of two convex (DC) functions. DCA is particularly central for structured nonconvex optimization in applications requiring precise criticality guarantees, efficient per-iteration majorization, and provable global convergence. The method operates by iteratively constructing and minimizing convex surrogates, leveraging convex analytical properties of the constituent terms. In smooth and strongly convex regimes, DCA admits rigorous descent analysis, sharp convergence rates under various geometric conditions, and serves as a basis for a spectrum of acceleration schemes with practical and theoretical impact [1507.07375].

## 1. Mathematical Formulation and Classical DCA Iteration

Consider the unconstrained minimization of a DC function:
\[
\min_{x \in \mathbb{R}^m} \ \varphi(x) = g(x) - h(x)
\]
where $g, h : \mathbb{R}^m \to \mathbb{R}$ are closed, convex, and at least $C^2$ (in the smooth setting). If $\varphi$ is bounded below, one may assume $g$ and $h$ are strongly convex—this is always achievable by adding and subtracting a quadratic term. At iteration $k$, the classical DCA forms the affine majorant of $-h$ at $x^k$:
\[
-h(x) \leq -h(x^k) - \langle \nabla h(x^k), x-x^k \rangle
\]
and solves
\[
y_k := \arg \min_{x} \left\{g(x) - \langle \nabla h(x^k), x \rangle \right\}
\]
followed by the update $x^{k+1} := y_k$. The DCA step direction is $d_k := y_k - x^k$. The descent property is explicit: one has
\[
\varphi(y_k) \leq \varphi(x^k) - \rho \|y_k - x^k\|^2
\]
for $\rho > 0$ denoting the strong convexity modulus [1507.07375].

## 2. Accelerated Variants: Boosted DC Algorithms

The DCA step $d_k$ is a strict descent direction for $\varphi$ evaluated at $y_k$:
\[
\langle \nabla \varphi(y_k), d_k \rangle \leq -\rho \|d_k\|^2 < 0
\]
This motivates "Boosted DCA" (BDCA) accelerations that perform a line search along $d_k$ to maximize decrease:
- **BDCA-Backtracking** initializes $\lambda = \bar{\lambda}$ and backtracks until
\[
\varphi(y_k + \lambda d_k) \leq \varphi(y_k) - \alpha \lambda \|d_k\|^2
\]
- **BDCA-Quadratic** fits a quadratic model $\varphi_k(\lambda)$ through $(0, \varphi(y_k))$, the directional derivative at $0$, and $\bar{\lambda}$, then minimizes this interpolation to select $\lambda$.

Both variants then update $x^{k+1} = y_k + \lambda d_k$. These BDCA schemes inherit the global convergence properties of DCA but empirically and theoretically exhibit substantially faster rates [1507.07375].

## 3. Convergence Guarantees and Łojasiewicz Analysis

Let $\varphi$ possess the Łojasiewicz property at every cluster point $x^*$ (automatic for real-analytic $\varphi$):
\[
|\varphi(x) - \varphi(x^*)|^\theta \leq M \|\nabla \varphi(x)\|
\]
for some $\theta \in [0,1)$ locally near $x^*$. Under mild regularity (local Lipschitzness, boundedness below), BDCA generates $\{\varphi(x_k)\}$ converging monotonically to $\varphi^*$. Moreover:
- $x_k \to x^*$, a stationary point: $\nabla \varphi(x^*) = 0$.
- The convergence rate is determined by $\theta$:
  - $\theta = 0$: finite step convergence;
  - $0 < \theta \leq \frac{1}{2}$: linear convergence;
  - $\frac{1}{2} < \theta < 1$: sublinear, with explicit polynomial rates:
    \[
    \|x_k - x^*\| = O\left(k^{-\frac{1-\theta}{2\theta-1}}\right),\quad \varphi(x_k) - \varphi^* = O\left(k^{-\frac{1}{2\theta-1}}\right)
    \]
Proof techniques

Source: https://www.emergentmind.com/topics/difference-of-convex-functions-algorithm-dca-dca1fcfa-7051-4b04-8ac7-89ba012e6f1f