---
title: Greedy Coordinate Descent
url: https://www.emergentmind.com/topics/greedy-coordinate-descent
type: topic
---

# Greedy Coordinate Descent

Greedy Coordinate Descent (GCD) is a variant of coordinate descent algorithms where the coordinate to update at each iteration is chosen using a greedy rule, typically selecting the coordinate promising the largest decrease in a surrogate or true objective. In high-dimensional optimization—spanning convex, nonconvex, and composite problems—GCD offers attractive rates, practical speedups, and flexible algorithmic paradigms leveraging problem structure. The method generalizes both to block variants and hybrid strategies, and is closely related to the classical Gauss–Southwell rule. It forms the backbone of numerous state-of-the-art solvers for $\ell_1$-regularized learning, quadratic programming, discrete optimization, quantization of neural networks, and large-scale empirical risk minimization.

## 1. Problem Formulation and Greedy Selection Rules

The most general setting considers composite objectives of the form
\[
\min_{\alpha \in \mathbb{R}^n} F(\alpha) := f(\alpha) + \sum_{i=1}^n g_i(\alpha_i),
\]
where $f$ is convex and coordinatewise $L$-smooth, each $g_i$ is convex (or possibly nonconvex and separable), and typical choices include $\ell_1$ penalties and box constraints [1810.06999].

At each iteration, GCD evaluates, for each coordinate $i$, a potential reduction using a local surrogate objective. In the smooth convex case, the canonical Gauss–Southwell rule picks
\[
i_t = \arg\max_i |\nabla_i f(\alpha)|.
\]
For nonsmooth or composite problems, the optimal one-dimensional decrease is used:
\[
s_i(\alpha) := \min_{s \in \partial g_i(\alpha_i)} [\nabla_i f(\alpha) + s],\quad i_t = \arg\max_i s_i(\alpha).
\]
For quadratic problems with nonnegativity or box constraints, the greedy rule computes for each $i$:
\[
\Delta_i = \min_{t \in X_i} f(\ldots, t, \ldots) - f(\ldots, x_i, \ldots),
\]
and selects $i_t = \arg\min_i \Delta_i$ [2012.05943].

In block and hybrid schemes, a working set $B\subset\{1,\dots,n\}$ of size $k$ is chosen greedily by maximizing k-dimensional decrease proxies and the block subproblem is solved exactly or approximately [1706.06493, 1712.08859, 1206.3238].

## 2. Theoretical Guarantees and Convergence Rates

### Strongly Convex Case

For composite $F$ that is strongly convex with respect to $\|\cdot\|_1$ with modulus $\mu_1>0$, the Gauss–Southwell–proximal (greedy) rule yields linear convergence independent of ambient dimension:
\[
F(\alpha^{(t)})-F(\alpha^\star) \leq \left(1 - \frac{\Theta^2 \mu_1}{L}\right)^{t/2} [F(\alpha^{(0)})-F(\alpha^\star)],
\]
for exact or $\Theta$-approximate selection, $0<\Theta\le 1$ [1810.06999]. The proof stratifies steps as "good" (proximal updates do not cross kinks of nonsmooth $g_i$) and "bad;" at least half are good, and each good step contracts the residual objective by a fixed fraction. No explicit $n$ dependence appears in the rate.

### General Convex and Non-Strongly Convex Case

Without strong convexity, greedy CD achieves sublinear decay:
\[
F(\alpha^{(t)}) - F(\alpha^\star) = O\left(\frac{L D^2}{\Theta^2 t}\right),
\]
with $D$ the $\ell_1$ diameter of the level set [1810.06999]. For pure smooth objectives, the rate is $O(1/k)$ in the number of iterations, matching classical coordinate descent but with smaller constants due to greedy selection [1512.05808, 1806.02476].

### Block, Hybrid, and Constrained Extensions

In block-greedy and hybrid-combinatorial schemes, convergence to block-k stationary points is guaranteed. Such stationary points are strictly stronger than those achievable by coordinatewise optimality, yielding fewer spurious local minima in nonconvex settings [1706.06493]. For equality- and box-constrained optimization, greedy two-coordinate schemes achieve linear convergence under a (proximal) Polyak–Łojasiewicz (PL) condition, with rates again independent of $n$ [2307.01169]. For nonnegativity-constrained QPs, GCD achieves global convergence, and with positive definite Hessians, the rate is $O\left(\left(1-\frac{\mu}{n L_{\rm avg}}\right)^k\right)$ [2012.05943].

## 3. Efficient Implementation and Algorithmic Advances

### Maximum Inner Product Search (MIPS)

For problems admitting the structure $f(\alpha)=l(A\alpha)+c^\top\alpha$ and $g_i(\alpha_i)=\lambda|\alpha_i|$, the greedy coordinate update can be rephrased as a maximum inner product search (MIPS):
\[
i_{\text{GS}} = \arg\max_{\tilde{A} \in B} \langle \tilde{A}, q\rangle,
\]
where $B$ is a dynamically maintained active set of lifted feature vectors [1810.06999]. Modern nearest neighbor algorithms (e.g., LSH, HNSW) can implement these queries in sublinear time per iteration, reducing wall-clock cost close to that of a single coordinate-gradient calculation.

### Block/Parallel Greedy Schemes

Block-greedy and thread-greedy coordinate descent select the best coordinate per block or per thread, with each update executed in parallel [1212.4174, 1206.6409]. For block sizes fitted to hardware architecture and blocks formed via clustering for low cross-block correlation, empirical speedups are substantial, especially when $\ell_1$-regularized solutions are denser [1212.4174].

### Accelerated and Stochastic Greedy Methods

Recent advances include semi-greedy and fully greedy accelerated coordinate descent (ASCD, AGCD), which combine Nesterov-style momentum with Gauss–Southwell selection. ASCD achieves $O(1/k^2)$ convergence and accelerated linear convergence under strong convexity, while AGCD typically performs even better in practice, though theoretical guarantees require additional mild conditions [1806.02476]. Mini-batch, block, and stochastic strategies further reduce per-iteration cost or hardware burden [1702.07842].

### Large-Scale and Application-Specific Innovations

Greedy CD has been specialized to large-scale Gaussian process regression through block selection solving zero-norm constrained subproblems greedily [1206.3238], LLM quantization via a tailored greedy descent over discrete codebooks [2406.17542], and large least squares by double-greedy subspace selection and orthogonalization [2203.02153]. For distributed settings (e.g., feature-wise parallelism in Hadoop clusters), greedy block selection yields faster convergence in both cycle count and wall-clock time, drastically reducing expensive inter-node communication [1405.4544].

## 4. Practical Performance and Applications

Greedy coordinate descent's practical impact is evidenced in several settings:

- **Sparse regression ($\ell_1$-regularized least squares, Lasso):** Greedy CD methods, including soft-thresholding coordinate updates and hybrid Ray-Refinement strategies, dominate traditional CD in number of sweeps especially for low-$\lambda$ regimes [1512.05808].
- **Large-scale linear SVMs:** Dual coordinate methods with greedy selection attain faster sparsity and reduced computation [1810.06999].
- **Nonnegative matrix factorization and NQP:** On pure NQP, GCD is orders of magnitude faster than CCD, RCD, and accelerated gradient methods, and achieves rapid convergence in matrix factorization quality [2012.05943].
- **Deep model quantization:** Greedy coordinate selection at the quantization-code level (as in CDQuant) consistently drives lower layerwise reconstruction error and outperforms cyclic/CD variants such as GPTQ, while scaling to $10^{11}$-parameter models [2406.17542].
- **High-dimensional empirical risk minimization (DP/ERM):** In private optimization settings, GCD leverages structural sparsity to reduce the penalty incurred by differentially private selection, yielding a utility bound logarithmic rather than polynomial in dimension [2207.01560].
- **Inverse problems and system-solving:** Greedy versions of Gauss–Seidel and block Kaczmarz schemes exhibit provably faster descent and wall-clock efficiency for linear and quadratic systems, especially when variable selection is judiciously adapted to problem structure [2004.03692, 1404.6635, 2203.02153].

## 5. Limitations, Variants, and Open Directions

Key limitations and active directions include:

- **Selection cost:** Greedy selection requires full or blockwise evaluation of decrease proxies, inducing at least $O(n)$ (or $O(n^2)$ for blocks) per iteration unless structure/MIPS tricks are available [1810.06999, 1206.3238, 1712.08859].
- **Scalability to very high dimensions:** Parallel/thread-greedy and block-greedy schemes ameliorate per-iteration overhead but may entail tradeoffs in load balancing and atomic update costs, especially for highly sparse or clustered features [1212.4174, 1206.6409].
- **Nonconvex and discrete optimization:** While greedy CD methods drive toward block-k stationary points, global optima are not guaranteed without exhaustive combinatorial search. The combinatorial subproblem dimension is a practical bottleneck (usually $k\leq 20$ is feasible) [1706.06493].
- **Worst-case bounds:** For some accelerated or block-greedy variants, worst-case theoretical rates are known only under additional technical conditions or for specific classes of problems (e.g., strong or quadratic growth), though empirical speedups are robust [1806.02476, 1712.08859].
- **Hybridization and adaptivity:** Combining greedy updates with stochastic, block, or cyclic schemes (e.g., switching to UCD at late stages) balances early rapid decrease with late-stage efficiency; hybrid rules remain an area of active algorithmic development [1810.06999, 1712.08859].

## 6. Summary Table: Representative Variants and Guarantees

| Variant / Application            | Greedy Rule Type   | Theoretical Rate                    | Specialized Implementation                   | Reference      |
|----------------------------------|--------------------|-------------------------------------|----------------------------------------------|---------------|
| Composite convex (sparse SVM/L1) | Gauss–Southwell-prox | Linear ($n$-independent, strong conv.) | MIPS search, sublinear iteration cost        | [1810.06999]  |
| Non-neg. quadratic programming   | Optimal decrease per coordinate | $O(1-\mu/(nL_{avg}))$ linear     | Gradient maintenance for $O(n)$ update       | [2012.05943]  |
| Hybrid discrete/sparse opt.      | Block-k greedy      | Block-k stationary, linear (binary)  | Exhaustive block search ($k$ small)          | [1706.06493]  |
| Distributed block CD (L1-class.) | Surrogate decrease  | Q-linear under strong convexity      | Local greedy in block, AllReduce, line search| [1405.4544]   |
| Quantized LLMs (CDQuant)         | Greedy discrete coordinate | Finite, monotonic, local opt.   | Full/Block search, per-row Hessian caching   | [2406.17542]  |
| Accelerated GCD (ASCD/AGCD)      | Greedy (with momentum) | $O(1/k^2)$ (semi-greedy), heuristic  | Hybrid random/greedy update policy           | [1806.02476]  |
| Gaussian process regression      | Block greedy via obj. decrease | Linear, global opt.           | Progressive block building, kernel subsampling| [1206.3238]   |
| 2-coordinate equality-constr.    | Max. gradient gap   | Linear, $n$-independent (PL)         | Sorting-based selection, block steepest descent| [2307.01169] |

## 7. Research Impact and Future Prospects

The greedy coordinate descent paradigm forms a unifying mechanism underlying many contemporary large-scale optimization methods, allowing practitioners to exploit problem structure, data sparsity, and low-dimensional active sets for accelerated convergence. The framework's flexibility supports parallel/distributed architectures, hybridized block/coordinate schemes, and integration with second-order updates (block cubic Newton), positioning GCD as a foundational, continually evolving family of algorithms for modern data-intensive optimization [1810.06999, 1712.08859, 2407.18150]. Continued efforts in reducing selection overhead, integrating global optima seeking (as in combinatorial hybrids), and exploiting adaptive block formation, as well as theoretical refinement for nonconvex and constrained settings, are central areas of future investigation.

Source: https://www.emergentmind.com/topics/greedy-coordinate-descent