Papers
Topics
Authors
Recent
Search
2000 character limit reached

Graded Fast Hard Thresholding Pursuit (GFHTP₁)

Updated 17 January 2026
  • GFHTP₁ is an iterative greedy algorithm that recovers sparse signals by incrementally identifying the support with hard thresholding, without requiring prior sparsity knowledge.
  • It robustly handles sparse regression with LAD loss by using masked subgradient steps and quantile-based outlier detection to ensure global linear convergence under RIP conditions.
  • The algorithm demonstrates competitive empirical performance in high-dimensional settings, scaling efficiently and achieving fast recovery on synthetic datasets and MNIST image tasks.

Graded Fast Hard Thresholding Pursuit (GFHTP1_1) is an iterative greedy selection algorithm for finding sparse solutions to underdetermined or highly contaminated linear systems, particularly under sparsity and outlier-robustness constraints. GFHTP1_1 applies to both general convex sparsity-constrained objectives and specialized statistical estimation tasks, notably sparse regression with least absolute deviations (LAD) loss. Its defining characteristics are a "graded" support scheme that incrementally identifies the sparsity pattern and a streamlined update rule comprising hard thresholding operations, all without requiring prior knowledge of the underlying solution’s sparsity level or extensive user parameter tuning (Yuan et al., 2013, Xu et al., 10 Jan 2026).

1. Problem Setting and Algorithmic Structure

GFHTP1_1 addresses the general class of sparsity-constrained convex optimization problems

min⁡x f(x)subject to∥x∥0≤k,\min_x\ f(x)\quad\text{subject to}\quad \|x\|_0 \leq k,

where f:Rp→Rf:\mathbb{R}^p\to\mathbb{R} is a smooth convex function and kk is a prescribed sparsity (Yuan et al., 2013). In the context of robust signal recovery, it can be specialized to minimizing the ℓ1\ell_1 (LAD) loss under a sparsity constraint,

min⁡x∥b−Ax∥1subject to∥x∥0≤s,\min_x \|b - A x\|_1 \quad \text{subject to}\quad \|x\|_0 \leq s,

where A∈Rm×nA\in\mathbb{R}^{m\times n} is known, bb is the observation, 1_10 the sparse parameter of interest, and 1_11 the (unknown) true sparsity (Xu et al., 10 Jan 2026).

The GFHTP1_12 update consists of two core steps per iteration:

  • Gradient or subgradient step appropriate to the loss 1_13;
  • Hard thresholding to project onto the 1_14-sparse set, using the operator 1_15 which keeps the 1_16 entries of 1_17 with largest absolute value and sets the remainder to zero.

In the "graded" version for robust LAD recovery, the support size is incrementally grown without knowledge of 1_18; i.e., at each outer iteration 1_19, the support is of size 1_10, enabling automatic sparsity discovery (Xu et al., 10 Jan 2026).

2. Algorithmic Details and Implementation

GFHTP1_11 for Smooth Convex Objectives

For 1_12 smooth and convex, the update at iteration 1_13 is:

  1. 1_14
  2. 1_15
  3. Terminate when 1_16.

No debiasing or grading step is used; this yields the "fast" variant compared to GraHTP with debiasing (Yuan et al., 2013).

GFHTP1_17 for LAD with Outliers

For the LAD problem with high contamination, the update scheme involves:

  • Outer iteration: At step 1_18, compute quantile 1_19 of current residuals, identify "small" residuals, and mask large outliers.
  • Subgradient step: Use truncated subgradient (weighted by the indicator of small residuals).
  • Support update: Apply min⁡x f(x)subject to∥x∥0≤k,\min_x\ f(x)\quad\text{subject to}\quad \|x\|_0 \leq k,0 operator over min⁡x f(x)subject to∥x∥0≤k,\min_x\ f(x)\quad\text{subject to}\quad \|x\|_0 \leq k,1, where min⁡x f(x)subject to∥x∥0≤k,\min_x\ f(x)\quad\text{subject to}\quad \|x\|_0 \leq k,2 is the masked subgradient.
  • Inner loop: Perform min⁡x f(x)subject to∥x∥0≤k,\min_x\ f(x)\quad\text{subject to}\quad \|x\|_0 \leq k,3 restricted subgradient steps on the candidate support min⁡x f(x)subject to∥x∥0≤k,\min_x\ f(x)\quad\text{subject to}\quad \|x\|_0 \leq k,4.
  • Step-size: Scaled by median/truncated norm, parameterized by min⁡x f(x)subject to∥x∥0≤k,\min_x\ f(x)\quad\text{subject to}\quad \|x\|_0 \leq k,5 (fixed, e.g., 6) for stability (Xu et al., 10 Jan 2026).

The process continues until the masked LAD loss drops below a tolerance. The entire procedure is parameter-free except for a moderate-scale step-size constant min⁡x f(x)subject to∥x∥0≤k,\min_x\ f(x)\quad\text{subject to}\quad \|x\|_0 \leq k,6, with no need for manual specification of sparsity min⁡x f(x)subject to∥x∥0≤k,\min_x\ f(x)\quad\text{subject to}\quad \|x\|_0 \leq k,7.

3. Theoretical Guarantees

General Convex Case

GFHTPmin⁡x f(x)subject to∥x∥0≤k,\min_x\ f(x)\quad\text{subject to}\quad \|x\|_0 \leq k,8 achieves linear convergence under a graded restricted contraction (Condition Cmin⁡x f(x)subject to∥x∥0≤k,\min_x\ f(x)\quad\text{subject to}\quad \|x\|_0 \leq k,9). Specifically, if the step-size f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}0 and f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}1, then

f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}2

yielding geometric convergence and exact finite-step recovery when f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}3 is an exact minimizer (Yuan et al., 2013).

Robust LAD Setting

Assuming the design matrix f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}4 satisfies the restricted 1-isometry property (RIPf:Rp→Rf:\mathbb{R}^p\to\mathbb{R}5) of appropriate order and contamination fraction f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}6 is bounded below f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}7, GFHTPf:Rp→Rf:\mathbb{R}^p\to\mathbb{R}8 shows:

  • Global linear convergence: For any f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}9-sparse signal kk0,

kk1

after kk2 outer steps with suitable kk3 and quantile parameters.

  • Exact support/signal recovery for flat signals: For sufficiently "flat" kk4 and Gaussian kk5, exact support identification and reconstruction occur after kk6 outer steps with high probability (Xu et al., 10 Jan 2026).

The proof strategy leverages contraction from inner subgradient updates and a support-matching induction relying on quantile concentration and RIPkk7 structural control.

4. Computational Complexity and Practical Considerations

The per-iteration complexity of GFHTPkk8 is as follows:

Operation Complexity (per iter) Context
Gradient or subgradient computation kk9 LAD and general convex loss
Quantile computation (LAD) ℓ1\ell_10 Needed to mask outliers (LAD only)
Hard thresholding ℓ1\ell_11 ℓ1\ell_12 or ℓ1\ell_13 Partial sort or selection
Inner loop (LAD, support size ℓ1\ell_14) ℓ1\ell_15 Per inner subgradient step

For most settings, with ℓ1\ell_16 inner steps and ℓ1\ell_17, total complexity is ℓ1\ell_18 per run (Xu et al., 10 Jan 2026). Compared to PSGD and AIHT, GFHTPℓ1\ell_19 is often faster when the true sparsity min⁡x∥b−Ax∥1subject to∥x∥0≤s,\min_x \|b - A x\|_1 \quad \text{subject to}\quad \|x\|_0 \leq s,0 is unknown or the data is heavily contaminated.

Parameter selection is minimized: only the step-size scaling min⁡x∥b−Ax∥1subject to∥x∥0≤s,\min_x \|b - A x\|_1 \quad \text{subject to}\quad \|x\|_0 \leq s,1 is user-set (in a moderate theoretical range, e.g., min⁡x∥b−Ax∥1subject to∥x∥0≤s,\min_x \|b - A x\|_1 \quad \text{subject to}\quad \|x\|_0 \leq s,2 suffices in experiments), and the quantile parameter min⁡x∥b−Ax∥1subject to∥x∥0≤s,\min_x \|b - A x\|_1 \quad \text{subject to}\quad \|x\|_0 \leq s,3 is typically set near 0.5. The stopping criterion is typically based on the relative iterate change or capped at a maximum number of iterations.

5. Empirical Performance and Applications

GFHTPmin⁡x∥b−Ax∥1subject to∥x∥0≤s,\min_x \|b - A x\|_1 \quad \text{subject to}\quad \|x\|_0 \leq s,4 shows strong empirical performance on both synthetic and real data:

  • Synthetic sparse signal recovery: In experiments with min⁡x∥b−Ax∥1subject to∥x∥0≤s,\min_x \|b - A x\|_1 \quad \text{subject to}\quad \|x\|_0 \leq s,5, min⁡x∥b−Ax∥1subject to∥x∥0≤s,\min_x \|b - A x\|_1 \quad \text{subject to}\quad \|x\|_0 \leq s,6, sparsity min⁡x∥b−Ax∥1subject to∥x∥0≤s,\min_x \|b - A x\|_1 \quad \text{subject to}\quad \|x\|_0 \leq s,7–min⁡x∥b−Ax∥1subject to∥x∥0≤s,\min_x \|b - A x\|_1 \quad \text{subject to}\quad \|x\|_0 \leq s,8, and outlier fractions min⁡x∥b−Ax∥1subject to∥x∥0≤s,\min_x \|b - A x\|_1 \quad \text{subject to}\quad \|x\|_0 \leq s,9–A∈Rm×nA\in\mathbb{R}^{m\times n}0, GFHTPA∈Rm×nA\in\mathbb{R}^{m\times n}1 achieves near-perfect support recovery (success rate A∈Rm×nA\in\mathbb{R}^{m\times n}2) and low relative error, with CPU time competitive with or superior to other greedy algorithms (AIHT, PSGD).
  • MNIST image recovery: For image vectors observed through random Gaussian projections and A∈Rm×nA\in\mathbb{R}^{m\times n}3 outlier corruption, GFHTPA∈Rm×nA\in\mathbb{R}^{m\times n}4 reaches SNR A∈Rm×nA\in\mathbb{R}^{m\times n}580 dB in A∈Rm×nA\in\mathbb{R}^{m\times n}69 ms, compared to PSGD's SNR of A∈Rm×nA\in\mathbb{R}^{m\times n}75 dB in over 1 s.

These results underline the robustness, computational efficiency, and adaptivity of GFHTPA∈Rm×nA\in\mathbb{R}^{m\times n}8 in outlier-prone and high-dimensional settings (Xu et al., 10 Jan 2026, Yuan et al., 2013).

6. Advantages, Limitations, and Extensions

GFHTPA∈Rm×nA\in\mathbb{R}^{m\times n}9 provides the following core advantages:

  • No a priori knowledge of sparsity bb0 needed; automatic graded support growth.
  • Nearly parameter-free, requiring only a single moderate step-size scale.
  • Provable linear convergence under RIPbb1-like structural assumptions.
  • Robustness to large-magnitude outliers via adaptive masked subgradients.
  • Efficient for high-dimensional or contaminated regimes.

Limitations include the bb2 overhead of quantile computation for very large datasets, a restriction to random Gaussian-type bb3 for current theory, and possibly conservative theoretical step-size ranges.

Potential extensions outlined include:

  • Adaptive or line-searched step-size schemes.
  • Momentum or Nesterov acceleration in the inner optimization loops.
  • Generalizations to nonlinear or low-rank matrix recovery settings.
  • Distributed and online adaptations for streaming scenarios.
  • Data-driven automation for the quantile truncation parameter.

A plausible implication is the broader applicability of the "graded" hard thresholding paradigm to a range of nonconvex, high-dimensional, and robust estimation problems beyond bb4 loss. However, extensions to structured or compressive measurement matrices and rigorous theoretical guarantees for such cases remain prominent avenues for investigation (Xu et al., 10 Jan 2026).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (2)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Graded Fast Hard Thresholding Pursuit (GFHTP$_1$).