---
title: 'Greedy-Threshold Methods: Theory & Applications'
url: https://www.emergentmind.com/topics/greedy-threshold
type: topic
---

# Greedy-Threshold Methods: Theory & Applications

Greedy-threshold denotes a family of constructions in which a greedy procedure is governed by a threshold condition rather than by an exact maximizer. In the classical Thresholding Greedy Algorithm (TGA), the threshold is implicit in the requirement that the selected coefficients are the largest in modulus; in weak and generalized variants, admissibility is relaxed by parameters such as $\tau$, $\lambda$, or a comparison function $f(m)$. In algorithmic settings outside Banach-space approximation, the same pattern appears when one accepts any atom, row, task, or candidate whose score exceeds a prescribed, adaptive, or capped threshold, and in some analyses the threshold instead marks a regime boundary below which a greedy policy is exactly optimal and above which it is not. This suggests a unifying view of greedy-threshold methods as threshold-controlled relaxations, generalizations, or phase-transition analyses of greediness [2201.12029], [2302.05758], [1307.1949], [1604.05993], [1909.07895].

## 1. Foundational definitions

In the approximation-theoretic literature, the basic object is a basis $\mathcal B=(e_n)_{n=1}^\infty$ with biorthogonal functionals $(e_n^*)$. For
\[
x=\sum_{n=1}^\infty e_n^*(x)\,e_n,
\]
a finite set $A\subset\mathbb N$ of cardinality $m$ is called a greedy set when the selected coefficients dominate the unselected ones in modulus. One common formulation is
\[
\min_{n\in A}|e_n^*(x)|>\max_{n\notin A}|e_n^*(x)|,
\]
and the associated greedy operator is
\[
G_m(x)=\sum_{n\in A}e_n^*(x)\,e_n.
\]
Equivalent expositions use the non-strict relation $\ge$ together with a tie-breaking rule by natural ordering. The TGA is then the sequence $(G_m)_{m\ge 1}$ obtained by repeatedly selecting the largest coefficients in modulus [2205.00268], [2201.12029].

A first generalization introduces an explicit threshold parameter $\tau\in(0,1]$. The $\tau$-Thresholding Greedy Algorithm chooses any set $A$ of size $m$ satisfying
\[
\min_{n\in A}|a_n|\ge \tau \max_{k\notin A}|a_k|,
\qquad a_n=e_n^*(x),
\]
and forms
\[
G_m^{(\tau)}(x)=\sum_{n\in A} e_n^*(x)\,e_n.
\]
When $\tau=1$, this is the classical TGA. A second generalization enlarges the greedy sum itself: instead of comparing the error after $\lfloor \lambda m\rfloor$ greedy terms with best $m$-term benchmarks, one defines
\[
Y_{[\lambda m]}(x)=\sup_{A\in G(x,\lfloor\lambda m\rfloor)}\|x-P_A(x)\|,
\qquad \lambda\ge 1,
\]
leading to $\lambda$-almost greedy and $\lambda$-partially greedy bases [2302.05758], [2205.00268].

A further axis of generalization replaces the benchmark size $m$ by a function $f(m)$. For $f$ in the class $\mathcal F$ of continuous, increasing, concave maps with $f(0)=0$, $f(1)\le 1$, and $f'>0$ on $(0,\infty)$, a basis is called $f$-greedy if
\[
\|x-P_A(x)\|\le C\,\sigma_{f(m)}(x),
\]
and $f$-almost greedy if
\[
\|x-P_A(x)\|\le C\,\tilde\sigma_{f(m)}(x),
\]
for all $x$, $m$, and greedy sets $A$ of size $m$ [2209.09628]. These variants make explicit that “threshold” may act on coefficient admissibility, on greedy-sum size, or on the benchmark against which greedy error is measured.

## 2. Structural theory in Banach and quasi-Banach spaces

The classical characterization of greedy bases is the Konyagin–Temlyakov theorem: a basis is greedy if and only if it is unconditional and democratic. The almost-greedy property replaces the best linear $m$-term error $\sigma_m(x)$ by the best projection error
\[
\tilde\sigma_m(x)=\inf\{\|x-P_A(x)\|:|A|=m\},
\]
and almost-greedy bases are exactly those that are quasi-greedy and democratic. In parallel, quasi-greediness is equivalent to convergence of greedy approximants and to uniform boundedness of greedy operators [2201.12029], [1903.11651].

Threshold enlargement does not behave uniformly across all greedy-type classes. For almost-greedy bases, the enlargement parameter $\lambda$ does not produce a genuinely new class: a basis is almost greedy if and only if it is $\lambda$-almost greedy for some, equivalently all, $\lambda>1$. By contrast, for each $\lambda>1$ there exists an unconditional basis that is $\lambda$-partially greedy but is not $1$-partially greedy. The characterization of $\lambda$-partially greedy bases is
\[
\text{$\lambda$-partially greedy}
\iff
\text{quasi-greedy}+\text{$\lambda$-max-conservative},
\]
where $\lambda$-max-conservativity is expressed by
\[
\|1_A\|\le C\|1_B\|
\]
for sets $A<B$ with
\[
\#B-\#A<\max A,\qquad \#B=\#A+\lfloor(\lambda-1)\max A\rfloor.
\]
This distinction corrects a common oversimplification: allowing larger greedy sums preserves almost-greediness but strictly weakens partial greed when $\lambda>1$ [2205.00268].

The $f$-theory yields a different phenomenon. For every non-identity $f\in\mathcal F$,
\[
f\text{-greedy}\Longleftrightarrow f\text{-almost greedy}.
\]
The paper also characterizes $f$-almost greedy as quasi-greedy plus Property $(A,f)$, and $f$-greedy as $f$-unconditional plus Property $(A,f)$. This collapses the distinction between best linear and best threshold approximants once the comparison size is strictly sublinear. A plausible implication is that sublinear benchmark growth regularizes the gap between linear and coordinate projection models of approximation [2209.09628].

Several later refinements restrict the admissible comparison supports. Consecutive greedy bases compare greedy error only with approximants supported on intervals of $\mathbb N$, while CGPCC restricts further to interval-supported polynomials with constant coefficients. On Schauder bases, consecutive-greedy is equivalent to almost-greedy. The sequential theory replaces intervals by families $F_{(a_n)}$ generated by a fixed gap sequence; the $F_{(a_n)}$-almost greedy property and the $F_{(a_n)}$-strong partially greedy property are equivalent to their classical counterparts if and only if $(a_n)$ is bounded. Under additional conditions, the $F_{(a_n)}$-minimum partially greedy property lies strictly between almost-greedy and strong partially greedy [2302.05758], [2310.16947].

Quantitative control is given by Lebesgue-type constants and related parameters. The Lebesgue constants
\[
L_m[x,X]=\sup_{f\ne 0}\frac{\|f-G_m(f)\|}{\sigma_m(f)}
\]
satisfy
\[
L_m[x,X]\approx \max\{k_m[x,X],A_m[x,X]\},
\]
where $k_m$ are unconditionality parameters and $A_m$ are the squeeze-symmetry constants. This answers Temlyakov’s 2011 question about a natural greedy-type parameter complementing $(k_m)$ in an exact linear rule for $(L_m)$. Rate-of-convergence results then bound the TGA error by a product of $\|f\|_X$ and $|f|_{A_1(D)}$, and in $L_p$ the trigonometric system and the Haar basis admit sharp rates [2104.10912], [2304.09586]. Related work also characterizes greedy bases by approximation with polynomials of constant coefficients, including the RGPCC and URGPCC formulations, and extends these equivalences to quasi-Banach spaces [2309.14607].

## 3. Sparse recovery and thresholded pursuit

In compressive sensing, thresholding often replaces the exact “largest inner product” step of a greedy pursuit. Orthogonal Matching Pursuit with Thresholding (OMPT) takes measurements $f\in\mathbb R^n$, a normalized dictionary $\Phi=[\phi_1\cdots \phi_d]$, and a threshold parameter $t\in(0,1)$. At iteration $k$, instead of selecting the atom of maximum correlation with the residual $r^k$, OMPT chooses any index $i\notin\Lambda^k$ satisfying
\[
|\langle r^k,\phi_i\rangle|\ge t\|r^k\|_2.
\]
The support is updated, a least-squares projection is performed on the selected support, and the residual is recomputed. The analysis uses the global $2$-coherence
\[
\nu_k(\Phi):=
\max_{i\in[d]}\max_{\Lambda\subset[d]\setminus\{i\},\,|\Lambda|\le k}
\Bigl(\sum_{j\in\Lambda}\langle \phi_i,\phi_j\rangle^2\Bigr)^{1/2},
\]
which bridges mutual coherence and the restricted isometry constant via
\[
M=\nu_1\le \nu_k\le \delta_{k+1}\le \sqrt{k}\,\nu_k\le kM.
\]
Under
\[
\delta_k+\sqrt{k}\,\nu_k<1
\]
and
\[
\frac{\nu_k}{\sqrt{1-\delta_k}}<t\le \frac{\sqrt{1-\delta_k}}{\sqrt{k}},
\]
OMPT recovers a $k$-sparse signal exactly in $k$ steps, and the paper states that these conditions coincide with the best known bounds for standard OMP while reducing the cost of the selection step [1307.1949].

Thresholding Greedy Pursuit (TGP) applies the same principle to a CoSaMP-style framework. At each iteration it computes the proxy
\[
y=A^*r,\qquad y\leftarrow y/\|r\|_2,
\]
keeps all coordinates above a scalar threshold,
\[
T=\{i:|y_i|>\tau\},
\]
merges them into the support estimate, solves a least-squares problem on the merged support, and updates the residual. The theory emphasizes asymptotic sparse recovery with additive noise. With
\[
\tau\ge c_0\frac{\sqrt{\log N}}{\sqrt N},
\qquad
c_0=\sqrt{2(\gamma+\kappa)},
\]
the “No Phantom Signal” theorem yields $\Omega=\emptyset$ when $x=0$ with probability at least $1-2/N^\kappa$, and the “No False Discoveries” theorem gives
\[
\Omega\subseteq \operatorname{supp}(x)
\]
under $M\le 1/(4\mu)$. In the exact-recovery regime, if
\[
M\le \min\Bigl\{\frac1{4\mu},\frac{\sqrt N}{4c_0\sqrt{\log N}}\Bigr\},
\qquad
\|e\|_2\le 0.03\,\sqrt M\,\min_{i\in\operatorname{supp}(x)}|x_i|,
\]
and
\[
\tau=\sqrt{\frac43\Bigl(\frac\mu4+\frac{c_0^2\log N}{N}\Bigr)},
\]
then TGP recovers exactly
\[
\Omega=\operatorname{supp}(x)
\]
with probability $1-2/N^\kappa$. The paper stresses that the threshold can be chosen without the knowledge of sparsity level of the signal and strength of the noise [2103.11893].

Taken together, OMPT and TGP show that thresholding is not merely a heuristic weakening of greedy pursuit. In both cases, the thresholded rule replaces a more expensive exact maximization while retaining recovery guarantees under coherence- or concentration-based assumptions [1307.1949], [2103.11893].

## 4. Learning and nonlinear inverse problems

Orthogonal Greedy Learning (OGL) provides a learning-theoretic version of the same idea. Classical OGL selects
\[
g_k=\arg\max_{g\in\mathcal D_n}|\langle r_{k-1},g\rangle_m|,
\qquad
r_{k-1}=y-f_{\mathbf z}^{k-1},
\]
and updates the predictor by orthogonal projection onto the span of selected atoms. The $\delta$-greedy-threshold metric declares an atom $g$ to be active when
\[
\frac{|\langle r_{k-1},g\rangle_m|}{\|r_{k-1}\|_m}>\delta,
\]
or, in the closely related formulation,
\[
\frac{|\langle r_{k-1},g\rangle_m|}{\|r_{k-1}\|_m}\ge \delta,
\qquad 0<\delta\le \tfrac12.
\]
The $\delta$-Threshold OGL algorithm chooses any active atom, projects orthogonally onto the enlarged span, and stops when no active atom remains or when
\[
\|r_{k+1}\|_m\le \delta \|y\|_m.
\]
Under the source condition
\[
f_\rho\in \mathcal L_{1,\mathcal D_n}^r,
\]
the output $f_{\mathbf z}^\delta$ satisfies, with probability at least $1-t$,
\[
\mathbb E\bigl[\|f_{\mathbf z}^\delta-f_\rho\|_\rho^2\bigr]
\le
C\mathcal B^2\Bigl\{
m^{-1}\delta^{-2}\log m\log\tfrac1\delta\log\tfrac2t
+
\delta^2
+
n^{-2r}
\Bigr\},
\]
and choosing
\[
\delta\asymp m^{-1/4},\qquad n\gtrsim m^{1/(4r)}
\]
yields the rate $\mathcal O(m^{-1/2}(\log m)^2)$. The papers emphasize that steepest gradient descent is not the unique greedy criterion of OGL and that the threshold parameter acts simultaneously as a greediness control and an adaptive stopping rule [1604.05993], [1411.3553].

A related pattern appears in nonlinear Kaczmarz methods. The greedy capped nonlinear Kaczmarz framework introduces a capped set of promising rows at each iterate. In the distance–residual capped method, one defines
\[
\varepsilon_k
=
\frac12\Bigl[
\frac{1}{\|f(x_k)\|_2^2}\max_{i\in[m]} |f_i(x_k)|^2\|\nabla f_i(x_k)\|_2^2
+
\frac1{\|f'(x_k)\|_F^2}
\Bigr]
\]
and the capped set
\[
U_k=\Bigl\{
i\in[m]:
|f_i(x_k)|^2\|\nabla f_i(x_k)\|_2^2
\ge
\varepsilon_k\|f(x_k)\|_2^2\|\nabla f_i(x_k)\|_2^2
\Bigr\}.
\]
In the residual–distance capped method, one defines
\[
\delta_k=\frac12\Bigl[
\max_{i\in[m]}\frac{|f_i(x_k)|^2}{\|f(x_k)\|_2^2}
+
\frac1m
\Bigr],
\qquad
I_k=\{i\in[m]: |f_i(x_k)|^2\ge \delta_k\|f(x_k)\|_2^2\}.
\]
Randomization is then performed only within $U_k$ or $I_k$, with probabilities determined by residual- or distance-based scores. Under the local tangential cone condition, the resulting single-row and block methods satisfy expected contraction bounds that the paper states are strictly smaller than the NRK factor. This is an explicit exploration–exploitation design: discard low-score rows by a threshold, then sample within the retained set [2210.00653].

## 5. Threshold-greedy optimization and decentralized control

For monotone submodular task allocation under a partition matroid, the Decreasing Threshold Task Allocation (DTTA) algorithm initializes a global threshold
\[
d=\max_{(j,a)}\Delta f_a(j\mid \emptyset)
\]
and then sweeps
\[
\theta_0=d,\qquad \theta_{t+1}=(1-\epsilon)\theta_t,
\qquad
\text{until }\theta_t<\frac{\epsilon}{r}d.
\]
At each threshold level, each robot proposes any task with marginal gain at least $\theta$, conflicts are resolved by a coordination step, and many tasks can be allocated in parallel. The resulting guarantee is
\[
f(S)>\frac{1-\epsilon}{2-\epsilon^2}f(OPT)>\Bigl(\frac12-\epsilon\Bigr)f(OPT),
\]
with time complexity
\[
O\Bigl(\frac{r}{\epsilon}\ln\frac r\epsilon\Bigr).
\]
The lazy variant LDTTA maintains cached marginals in a max-heap and uses submodularity to avoid recomputing most gains, while retaining the same $(\frac12-\epsilon)$ approximation guarantee [1909.01239].

Noisy submodular maximization uses thresholding in a statistically adaptive form. The Confident Sample subroutine decides whether an unknown mean is approximately above or below a threshold $w$ with probability at least $1-\delta$, and the monotone cardinality algorithm CTG integrates it into a threshold schedule based on
\[
h(\alpha)=\frac{\ln(k/\alpha)}{\alpha}.
\]
With probability at least $1-\delta$, CTG returns $S$ with
\[
|S|\le k
\quad\text{and}\quad
f(S)\ge (1-e^{-1}-\alpha)f(OPT)-2k\epsilon.
\]
The same framework extends to unconstrained non-monotone maximization through CDG and to matroid constraints through CCTG, the latter obtaining
\[
F(x)\ge (1-e^{-1}-2\epsilon)f(OPT)-\epsilon R.
\]
Here the threshold is not fixed solely by optimization geometry; it is coupled to adaptive sampling and confidence control [2312.00155].

Threshold control can also be the object being optimized. In persistent monitoring on graphs, each agent carries a threshold matrix $\Theta^a$ that determines dwell times and next-hop decisions by testing whether the current node uncertainty falls below $\theta_{ii}^a$ and whether some neighboring uncertainty exceeds $\theta_{ij}^a$. Because local optimization over thresholds is highly initialization-dependent, the paper develops an off-line greedy initialization based on asymptotic cycle analysis and a target-cycle expansion operation. The resulting method constructs a high-performing set of initial thresholds and then refines them by on-line IPA-driven gradient descent. This use of greediness is secondary to the threshold controller itself, but it shows that threshold-greedy methodology also appears as an initialization strategy for hybrid distributed control [1911.02658].

## 6. Thresholds as optimality boundaries and phase transitions

Not all greedy-threshold results are algorithmic selection rules. In some settings, the threshold is a regime boundary that exactly characterizes when a greedy policy is optimal. For a battery limited energy harvesting communication system with i.i.d. arrivals, finite battery capacity $c$, and reward $r(g)$ increasing, concave, and continuously differentiable, the greedy policy uses all available battery energy in each slot. Its baseline throughput is
\[
\underline\gamma(c)=E[r(\min\{X,c\})].
\]
The main theorem states that there exists a critical battery size
\[
c^*=\max\Bigl\{c\ge 0:\ r'(c)\ge \rho(c)\,E[r'(X)\mid X<c]\Bigr\}
\]
such that
\[
\gamma^*(c)=\underline\gamma(c)
\iff
c\le c^*.
\]
For the AWGN reward,
\[
r(g)=\tfrac12\log(1+g),
\]
this becomes
\[
c^*=
\max\Bigl\{c\ge 0:\ \frac1{1+c}\ge \rho(c)\,E\bigl[\tfrac1{1+X}\mid X<c\bigr]\Bigr\}.
\]
The paper also provides lower and upper bounds on $c^*$ and asymptotic formulas for geometric, Poisson, uniform, exponential, and Rayleigh arrivals. A common misconception is that greediness should improve with larger storage; this result shows the opposite structure: greedy is exactly optimal only on the interval $[0,c^*]$ [1909.07895].

A second regime-threshold interpretation appears in online edge coloring. The folklore greedy algorithm uses at most
\[
2\Delta-1
\]
colors. The sharp-threshold results show that this guarantee is unimprovable for deterministic algorithms when $\Delta=O(\log n)$ and for randomized algorithms when $\Delta=O(\sqrt{\log n})$, matching the classical impossibility results. Beyond those regimes, however, the greedy barrier disappears: there is a deterministic online algorithm achieving $(1+o(1))\Delta$-colorings for all $\Delta=\omega(\log n)$, and a randomized algorithm achieving $(1+o(1))\Delta$-colorings already for $\Delta=\omega(\sqrt{\log n})$. The paper explicitly describes these as sharp thresholds for when greedy can be surpassed [2507.21560].

These phase-transition results use “threshold” in a different sense from TGA, OMPT, or DTTA. The threshold does not define admissible local moves; it delineates the structural regime in which greediness is exactly optimal, provably optimal only up to a factor, or asymptotically improvable. This suggests that greedy-threshold is best understood as a broader research motif rather than a single algorithmic template [1909.07895], [2507.21560].

Source: https://www.emergentmind.com/topics/greedy-threshold