- The paper introduces a Huberized, folded-concave estimator that robustly demixes sparse signals from nonlinear and contaminated observations.
- It proposes the NLD-PALM algorithm, which guarantees convergence under the Kurdyka–Łojasiewicz property and exhibits local R-linear convergence.
- The study establishes statistical recovery bounds and demonstrates significant error improvements over traditional convex methods in the presence of outliers.
Demixing Sparse Signals from Nonlinear Observations using Generalized Non-convex Regularization
The paper addresses the problem of sparse signal demixing from nonlinear observations, where the measurements are of the form $y_i = g({a_i}{\bPhi w^* + \bPsi z^*}) + e_i$ with m≪n. Here, $\bPhi$ and $\bPsi$ are incoherent orthonormal bases, g is a link function often corresponding to a nonlinear front-end (e.g., saturation, clipping, companding), and the noise term ei may be heavy-tailed or contaminated. The task is to recover both sparse components from significantly undersampled and potentially nonlinear data.
Traditional approaches are limited in this setting. Greedy algorithms such as Demixing Hard Thresholding (DHT) require exact sparsity knowledge and offer no robustness guarantees, while convex ℓ1 demixing suffers from shrinkage bias and fails to attain oracle performance, especially under strong noise or gross outliers.
This work proposes a Huberized, folded-concave regularized estimator for robust, bias-minimized demixing in both known- and unknown-link regimes, extending non-convex regularization theory (SCAD, MCP) to handle nonlinear and contaminated measurements.
Regularization Framework and Estimator
The estimator minimizes a Huberized, non-convex objective: $F(w, z) = \frac{1}{m} \sum_{i=1}^m h_H\left(y_i - g(a_i^\top (\bPhi w + \bPsi z))\right) + P_\lambda(w) + P_\lambda(z)$
where hH is the Huber loss (threshold H), and m≪n0 belongs to the folded-concave penalty class (e.g., SCAD, MCP), ensuring unbiased recovery for large coefficients and superior oracle behavior compared to m≪n1.
Huberization obviates compatibility conditions required by squared-loss analysis, ensuring concentration of empirical processes and robust behavior to only finite-variance symmetric noise. This is crucial in applications where gross outliers or heavy-tailed contamination are present.
Algorithmic Contributions: NLD-PALM
A core contribution is the Nonlinear Demixing Proximal Alternating Linearized Minimization (NLD-PALM) algorithm, a two-block proximal alternating scheme with per-block backtracking and over-relaxation parameter m≪n2. The algorithm guarantees whole-iterate convergence to critical points under the Kurdyka–Łojasiewicz property, with local m≪n3-linear convergence when RSC holds. The requirement m≪n4 is necessary for sufficient decrease with non-convex penalties, distinguishing the method from fixed-step convex proximal algorithms.
Proximal mappings for SCAD, MCP, and m≪n5 are closed-form, and the methodology ensures boundedness of iterates even under non-coercive penalties and loss functions.
Statistical Recovery Theory
Restricted strong convexity (RSC) of the Huberized loss is established via an exact sign-definite decomposition, ensuring curvature at every sample level. The key result is: m≪n6
where m≪n7 and m≪n8 aggregates problem and penalty constants.
The main error bounds at stationary points demonstrate that for an appropriate choice of m≪n9, the estimator achieves: $\bPhi$0
with $\bPhi$1 total sparsity, $\bPhi$2 ambient dimension, and $\bPhi$3 sample count. The oracle regime is established under a beta-min condition ($\bPhi$4 large enough), yielding $\bPhi$5 rates without $\bPhi$6 or shrinkage bias—unattainable by convex $\bPhi$7 demixing.
The estimator requires no knowledge of the sparsity levels, and its guarantees apply under symmetric noise with only finite variance.
Robustness and Unknown-Link Recovery
Huberization confers strong robustness to gross outliers and heavy tails. Experiments with $\bPhi$8 gross outlier contamination show $\bPhi$9 improvement in median relative error over squared-loss methods. The estimator's guarantees extend to unknown monotone links via a surrogate linear objective and clipped Plan–Vershynin decoupling, achieving comparable rates modulo a $\bPsi$0 irreducible floor induced by the nonlinearity.

Figure 1: Phase transition and median error versus sample size $\bPsi$1; SCAD achieves earlier transition, significantly lower error, and avoids shrinkage bias compared to $\bPsi$2 and DHT.
Numerical Experiments
Comprehensive experiments at $\bPsi$3 and various sparsity and noise regimes validate the theoretical findings. SCAD/MCP estimators attain phase transition approximately $\bPsi$4--$\bPsi$5 earlier than DHT, with SCAD's median error $\bPsi$6--$\bPsi$7 versus $\bPsi$8--$\bPsi$9 for DHT in scarce sampling regimes. g0 fails to attain exact recovery, instead plateauing at its shrinkage bias as asserted by theory.
Performance is robust across noise regimes: median relative error as low as g1 under g2 gross outliers and parity with squared-loss under clean noise. Measured error scales as g3 and g4 up to the information-theoretic floor.
A saturation-demixing application demonstrates the method's ability to successfully separate spikes and smooth backgrounds even under saturating amplifier nonlinearity and contamination.

Figure 2: Saturation demixing of spike-plus-smooth signals under nonlinear g5 observations and gross outliers; Huberized SCAD achieves relative error g6, outperforming squared-loss SCAD at g7.
Practical and Theoretical Implications
The findings substantiate that Huberized, folded-concave regularization is capable of robust, bias-minimized sparse demixing from highly noisy, nonlinear, and undersampled measurements. The theoretical guarantees hold at every localized stationary point, independent of sparsity knowledge, challenging the necessity of convexity in high-dimensional regression and signal recovery.
Practically, these methods enable recovery in compressed sensing, imaging, and source separation applications with nonlinear sensing front-ends. This is relevant for real-world acquisition and demixing tasks where saturation, quantization, and contamination are inherent.
Theoretically, the work opens new avenues for extending non-convex statistical theory to broader penalty classes (g8 quasi-norms), periodic and quantized links, and more generalized superposition models. Future developments may include deep-unfolded NLD-PALM architectures with trainable thresholding, fully removing dimensional restrictions for support recovery via primal–dual witness methods.
Conclusion
This paper establishes a robust, non-convex regularization approach for nonlinear sparse demixing, with provable convergence and statistical recovery guarantees. Huberized SCAD/MCP estimators provide earlier phase transitions and oracle rates, outperforming traditional DHT and g9 methods, especially in the presence of gross contamination and nonlinear front-ends. This framework is practically relevant and lays the foundation for deeper statistical and algorithmic exploration in nonlinear observation models and non-convex regularization.