Papers
Topics
Authors
Recent
Search
2000 character limit reached

Demixing Sparse Signals from Nonlinear Observations using Generalized Non-convex Regularization

Published 12 Jul 2026 in stat.ML, cs.LG, and eess.SP | (2607.10618v1)

Abstract: We consider the recovery of a pair of sparse vectors from a limited number of nonlinear observations of their superposition: $y_i=g(\inner{\ba_i}{\bPhi\bw<sup>\ast+\bPsi\bz<sup>\ast})+e_i$, i=1,,mi=1,\dots,m, with mnm\ll n, incoherent orthonormal bases $\bPhi,\bPsi$, a scalar link gg, and noise eie_i that may be heavy-tailed or contaminated. We propose a regularization-based framework combining a Huberized data fidelity with generalized folded-concave penalties (SCAD, MCP), and a two-block proximal alternating algorithm with backtracking (NLD-PALM) whose whole iterate sequence provably converges to critical points under the Kurdyka--Łojasiewicz property, with local linear rates. On the statistical side we establish restricted strong convexity of the Huberized nonlinear loss through an exact sign-definite decomposition, and derive estimation error bounds of order σslog(n)/mσ\sqrt{s\log(n)/m} that hold at \emph{every} localized stationary point, an oracle rate σs/mσ\sqrt{s/m} free of logn\log n and shrinkage bias under a beta-min condition, and a co-equal recovery theorem for \emph{unknown} monotone links via a linear surrogate and a clipped Plan--Vershynin decoupling. The estimator requires no knowledge of the sparsity levels, and its guarantees hold under symmetric noise with only finite variance. Experiments at n=512n=512 under a frozen data-driven regularization rule show an earlier phase transition than convex 1\ell_1 demixing and greedy hard-thresholding baselines, a 35×35\times accuracy advantage over squared-loss estimation under 5%5\% gross outliers, and successful demixing of spike-plus-background signals observed through a saturating amplifier.

Authors (1)

Summary

  • The paper introduces a Huberized, folded-concave estimator that robustly demixes sparse signals from nonlinear and contaminated observations.
  • It proposes the NLD-PALM algorithm, which guarantees convergence under the Kurdyka–Łojasiewicz property and exhibits local R-linear convergence.
  • The study establishes statistical recovery bounds and demonstrates significant error improvements over traditional convex methods in the presence of outliers.

Demixing Sparse Signals from Nonlinear Observations using Generalized Non-convex Regularization

Problem Formulation and Motivation

The paper addresses the problem of sparse signal demixing from nonlinear observations, where the measurements are of the form $y_i = g({a_i}{\bPhi w^* + \bPsi z^*}) + e_i$ with mnm \ll n. Here, $\bPhi$ and $\bPsi$ are incoherent orthonormal bases, gg is a link function often corresponding to a nonlinear front-end (e.g., saturation, clipping, companding), and the noise term eie_i may be heavy-tailed or contaminated. The task is to recover both sparse components from significantly undersampled and potentially nonlinear data.

Traditional approaches are limited in this setting. Greedy algorithms such as Demixing Hard Thresholding (DHT) require exact sparsity knowledge and offer no robustness guarantees, while convex 1\ell_1 demixing suffers from shrinkage bias and fails to attain oracle performance, especially under strong noise or gross outliers.

This work proposes a Huberized, folded-concave regularized estimator for robust, bias-minimized demixing in both known- and unknown-link regimes, extending non-convex regularization theory (SCAD, MCP) to handle nonlinear and contaminated measurements.

Regularization Framework and Estimator

The estimator minimizes a Huberized, non-convex objective: $F(w, z) = \frac{1}{m} \sum_{i=1}^m h_H\left(y_i - g(a_i^\top (\bPhi w + \bPsi z))\right) + P_\lambda(w) + P_\lambda(z)$ where hHh_H is the Huber loss (threshold HH), and mnm \ll n0 belongs to the folded-concave penalty class (e.g., SCAD, MCP), ensuring unbiased recovery for large coefficients and superior oracle behavior compared to mnm \ll n1.

Huberization obviates compatibility conditions required by squared-loss analysis, ensuring concentration of empirical processes and robust behavior to only finite-variance symmetric noise. This is crucial in applications where gross outliers or heavy-tailed contamination are present.

Algorithmic Contributions: NLD-PALM

A core contribution is the Nonlinear Demixing Proximal Alternating Linearized Minimization (NLD-PALM) algorithm, a two-block proximal alternating scheme with per-block backtracking and over-relaxation parameter mnm \ll n2. The algorithm guarantees whole-iterate convergence to critical points under the Kurdyka–Łojasiewicz property, with local mnm \ll n3-linear convergence when RSC holds. The requirement mnm \ll n4 is necessary for sufficient decrease with non-convex penalties, distinguishing the method from fixed-step convex proximal algorithms.

Proximal mappings for SCAD, MCP, and mnm \ll n5 are closed-form, and the methodology ensures boundedness of iterates even under non-coercive penalties and loss functions.

Statistical Recovery Theory

Restricted strong convexity (RSC) of the Huberized loss is established via an exact sign-definite decomposition, ensuring curvature at every sample level. The key result is: mnm \ll n6 where mnm \ll n7 and mnm \ll n8 aggregates problem and penalty constants.

The main error bounds at stationary points demonstrate that for an appropriate choice of mnm \ll n9, the estimator achieves: $\bPhi$0 with $\bPhi$1 total sparsity, $\bPhi$2 ambient dimension, and $\bPhi$3 sample count. The oracle regime is established under a beta-min condition ($\bPhi$4 large enough), yielding $\bPhi$5 rates without $\bPhi$6 or shrinkage bias—unattainable by convex $\bPhi$7 demixing.

The estimator requires no knowledge of the sparsity levels, and its guarantees apply under symmetric noise with only finite variance.

Huberization confers strong robustness to gross outliers and heavy tails. Experiments with $\bPhi$8 gross outlier contamination show $\bPhi$9 improvement in median relative error over squared-loss methods. The estimator's guarantees extend to unknown monotone links via a surrogate linear objective and clipped Plan–Vershynin decoupling, achieving comparable rates modulo a $\bPsi$0 irreducible floor induced by the nonlinearity.

Figure 1

Figure 1: Phase transition and median error versus sample size $\bPsi$1; SCAD achieves earlier transition, significantly lower error, and avoids shrinkage bias compared to $\bPsi$2 and DHT.

Numerical Experiments

Comprehensive experiments at $\bPsi$3 and various sparsity and noise regimes validate the theoretical findings. SCAD/MCP estimators attain phase transition approximately $\bPsi$4--$\bPsi$5 earlier than DHT, with SCAD's median error $\bPsi$6--$\bPsi$7 versus $\bPsi$8--$\bPsi$9 for DHT in scarce sampling regimes. gg0 fails to attain exact recovery, instead plateauing at its shrinkage bias as asserted by theory.

Performance is robust across noise regimes: median relative error as low as gg1 under gg2 gross outliers and parity with squared-loss under clean noise. Measured error scales as gg3 and gg4 up to the information-theoretic floor.

A saturation-demixing application demonstrates the method's ability to successfully separate spikes and smooth backgrounds even under saturating amplifier nonlinearity and contamination.

Figure 2

Figure 2: Saturation demixing of spike-plus-smooth signals under nonlinear gg5 observations and gross outliers; Huberized SCAD achieves relative error gg6, outperforming squared-loss SCAD at gg7.

Practical and Theoretical Implications

The findings substantiate that Huberized, folded-concave regularization is capable of robust, bias-minimized sparse demixing from highly noisy, nonlinear, and undersampled measurements. The theoretical guarantees hold at every localized stationary point, independent of sparsity knowledge, challenging the necessity of convexity in high-dimensional regression and signal recovery.

Practically, these methods enable recovery in compressed sensing, imaging, and source separation applications with nonlinear sensing front-ends. This is relevant for real-world acquisition and demixing tasks where saturation, quantization, and contamination are inherent.

Theoretically, the work opens new avenues for extending non-convex statistical theory to broader penalty classes (gg8 quasi-norms), periodic and quantized links, and more generalized superposition models. Future developments may include deep-unfolded NLD-PALM architectures with trainable thresholding, fully removing dimensional restrictions for support recovery via primal–dual witness methods.

Conclusion

This paper establishes a robust, non-convex regularization approach for nonlinear sparse demixing, with provable convergence and statistical recovery guarantees. Huberized SCAD/MCP estimators provide earlier phase transitions and oracle rates, outperforming traditional DHT and gg9 methods, especially in the presence of gross contamination and nonlinear front-ends. This framework is practically relevant and lays the foundation for deeper statistical and algorithmic exploration in nonlinear observation models and non-convex regularization.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 6 likes about this paper.