Papers
Topics
Authors
Recent
Search
2000 character limit reached

Low-Rank, Sparsity & Smoothness Priors

Updated 9 July 2026
  • LSSP is a framework that combines global low-rank structure, localized sparsity, and smoothness constraints to regularize inverse problems and decompose matrices or tensors.
  • It is applied in areas like dynamic MRI and hyperspectral denoising where enforcing low complexity while preserving local regularity improves recovery and model accuracy.
  • Optimization techniques such as proximal methods, factorization, and IRLS are employed to guarantee efficient convergence and robust recovery under structured regularization conditions.

Low-rank, sparsity, and smoothness priors (LSSP) denote, in the cited literature, a family of regularization principles for inverse problems and matrix or tensor decomposition in which global correlation is represented by low-rank structure, localized innovation or corruption by sparsity, and regularity by smoothness constraints such as total variation, local temporal consistency, or graph-Laplacian smoothness. Typical instances are linear inverse models y=Φx0+wy=\Phi x_0+w and superposition models M=L0+S0\mathbf{M}=\mathbf{L}_0+\mathbf{S}_0, where the estimator is forced toward a low-complexity model or manifold while preserving piecewise-regular behavior and isolating sparse components (Vaiter et al., 2014, Peng et al., 2022, Ting et al., 2024).

1. Geometric and variational foundations

The broadest unifying viewpoint in this area is low-complexity regularization. In the convex inverse-problem setting, one solves

xArgminxRN12λΦxy2+J(x),x^\star \in \operatorname*{Argmin}_{x \in \mathbb{R}^N} \frac{1}{2\lambda}\|\Phi x - y\|^2 + J(x),

or, in the noiseless limit,

xArgminxRNJ(x)s.t.Φx=y.x^\star \in \operatorname*{Argmin}_{x \in \mathbb{R}^N} J(x)\quad \text{s.t.}\quad \Phi x=y.

Here JJ promotes a low-dimensional structure, and the central object is the model tangent subspace

Tx=Lin(J(x)).T_x=\operatorname{Lin}(\partial J(x))^\perp.

This construction supports sparsity, group sparsity, total variation, analysis sparsity, and low-rank regularization under a common geometric language of partial smoothness (Vaiter et al., 2014).

Partial smoothness formalizes the idea that a regularizer is smooth along the correct model and nonsmooth transversally. In the linear-manifold formulation, a convex finite-valued function JJ is partly smooth at xx relative to M\mathcal M if M\mathcal M is a M=L0+S0\mathbf{M}=\mathbf{L}_0+\mathbf{S}_00-manifold around M=L0+S0\mathbf{M}=\mathbf{L}_0+\mathbf{S}_01, M=L0+S0\mathbf{M}=\mathbf{L}_0+\mathbf{S}_02 is M=L0+S0\mathbf{M}=\mathbf{L}_0+\mathbf{S}_03, the tangent space satisfies M=L0+S0\mathbf{M}=\mathbf{L}_0+\mathbf{S}_04, and the subdifferential mapping is continuous at M=L0+S0\mathbf{M}=\mathbf{L}_0+\mathbf{S}_05 relative to M=L0+S0\mathbf{M}=\mathbf{L}_0+\mathbf{S}_06. This framework directly covers M=L0+S0\mathbf{M}=\mathbf{L}_0+\mathbf{S}_07, group sparsity, analysis M=L0+S0\mathbf{M}=\mathbf{L}_0+\mathbf{S}_08, total variation, fused Lasso, and finite-valued polyhedral gauges; it is also closed under addition and under pre-composition by a linear operator, which is the mechanism by which mixed regularization and analysis-type priors are incorporated (Vaiter et al., 2013).

A key qualification is that not every low-rank prior fits the same linear-manifold template. The linear-manifold theory explicitly states that the nuclear norm is partly smooth but lies on the rank manifold, which is not linear, so it is outside that particular framework. By contrast, the broader review of low-complexity regularization treats low rank as a natural extension of sparsity to matrix-valued data and places nuclear-norm regularization within the more general manifold-based version of partial smoothness (Vaiter et al., 2013, Vaiter et al., 2014).

2. Canonical LSSP formulations

LSSP models appear in at least three recurrent forms: additive regularization, superposition decomposition, and structured regularization of gradients or transforms.

The additive form places multiple priors on the same unknown. In dynamic MRI, the standard low-rank plus sparse model

M=L0+S0\mathbf{M}=\mathbf{L}_0+\mathbf{S}_09

is extended by a smoothness term on the background: xArgminxRN12λΦxy2+J(x),x^\star \in \operatorname*{Argmin}_{x \in \mathbb{R}^N} \frac{1}{2\lambda}\|\Phi x - y\|^2 + J(x),0 with

xArgminxRN12λΦxy2+J(x),x^\star \in \operatorname*{Argmin}_{x \in \mathbb{R}^N} \frac{1}{2\lambda}\|\Phi x - y\|^2 + J(x),1

The paper’s interpretation is explicit: the nuclear norm captures global temporal correlation, the xArgminxRN12λΦxy2+J(x),x^\star \in \operatorname*{Argmin}_{x \in \mathbb{R}^N} \frac{1}{2\lambda}\|\Phi x - y\|^2 + J(x),2 term captures sparse dynamic activity after a temporal Fourier transform, and the extra penalty enforces piecewise local consistency between neighboring frames (Ting et al., 2024).

The superposition form separates an observation into a structured component and sparse corruption. A general low-rank-plus-sparse matrix recovery model is treated in factorized form for objectives satisfying restricted strong convexity and smoothness conditions, and concrete instances include robust matrix sensing and robust PCA. In the exact low-rank-and-smooth-plus-sparse direction, 3DCTV-RPCA solves

xArgminxRN12λΦxy2+J(x),x^\star \in \operatorname*{Argmin}_{x \in \mathbb{R}^N} \frac{1}{2\lambda}\|\Phi x - y\|^2 + J(x),3

where

xArgminxRN12λΦxy2+J(x),x^\star \in \operatorname*{Argmin}_{x \in \mathbb{R}^N} \frac{1}{2\lambda}\|\Phi x - y\|^2 + J(x),4

The decisive modeling step is to regularize the gradient maps xArgminxRN12λΦxy2+J(x),x^\star \in \operatorname*{Argmin}_{x \in \mathbb{R}^N} \frac{1}{2\lambda}\|\Phi x - y\|^2 + J(x),5 by nuclear norms, so that local smoothness and correlation of directional derivatives are built into a single convex regularizer rather than appended as an auxiliary term (Zhang et al., 2017, Peng et al., 2022).

A related construction appears in hyperspectral denoising. Conventional SSTV uses only xArgminxRN12λΦxy2+J(x),x^\star \in \operatorname*{Argmin}_{x \in \mathbb{R}^N} \frac{1}{2\lambda}\|\Phi x - y\|^2 + J(x),6 penalties on directional gradients,

xArgminxRN12λΦxy2+J(x),x^\star \in \operatorname*{Argmin}_{x \in \mathbb{R}^N} \frac{1}{2\lambda}\|\Phi x - y\|^2 + J(x),7

LRSTV augments this with a low-rank penalty on the gradient tensors: xArgminxRN12λΦxy2+J(x),x^\star \in \operatorname*{Argmin}_{x \in \mathbb{R}^N} \frac{1}{2\lambda}\|\Phi x - y\|^2 + J(x),8 The intended effect is to preserve the TV-style sparse-gradient prior while exploiting the fact that the gradient tensors are also approximately low-rank under FFT along the spectral dimension (Zeng et al., 2022).

Image recovery provides a simpler two-term variant: xArgminxRN12λΦxy2+J(x),x^\star \in \operatorname*{Argmin}_{x \in \mathbb{R}^N} \frac{1}{2\lambda}\|\Phi x - y\|^2 + J(x),9 later relaxed by replacing rank with a concave, monotonically increasing singular-value surrogate xArgminxRNJ(x)s.t.Φx=y.x^\star \in \operatorname*{Argmin}_{x \in \mathbb{R}^N} J(x)\quad \text{s.t.}\quad \Phi x=y.0. In this formulation, the rank term is interpreted as an xArgminxRNJ(x)s.t.Φx=y.x^\star \in \operatorname*{Argmin}_{x \in \mathbb{R}^N} J(x)\quad \text{s.t.}\quad \Phi x=y.1-type prior on singular values, while anisotropic TV supplies the smoothness prior (Goyal et al., 2020).

3. Algorithmic mechanisms

Optimization strategies in LSSP are correspondingly diverse, but several recurring patterns dominate: factorization-based first-order methods, proximal thresholding, iteratively reweighted least squares, and operator-splitting schemes.

For low-rank plus sparse matrix recovery, a unified factorized framework uses projected gradient descent together with a double thresholding operator. Its theoretical novelty is a structural Lipschitz gradient condition for low-rank plus sparse matrices, introduced to prove local linear convergence while matching the best-known robustness guarantee with respect to sparsity tolerance. The method is designed for general objective functions satisfying restricted strong convexity and smoothness, rather than for a single decomposition model (Zhang et al., 2017).

IRLS-based methods pursue a different strategy: they smooth a nonsmooth objective and then alternate between variable updates and weight updates. For the Schatten-xArgminxRNJ(x)s.t.Φx=y.x^\star \in \operatorname*{Argmin}_{x \in \mathbb{R}^N} J(x)\quad \text{s.t.}\quad \Phi x=y.2 plus xArgminxRNJ(x)s.t.Φx=y.x^\star \in \operatorname*{Argmin}_{x \in \mathbb{R}^N} J(x)\quad \text{s.t.}\quad \Phi x=y.3 low-rank representation problem, the smoothed objective is

xArgminxRNJ(x)s.t.Φx=y.x^\star \in \operatorname*{Argmin}_{x \in \mathbb{R}^N} J(x)\quad \text{s.t.}\quad \Phi x=y.4

with weight matrices

xArgminxRNJ(x)s.t.Φx=y.x^\star \in \operatorname*{Argmin}_{x \in \mathbb{R}^N} J(x)\quad \text{s.t.}\quad \Phi x=y.5

This is a joint treatment of low-rank and sparse penalties after smoothing, and the smoothing is explicitly numerical rather than a structural prior on the underlying signal (Lu et al., 2014).

Strict low-rank optimization develops the singular-value analogue of hard-threshold sparse learning. The objective

xArgminxRNJ(x)s.t.Φx=y.x^\star \in \operatorname*{Argmin}_{x \in \mathbb{R}^N} J(x)\quad \text{s.t.}\quad \Phi x=y.6

is solved by proximal gradient descent

xArgminxRNJ(x)s.t.Φx=y.x^\star \in \operatorname*{Argmin}_{x \in \mathbb{R}^N} J(x)\quad \text{s.t.}\quad \Phi x=y.7

where xArgminxRNJ(x)s.t.Φx=y.x^\star \in \operatorname*{Argmin}_{x \in \mathbb{R}^N} J(x)\quad \text{s.t.}\quad \Phi x=y.8 zeros out singular values below threshold. The accelerated version adds a support projection operator on the singular values, and the support set is shown to shrink monotonically: xArgminxRNJ(x)s.t.Φx=y.x^\star \in \operatorname*{Argmin}_{x \in \mathbb{R}^N} J(x)\quad \text{s.t.}\quad \Phi x=y.9 This converts rank minimization into a sparsity mechanism on the singular spectrum (Zhang et al., 2023).

Application-driven models usually adopt proximal splitting or ADMM. SR-L+S for dynamic MRI uses a proximal gradient method with closed-form updates: JJ0 3DCTV-RPCA uses multi-block ADMM, where each gradient auxiliary variable is updated by singular value thresholding and the main linear system is diagonalized by 3D FFT because the difference operators induce a block-circulant structure. LRSTV for hyperspectral denoising combines ALM, ADMM, HOOI, shrinkage, singular value thresholding, and FFT-based solution of the JJ1-subproblem (Ting et al., 2024, Peng et al., 2022, Zeng et al., 2022).

4. Recovery theory, identifiability, and convergence

Theoretical analysis in this literature separates into geometric recovery guarantees for convex regularization and convergence guarantees for the algorithms that implement these priors.

For convex low-complexity regularization, a recurring requirement is restricted injectivity,

JJ2

augmented by a dual nondegeneracy or irrepresentability condition. Under these hypotheses, the cited theory gives exact recovery in the noiseless case, robust recovery in the noisy case, and model selection consistency in the sense JJ3. In the partial-smoothness review, the minimal norm certificate and linearized pre-certificate are used to characterize manifold identification; once the nondegeneracy condition holds, forward-backward splitting identifies the correct manifold in finitely many iterations and then converges linearly (Vaiter et al., 2013, Vaiter et al., 2014).

Nonconvex and factorized models admit more algorithm-specific guarantees. The unified low-rank-plus-sparse matrix recovery framework proves local linear convergence under restricted strong convexity and smoothness together with the structural Lipschitz gradient condition. The generalized IRLS method proves monotonic descent of the smoothed objective,

JJ4

boundedness of the iterates, vanishing successive differences, and stationarity of limit points; when JJ5, the smoothed problem is convex, so the stationary point is globally optimal. The strict low-rank PGD line proves monotone decrease to a critical point for basic PGD and an asymptotic JJ6 objective rate for accelerated variants, with monotone shrinkage of the singular-value support set (Zhang et al., 2017, Lu et al., 2014, Zhang et al., 2023).

The strongest exact decomposition result in the supplied corpus is for matrices that are jointly low-rank and locally smooth. In 3DCTV-RPCA, if the gradient maps JJ7 satisfy incoherence bounds, the sparse support is uniformly random with cardinality JJ8, and

JJ9

then with probability at least Tx=Lin(J(x)).T_x=\operatorname{Lin}(\partial J(x))^\perp.0 the solution is exact: Tx=Lin(J(x)).T_x=\operatorname{Lin}(\partial J(x))^\perp.1 provided

Tx=Lin(J(x)).T_x=\operatorname{Lin}(\partial J(x))^\perp.2

The proof uses a modification of the golfing scheme and constructs the dual certificate on the gradient maps rather than directly on Tx=Lin(J(x)).T_x=\operatorname{Lin}(\partial J(x))^\perp.3 (Peng et al., 2022).

5. Major application domains

Dynamic MRI is one of the clearest applied realizations of explicit LSSP modeling. In SR-L+S, low rank is assigned to the background component, sparsity to the dynamic component after a temporal Fourier transform, and smoothness to first-order temporal differences of the background. On the PINCAT synthetic cardiac phantom, the best scores were reported for SR-L+S-Tx=Lin(J(x)).T_x=\operatorname{Lin}(\partial J(x))^\perp.4: SER Tx=Lin(J(x)).T_x=\operatorname{Lin}(\partial J(x))^\perp.5 dB, PSNR Tx=Lin(J(x)).T_x=\operatorname{Lin}(\partial J(x))^\perp.6 dB, SSIM Tx=Lin(J(x)).T_x=\operatorname{Lin}(\partial J(x))^\perp.7, and HFEN Tx=Lin(J(x)).T_x=\operatorname{Lin}(\partial J(x))^\perp.8. On in vivo cardiac perfusion MRI, SR-L+S-Tx=Lin(J(x)).T_x=\operatorname{Lin}(\partial J(x))^\perp.9 again performed best with SER JJ0 dB, PSNR JJ1 dB, SSIM JJ2, and HFEN JJ3. The accompanying interpretation is that low rank captures global temporal correlation, whereas the smoothness term captures local temporal consistency that low rank alone may miss (Ting et al., 2024).

Hyperspectral denoising supplies a second major application class. LRSTV starts from the claim that SSTV only exploits gradient sparsity, while the gradient tensors are also approximately low-rank under FFT along the spectral dimension. Quantitatively, the paper reports that on Washington DC Mall under mixed noise with Gaussian variance JJ4, salt-and-pepper noise JJ5, dead lines, and stripes, the proposed method reaches JJ6 dB, compared with JJ7 dB for SSTV and JJ8 dB for LRTDTV; under heavier Gaussian noise JJ9, it reaches xx0 dB, compared with xx1 dB for SSTV and xx2 dB for LRTDTV. A separate summary statement in the abstract is that the proposed model can get xx3 dB improvement of PSNR (Zeng et al., 2022).

The exact low-rank-and-smooth-plus-sparse perspective is also effective in robust decomposition tasks. In synthetic phase-transition experiments, 3DCTV-RPCA reports a successful recovery area of xx4 versus xx5 for PCP, and the residuals drop sharply in the first xx6 iterations and are near zero by xx7 iterations. The same model is then validated on hyperspectral image denoising, multispectral image denoising, and surveillance-video background modeling, where it is described as best or near-best in most noise settings and competitive in AUC for foreground detection on the Li dataset (Peng et al., 2022).

Image recovery from partial observations provides a more classical matrix setting. IRNN_TV combines nonconvex singular-value shrinkage with anisotropic total variation and reports consistent gains over low-rank-only IRNN, LMaFit, and TV-only TFOCS in both image completion and image completion with noise. The reported qualitative pattern is that low rank alone can fail when entire rows or columns are missing, whereas the combined model reconstructs local features and sharper structures while still exploiting global low-rank structure (Goyal et al., 2020).

6. Scope, misconceptions, and adjacent models

A recurring misconception is that “smoothness” has a unique meaning across this literature. In some works it is an explicit structural prior, such as temporal finite differences in dynamic MRI, anisotropic TV in image recovery, or graph-Laplacian smoothness in graph learning. In the generalized IRLS framework, by contrast, the smoothing parameter xx8 is introduced to make the objective differentiable and easier to optimize; the paper explicitly states that this is optimization smoothing or regularization, not a structural smoothness prior on the unknown signal (Ting et al., 2024, Goyal et al., 2020, Chepuri et al., 2016, Lu et al., 2014).

A second misconception is that every low-rank-plus-sparse model is automatically an LSSP model. Several adjacent works cover only part of the triad. Tuning-free Bayesian matrix and tensor completion with Horseshoe and Horseshoe+ priors is a rank-sparsity model in which shrinkage acts on factor columns, but the paper does not introduce an explicit smoothness prior in the usual sense of penalizing differences across neighboring indices. Sparse graph learning under a smoothness prior combines graph smoothness and explicit edge sparsity, but does not formulate low rank as a primary objective. Speech enhancement with an online estimated dictionary imposes low-rank and sparsity constraints on speech activation and background noise, yet states that there is no explicit smoothness regularizer. Interventional MRI with framelet-based LS decomposition uses low rank, temporal sparsity, and additional spatial sparsity through framelets, but the final proposed objective contains no separate smoothness term, only an implicit piecewise-smoothness mechanism through framelet sparsity (Gilbert et al., 2019, Chepuri et al., 2016, Sun et al., 2016, He et al., 2021).

A third point of contention concerns whether low rank and smoothness should be encoded separately or through a single integrated regularizer. SR-L+S adds a smoothness term to a conventional L+S objective; 3DCTV-RPCA integrates low-rankness and local smoothness by penalizing the nuclear norms of directional gradient maps; LRSTV likewise combines xx9 gradient sparsity with a tensor nuclear norm on the gradients themselves. This suggests that there is no single canonical LSSP formulation: some models treat low rank, sparsity, and smoothness as separable priors with separate trade-off parameters, while others fuse them into a joint regularizer on transformed or differentiated representations (Ting et al., 2024, Peng et al., 2022, Zeng et al., 2022).

The main conceptual lesson of the cited work is therefore not the existence of one universal objective, but the emergence of a common design principle: use low-rank structure to encode global correlation, sparsity to isolate innovation or corruption, and smoothness to preserve local regularity, then exploit the geometry of the resulting low-complexity model for identifiability, stability, and fast first-order optimization (Vaiter et al., 2014, Zhang et al., 2017).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Low-Rank, Sparsity, and Smoothness Priors (LSSP).