Papers
Topics
Authors
Recent
Search
2000 character limit reached

Statistical inverse learning and 1\ell^1-regularization

Published 8 Jul 2026 in stat.ML, cs.LG, and math.ST | (2607.07468v1)

Abstract: We study the recovery of sparse functions from finite, noisy, and indirect observations in the framework of statistical inverse learning. The unknown is modeled as an element of <sup>1\ell<sup>1, and observations are generated through a possibly nonlinear forward operator A:<sup>1</sup>HA:\ell<sup>1\to</sup> H, where HH is a vector-valued reproducing kernel Hilbert space. We propose an <sup>1\ell<sup>1-regularized empirical risk minimizer and develop a theoretical analysis of its statistical properties. Under mild assumptions, we establish almost-sure consistency and derive non-asymptotic high-probability convergence rates in both the prediction and <sup>1\ell<sup>1 reconstruction norms. The rates depend on the source smoothness parameter rr, characterized by a variational source condition, and the effective dimension exponent bb, describing the polynomial spectral decay of the covariance operator. We further prove matching minimax lower bounds, showing that the obtained convergence rates are optimal. To relate the theory to practical sparsity models, we consider finitely smoothing operators of the form A=GSA=G\circ S, where SS is a synthesis operator, and show that approximation-space assumptions imply the required variational source conditions. In particular, we prove that membership in the approximation space ktk_t is equivalent to polynomial decay of the best nn-term approximation error. Finally, we verify the assumptions for two representative inverse problems: reaction coefficient identification in elliptic PDEs and sparse computed tomography. For filtered Radon transforms, we derive explicit effective-dimension asymptotics, yielding concrete convergence rates for standard image models and sparsifying systems.

Summary

  • The paper establishes consistency and matching minimax convergence rates for ℓ¹-regularized empirical risk minimization in nonlinear inverse problems with sparse, indirect, and noisy data.
  • Its rates depend on source smoothness and effective dimension, achieving an ℓ¹ reconstruction rate of n^{-r/(1+b-br)} under polynomial spectral decay and a variational source condition.
  • The paper links best n-term approximation, sparse approximation spaces, and variational source conditions, then validates the theory for elliptic coefficient recovery, filtered Radon transforms, and direct learning.

Setting and motivation

The paper develops a statistical learning theory for sparse recovery in nonlinear inverse problems. The unknown is a sequence f1f_\dagger \in \ell^1, observed through random, noisy, indirect data

yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,

where A:D(A)1HA : D(A) \cap \ell^1 \to \mathcal{H} is a possibly nonlinear forward operator into a vector-valued reproducing kernel Hilbert space (vv-RKHS) continuously embedded in H=L2(X,μ;Y)H = L^2(\mathcal{X},\mu;\mathcal{Y}). The estimator is the 1\ell^1-regularized empirical risk minimizer

fz,λ=argminfD(A)11ni=1nA(f)(xi)yiY2+λf1.f_{\mathbf{z},\lambda} = \arg\min_{f \in D(A)\cap\ell^1} \frac{1}{n}\sum_{i=1}^n \|A(f)(x_i)-y_i\|_{\mathcal{Y}}^2 + \lambda \|f\|_{\ell^1}.

The framework extends the effective-dimension analysis of Caponnetto–De Vito (2607.07468) and the linear statistical inverse learning of Blanchard–Mücke to a Banach-space setting with sparsity-promoting penalties. Compared with prior work on nonlinear Tikhonov regularization in vv-RKHSs [Rastogi et al., 2020], which yields rates nr/(2r+b+1)n^{-r/(2r+b+1)} under Hilbert-space penalties, the present paper obtains the rate nr/(1+bbr)n^{-r/(1+b-br)}, where rr is a source smoothness index from a variational source condition (VSC) and bb is the polynomial decay exponent of the effective dimension yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,0.

Assumptions

The analysis rests on six structural assumptions:

  • True solution: the conditional mean yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,1 equals yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,2.
  • Sub-exponential noise: a Bernstein-type moment condition with constants yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,3; notably, Gaussian white noise in infinite-dimensional output spaces is explicitly excluded.
  • Kernel regularity: uniformly bounded Hilbert–Schmidt norm yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,4 and measurability of kernel inner products.
  • Lipschitz continuity: yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,5 injective and Lipschitz in yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,6 with constant yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,7.
  • Variational source condition: for all admissible yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,8,

yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,9

with concave index function A:D(A)1HA : D(A) \cap \ell^1 \to \mathcal{H}0; the main results take A:D(A)1HA : D(A) \cap \ell^1 \to \mathcal{H}1.

  • Weighted bi-Lipschitz property: two-sided control between A:D(A)1HA : D(A) \cap \ell^1 \to \mathcal{H}2 and A:D(A)1HA : D(A) \cap \ell^1 \to \mathcal{H}3 for weights vanishing at infinity.

The model class A:D(A)1HA : D(A) \cap \ell^1 \to \mathcal{H}4 collects all distributions satisfying these assumptions with A:D(A)1HA : D(A) \cap \ell^1 \to \mathcal{H}5 and polynomial spectral decay.

Consistency

Under Assumptions 1–4 plus weak-to-weak sequential continuity of A:D(A)1HA : D(A) \cap \ell^1 \to \mathcal{H}6, and the parameter choice A:D(A)1HA : D(A) \cap \ell^1 \to \mathcal{H}7 with A:D(A)1HA : D(A) \cap \ell^1 \to \mathcal{H}8, the estimator converges almost surely: A:D(A)1HA : D(A) \cap \ell^1 \to \mathcal{H}9. The proof combines Borel–Cantelli application of high-probability bounds on H=L2(X,μ;Y)H = L^2(\mathcal{X},\mu;\mathcal{Y})0 and H=L2(X,μ;Y)H = L^2(\mathcal{X},\mu;\mathcal{Y})1 (of order H=L2(X,μ;Y)H = L^2(\mathcal{X},\mu;\mathcal{Y})2), almost-sure boundedness of the iterates in H=L2(X,μ;Y)H = L^2(\mathcal{X},\mu;\mathcal{Y})3, weak-* compactness via Banach–Alaoglu–Bourbaki, identification of the limit through injectivity, and norm convergence via Scheffé's lemma. Existence of minimizers holds under continuity/compactness of H=L2(X,μ;Y)H = L^2(\mathcal{X},\mu;\mathcal{Y})4, though uniqueness is not guaranteed for nonlinear H=L2(X,μ;Y)H = L^2(\mathcal{X},\mu;\mathcal{Y})5.

Upper convergence rates

The core estimate separates bias and variance through the Fenchel conjugate H=L2(X,μ;Y)H = L^2(\mathcal{X},\mu;\mathcal{Y})6. With confidence H=L2(X,μ;Y)H = L^2(\mathcal{X},\mu;\mathcal{Y})7,

H=L2(X,μ;Y)H = L^2(\mathcal{X},\mu;\mathcal{Y})8

For H=L2(X,μ;Y)H = L^2(\mathcal{X},\mu;\mathcal{Y})9 and the explicit choice 1\ell^10, this yields

1\ell^11

with high probability. Via an interpolation inequality across weighted sequence spaces 1\ell^12 with weights 1\ell^13, the reconstruction bound generalizes to

1\ell^14

and this rate holds uniformly over 1\ell^15 in 1\ell^16 expectation for all 1\ell^17 by tail integration. Two special cases are worth noting: for 1\ell^18 the exponent reduces to 1\ell^19, while for fz,λ=argminfD(A)11ni=1nA(f)(xi)yiY2+λf1.f_{\mathbf{z},\lambda} = \arg\min_{f \in D(A)\cap\ell^1} \frac{1}{n}\sum_{i=1}^n \|A(f)(x_i)-y_i\|_{\mathcal{Y}}^2 + \lambda \|f\|_{\ell^1}.0 it becomes fz,λ=argminfD(A)11ni=1nA(f)(xi)yiY2+λf1.f_{\mathbf{z},\lambda} = \arg\min_{f \in D(A)\cap\ell^1} \frac{1}{n}\sum_{i=1}^n \|A(f)(x_i)-y_i\|_{\mathcal{Y}}^2 + \lambda \|f\|_{\ell^1}.1 — independent of fz,λ=argminfD(A)11ni=1nA(f)(xi)yiY2+λf1.f_{\mathbf{z},\lambda} = \arg\min_{f \in D(A)\cap\ell^1} \frac{1}{n}\sum_{i=1}^n \|A(f)(x_i)-y_i\|_{\mathcal{Y}}^2 + \lambda \|f\|_{\ell^1}.2, reflecting that the weighted fz,λ=argminfD(A)11ni=1nA(f)(xi)yiY2+λf1.f_{\mathbf{z},\lambda} = \arg\min_{f \in D(A)\cap\ell^1} \frac{1}{n}\sum_{i=1}^n \|A(f)(x_i)-y_i\|_{\mathcal{Y}}^2 + \lambda \|f\|_{\ell^1}.3 error is dominated by prediction-norm behavior rather than sparsity.

Minimax lower bounds

Lower bounds are established by an information-theoretic argument: a packing construction over sign vectors (giving fz,λ=argminfD(A)11ni=1nA(f)(xi)yiY2+λf1.f_{\mathbf{z},\lambda} = \arg\min_{f \in D(A)\cap\ell^1} \frac{1}{n}\sum_{i=1}^n \|A(f)(x_i)-y_i\|_{\mathcal{Y}}^2 + \lambda \|f\|_{\ell^1}.4 codewords at Hamming distance at least fz,λ=argminfD(A)11ni=1nA(f)(xi)yiY2+λf1.f_{\mathbf{z},\lambda} = \arg\min_{f \in D(A)\cap\ell^1} \frac{1}{n}\sum_{i=1}^n \|A(f)(x_i)-y_i\|_{\mathcal{Y}}^2 + \lambda \|f\|_{\ell^1}.5), a family of discrete measures fz,λ=argminfD(A)11ni=1nA(f)(xi)yiY2+λf1.f_{\mathbf{z},\lambda} = \arg\min_{f \in D(A)\cap\ell^1} \frac{1}{n}\sum_{i=1}^n \|A(f)(x_i)-y_i\|_{\mathcal{Y}}^2 + \lambda \|f\|_{\ell^1}.6 adapted from Caponnetto–De Vito, KL-divergence estimates controlled by the upper Lipschitz property of fz,λ=argminfD(A)11ni=1nA(f)(xi)yiY2+λf1.f_{\mathbf{z},\lambda} = \arg\min_{f \in D(A)\cap\ell^1} \frac{1}{n}\sum_{i=1}^n \|A(f)(x_i)-y_i\|_{\mathcal{Y}}^2 + \lambda \|f\|_{\ell^1}.7, and Fano's inequality. Under a scale condition on the weights fz,λ=argminfD(A)11ni=1nA(f)(xi)yiY2+λf1.f_{\mathbf{z},\lambda} = \arg\min_{f \in D(A)\cap\ell^1} \frac{1}{n}\sum_{i=1}^n \|A(f)(x_i)-y_i\|_{\mathcal{Y}}^2 + \lambda \|f\|_{\ell^1}.8 over a block of fz,λ=argminfD(A)11ni=1nA(f)(xi)yiY2+λf1.f_{\mathbf{z},\lambda} = \arg\min_{f \in D(A)\cap\ell^1} \frac{1}{n}\sum_{i=1}^n \|A(f)(x_i)-y_i\|_{\mathcal{Y}}^2 + \lambda \|f\|_{\ell^1}.9 indices, any estimator suffers error exceeding nr/(2r+b+1)n^{-r/(2r+b+1)}0 with probability at least nr/(2r+b+1)n^{-r/(2r+b+1)}1, where balancing the packing size nr/(2r+b+1)n^{-r/(2r+b+1)}2 against the KL term gives exactly

nr/(2r+b+1)n^{-r/(2r+b+1)}3

Consequently no algorithm can beat this rate in nr/(2r+b+1)n^{-r/(2r+b+1)}4 over nr/(2r+b+1)n^{-r/(2r+b+1)}5, matching the upper bound. The remark verifying the weight-block condition shows it is compatible with polynomially decaying weights nr/(2r+b+1)n^{-r/(2r+b+1)}6 provided nr/(2r+b+1)n^{-r/(2r+b+1)}7.

From approximation spaces to source conditions

A central structural contribution connects sparse approximation theory to the VSC. For the space

nr/(2r+b+1)n^{-r/(2r+b+1)}8

the paper proves three linked facts:

  1. Equivalence: nr/(2r+b+1)n^{-r/(2r+b+1)}9 iff a power-type inequality nr/(1+bbr)n^{-r/(1+b-br)}0 holds (a variant of a lemma of Miller–Hohage).
  2. VSC: under the lower bi-Lipschitz bound, nr/(1+bbr)n^{-r/(1+b-br)}1 membership implies the VSC with index nr/(1+bbr)n^{-r/(1+b-br)}2; conversely, the VSC implies nr/(1+bbr)n^{-r/(1+b-br)}3 membership with nr/(1+bbr)n^{-r/(1+b-br)}4.
  3. Best nr/(1+bbr)n^{-r/(1+b-br)}5-term characterization: nr/(1+bbr)n^{-r/(1+b-br)}6 iff the best nr/(1+bbr)n^{-r/(1+b-br)}7-term approximation error satisfies nr/(1+bbr)n^{-r/(1+b-br)}8 in nr/(1+bbr)n^{-r/(1+b-br)}9, proved via a layer-cake argument.

Together these give the chain: polynomial decay of rr0 rr1 rr2 membership rr3 VSC with rr4 rr5 the optimal rate. An important consequence is that within this framework the exploitable smoothness index is confined to rr6: rr7 regularization extracts sparsity efficiently but cannot exploit smoothness beyond this range, unlike rr8 regularization. This is a genuine limitation of the method, not merely of the analysis.

Verified examples

Finitely smoothing operators. For composite operators rr9 with bb0 a wavelet or shearlet synthesis operator and bb1 satisfying a two-sided bb2 equivalence, both parts of the bi-Lipschitz assumption hold with explicit weights (bb3 for wavelets; shifted by bb4 for shearlets). Concrete image models then fix bb5: piecewise bb6 signals in 1D give bb7; bounded-variation images give bb8 (hence bb9, no rate); cartoon-like images under shearlets or curvelets give yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,00, i.e., yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,01.

Elliptic coefficient identification. Recovering a reaction coefficient yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,02 from the solution map of yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,03: the output space yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,04 is an RKHS whose covariance eigenvalues satisfy yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,05, giving yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,06 for yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,07.

Filtered Radon transform. The unfiltered Radon transform fails the framework: its natural RKHS kernel operators yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,08 are not Hilbert–Schmidt (since yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,09 is multiplication by a nonzero line-intersection length, hence noncompact). Smoothing by convolution along the detector variable restores compliance. Via the Fourier slice theorem, the covariance is a pseudodifferential operator with symbol yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,10; Weyl-law asymptotics then give:

Filter Eigenvalue decay Effective dimension
Boxcar yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,11 yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,12, so yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,13
Gaussian faster than any yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,14 slower than any yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,15, any yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,16

Combining these with the image models yields concrete rates: for cartoon-like images with boxcar filtering, yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,17 with yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,18; with Gaussian filtering, yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,19 with yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,20. For finitely supported (sparse) coefficients the rates improve to yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,21 (boxcar) and yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,22 (Gaussian). Numerical SVD computations with the ASTRA toolbox strip model confirm a pre-asymptotic polynomial decay slightly steeper than the unfiltered prediction yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,23, with the caveat that sketching-based low-rank SVD may be inaccurate for small singular values.

Direct learning. When yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,24, yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,25 reduces to the synthesis operator, the bi-Lipschitz property holds with equality and constant yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,26, and the framework recovers classical vector-valued kernel ridge regression as a special case — providing a clean consistency check against Caponnetto–De Vito and Lasso-type theory without restricted isometry or incoherence conditions.

Limitations and open questions

Several restrictions are acknowledged explicitly. The sub-exponential noise assumption excludes infinite-dimensional Gaussian white noise. The lower bi-Lipschitz bound rules out infinitely smoothing operators such as the backward heat equation and electrical impedance tomography; extending the theory to logarithmic source conditions or Sobolev-type variational inequalities for such problems remains open. The optimal parameter yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,27 depends on unknown quantities yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,28 and yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,29; while Corollary 4 offers a dual-function characterization amenable to data-driven selection analogous to the discrepancy principle, a rigorous adaptive procedure with finite-sample guarantees is not provided. Finally, whether the optimal rates remain achievable under online or stochastic proximal algorithms, and how the analysis extends to Fréchet-differentiable yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,30 via iteratively reweighted schemes, are questions the paper leaves unanswered.

Conclusion

The paper establishes a complete minimax theory for yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,31-regularized nonlinear statistical inverse learning in vv-RKHSs: almost-sure consistency, high-probability rates yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,32 in interpolated norms, and matching information-theoretic lower bounds over the class yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,33. Its distinctive contribution is the equivalence linking best yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,34-term approximation decay, the spaces yi=A(f)(xi)+εi,i=1,,n,y_i = A(f_\dagger)(x_i) + \varepsilon_i, \qquad i = 1,\dots,n,35, and variational source conditions, which makes the abstract assumptions verifiable for concrete sparsity models. Applications to elliptic coefficient identification and filtered Radon transforms demonstrate that the theory produces explicit, filter-dependent convergence guarantees for standard tomographic imaging models.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.