- The paper establishes consistency and matching minimax convergence rates for ℓ¹-regularized empirical risk minimization in nonlinear inverse problems with sparse, indirect, and noisy data.
- Its rates depend on source smoothness and effective dimension, achieving an ℓ¹ reconstruction rate of n^{-r/(1+b-br)} under polynomial spectral decay and a variational source condition.
- The paper links best n-term approximation, sparse approximation spaces, and variational source conditions, then validates the theory for elliptic coefficient recovery, filtered Radon transforms, and direct learning.
Setting and motivation
The paper develops a statistical learning theory for sparse recovery in nonlinear inverse problems. The unknown is a sequence f†∈ℓ1, observed through random, noisy, indirect data
yi=A(f†)(xi)+εi,i=1,…,n,
where A:D(A)∩ℓ1→H is a possibly nonlinear forward operator into a vector-valued reproducing kernel Hilbert space (vv-RKHS) continuously embedded in H=L2(X,μ;Y). The estimator is the ℓ1-regularized empirical risk minimizer
fz,λ=argf∈D(A)∩ℓ1minn1i=1∑n∥A(f)(xi)−yi∥Y2+λ∥f∥ℓ1.
The framework extends the effective-dimension analysis of Caponnetto–De Vito (2607.07468) and the linear statistical inverse learning of Blanchard–Mücke to a Banach-space setting with sparsity-promoting penalties. Compared with prior work on nonlinear Tikhonov regularization in vv-RKHSs [Rastogi et al., 2020], which yields rates n−r/(2r+b+1) under Hilbert-space penalties, the present paper obtains the rate n−r/(1+b−br), where r is a source smoothness index from a variational source condition (VSC) and b is the polynomial decay exponent of the effective dimension yi=A(f†)(xi)+εi,i=1,…,n,0.
Assumptions
The analysis rests on six structural assumptions:
- True solution: the conditional mean yi=A(f†)(xi)+εi,i=1,…,n,1 equals yi=A(f†)(xi)+εi,i=1,…,n,2.
- Sub-exponential noise: a Bernstein-type moment condition with constants yi=A(f†)(xi)+εi,i=1,…,n,3; notably, Gaussian white noise in infinite-dimensional output spaces is explicitly excluded.
- Kernel regularity: uniformly bounded Hilbert–Schmidt norm yi=A(f†)(xi)+εi,i=1,…,n,4 and measurability of kernel inner products.
- Lipschitz continuity: yi=A(f†)(xi)+εi,i=1,…,n,5 injective and Lipschitz in yi=A(f†)(xi)+εi,i=1,…,n,6 with constant yi=A(f†)(xi)+εi,i=1,…,n,7.
- Variational source condition: for all admissible yi=A(f†)(xi)+εi,i=1,…,n,8,
yi=A(f†)(xi)+εi,i=1,…,n,9
with concave index function A:D(A)∩ℓ1→H0; the main results take A:D(A)∩ℓ1→H1.
- Weighted bi-Lipschitz property: two-sided control between A:D(A)∩ℓ1→H2 and A:D(A)∩ℓ1→H3 for weights vanishing at infinity.
The model class A:D(A)∩ℓ1→H4 collects all distributions satisfying these assumptions with A:D(A)∩ℓ1→H5 and polynomial spectral decay.
Consistency
Under Assumptions 1–4 plus weak-to-weak sequential continuity of A:D(A)∩ℓ1→H6, and the parameter choice A:D(A)∩ℓ1→H7 with A:D(A)∩ℓ1→H8, the estimator converges almost surely: A:D(A)∩ℓ1→H9. The proof combines Borel–Cantelli application of high-probability bounds on H=L2(X,μ;Y)0 and H=L2(X,μ;Y)1 (of order H=L2(X,μ;Y)2), almost-sure boundedness of the iterates in H=L2(X,μ;Y)3, weak-* compactness via Banach–Alaoglu–Bourbaki, identification of the limit through injectivity, and norm convergence via Scheffé's lemma. Existence of minimizers holds under continuity/compactness of H=L2(X,μ;Y)4, though uniqueness is not guaranteed for nonlinear H=L2(X,μ;Y)5.
Upper convergence rates
The core estimate separates bias and variance through the Fenchel conjugate H=L2(X,μ;Y)6. With confidence H=L2(X,μ;Y)7,
H=L2(X,μ;Y)8
For H=L2(X,μ;Y)9 and the explicit choice ℓ10, this yields
ℓ11
with high probability. Via an interpolation inequality across weighted sequence spaces ℓ12 with weights ℓ13, the reconstruction bound generalizes to
ℓ14
and this rate holds uniformly over ℓ15 in ℓ16 expectation for all ℓ17 by tail integration. Two special cases are worth noting: for ℓ18 the exponent reduces to ℓ19, while for fz,λ=argf∈D(A)∩ℓ1minn1i=1∑n∥A(f)(xi)−yi∥Y2+λ∥f∥ℓ1.0 it becomes fz,λ=argf∈D(A)∩ℓ1minn1i=1∑n∥A(f)(xi)−yi∥Y2+λ∥f∥ℓ1.1 — independent of fz,λ=argf∈D(A)∩ℓ1minn1i=1∑n∥A(f)(xi)−yi∥Y2+λ∥f∥ℓ1.2, reflecting that the weighted fz,λ=argf∈D(A)∩ℓ1minn1i=1∑n∥A(f)(xi)−yi∥Y2+λ∥f∥ℓ1.3 error is dominated by prediction-norm behavior rather than sparsity.
Minimax lower bounds
Lower bounds are established by an information-theoretic argument: a packing construction over sign vectors (giving fz,λ=argf∈D(A)∩ℓ1minn1i=1∑n∥A(f)(xi)−yi∥Y2+λ∥f∥ℓ1.4 codewords at Hamming distance at least fz,λ=argf∈D(A)∩ℓ1minn1i=1∑n∥A(f)(xi)−yi∥Y2+λ∥f∥ℓ1.5), a family of discrete measures fz,λ=argf∈D(A)∩ℓ1minn1i=1∑n∥A(f)(xi)−yi∥Y2+λ∥f∥ℓ1.6 adapted from Caponnetto–De Vito, KL-divergence estimates controlled by the upper Lipschitz property of fz,λ=argf∈D(A)∩ℓ1minn1i=1∑n∥A(f)(xi)−yi∥Y2+λ∥f∥ℓ1.7, and Fano's inequality. Under a scale condition on the weights fz,λ=argf∈D(A)∩ℓ1minn1i=1∑n∥A(f)(xi)−yi∥Y2+λ∥f∥ℓ1.8 over a block of fz,λ=argf∈D(A)∩ℓ1minn1i=1∑n∥A(f)(xi)−yi∥Y2+λ∥f∥ℓ1.9 indices, any estimator suffers error exceeding n−r/(2r+b+1)0 with probability at least n−r/(2r+b+1)1, where balancing the packing size n−r/(2r+b+1)2 against the KL term gives exactly
n−r/(2r+b+1)3
Consequently no algorithm can beat this rate in n−r/(2r+b+1)4 over n−r/(2r+b+1)5, matching the upper bound. The remark verifying the weight-block condition shows it is compatible with polynomially decaying weights n−r/(2r+b+1)6 provided n−r/(2r+b+1)7.
From approximation spaces to source conditions
A central structural contribution connects sparse approximation theory to the VSC. For the space
n−r/(2r+b+1)8
the paper proves three linked facts:
- Equivalence: n−r/(2r+b+1)9 iff a power-type inequality n−r/(1+b−br)0 holds (a variant of a lemma of Miller–Hohage).
- VSC: under the lower bi-Lipschitz bound, n−r/(1+b−br)1 membership implies the VSC with index n−r/(1+b−br)2; conversely, the VSC implies n−r/(1+b−br)3 membership with n−r/(1+b−br)4.
- Best n−r/(1+b−br)5-term characterization: n−r/(1+b−br)6 iff the best n−r/(1+b−br)7-term approximation error satisfies n−r/(1+b−br)8 in n−r/(1+b−br)9, proved via a layer-cake argument.
Together these give the chain: polynomial decay of r0 r1 r2 membership r3 VSC with r4 r5 the optimal rate. An important consequence is that within this framework the exploitable smoothness index is confined to r6: r7 regularization extracts sparsity efficiently but cannot exploit smoothness beyond this range, unlike r8 regularization. This is a genuine limitation of the method, not merely of the analysis.
Verified examples
Finitely smoothing operators. For composite operators r9 with b0 a wavelet or shearlet synthesis operator and b1 satisfying a two-sided b2 equivalence, both parts of the bi-Lipschitz assumption hold with explicit weights (b3 for wavelets; shifted by b4 for shearlets). Concrete image models then fix b5: piecewise b6 signals in 1D give b7; bounded-variation images give b8 (hence b9, no rate); cartoon-like images under shearlets or curvelets give yi=A(f†)(xi)+εi,i=1,…,n,00, i.e., yi=A(f†)(xi)+εi,i=1,…,n,01.
Elliptic coefficient identification. Recovering a reaction coefficient yi=A(f†)(xi)+εi,i=1,…,n,02 from the solution map of yi=A(f†)(xi)+εi,i=1,…,n,03: the output space yi=A(f†)(xi)+εi,i=1,…,n,04 is an RKHS whose covariance eigenvalues satisfy yi=A(f†)(xi)+εi,i=1,…,n,05, giving yi=A(f†)(xi)+εi,i=1,…,n,06 for yi=A(f†)(xi)+εi,i=1,…,n,07.
Filtered Radon transform. The unfiltered Radon transform fails the framework: its natural RKHS kernel operators yi=A(f†)(xi)+εi,i=1,…,n,08 are not Hilbert–Schmidt (since yi=A(f†)(xi)+εi,i=1,…,n,09 is multiplication by a nonzero line-intersection length, hence noncompact). Smoothing by convolution along the detector variable restores compliance. Via the Fourier slice theorem, the covariance is a pseudodifferential operator with symbol yi=A(f†)(xi)+εi,i=1,…,n,10; Weyl-law asymptotics then give:
| Filter |
Eigenvalue decay |
Effective dimension |
| Boxcar |
yi=A(f†)(xi)+εi,i=1,…,n,11 |
yi=A(f†)(xi)+εi,i=1,…,n,12, so yi=A(f†)(xi)+εi,i=1,…,n,13 |
| Gaussian |
faster than any yi=A(f†)(xi)+εi,i=1,…,n,14 |
slower than any yi=A(f†)(xi)+εi,i=1,…,n,15, any yi=A(f†)(xi)+εi,i=1,…,n,16 |
Combining these with the image models yields concrete rates: for cartoon-like images with boxcar filtering, yi=A(f†)(xi)+εi,i=1,…,n,17 with yi=A(f†)(xi)+εi,i=1,…,n,18; with Gaussian filtering, yi=A(f†)(xi)+εi,i=1,…,n,19 with yi=A(f†)(xi)+εi,i=1,…,n,20. For finitely supported (sparse) coefficients the rates improve to yi=A(f†)(xi)+εi,i=1,…,n,21 (boxcar) and yi=A(f†)(xi)+εi,i=1,…,n,22 (Gaussian). Numerical SVD computations with the ASTRA toolbox strip model confirm a pre-asymptotic polynomial decay slightly steeper than the unfiltered prediction yi=A(f†)(xi)+εi,i=1,…,n,23, with the caveat that sketching-based low-rank SVD may be inaccurate for small singular values.
Direct learning. When yi=A(f†)(xi)+εi,i=1,…,n,24, yi=A(f†)(xi)+εi,i=1,…,n,25 reduces to the synthesis operator, the bi-Lipschitz property holds with equality and constant yi=A(f†)(xi)+εi,i=1,…,n,26, and the framework recovers classical vector-valued kernel ridge regression as a special case — providing a clean consistency check against Caponnetto–De Vito and Lasso-type theory without restricted isometry or incoherence conditions.
Limitations and open questions
Several restrictions are acknowledged explicitly. The sub-exponential noise assumption excludes infinite-dimensional Gaussian white noise. The lower bi-Lipschitz bound rules out infinitely smoothing operators such as the backward heat equation and electrical impedance tomography; extending the theory to logarithmic source conditions or Sobolev-type variational inequalities for such problems remains open. The optimal parameter yi=A(f†)(xi)+εi,i=1,…,n,27 depends on unknown quantities yi=A(f†)(xi)+εi,i=1,…,n,28 and yi=A(f†)(xi)+εi,i=1,…,n,29; while Corollary 4 offers a dual-function characterization amenable to data-driven selection analogous to the discrepancy principle, a rigorous adaptive procedure with finite-sample guarantees is not provided. Finally, whether the optimal rates remain achievable under online or stochastic proximal algorithms, and how the analysis extends to Fréchet-differentiable yi=A(f†)(xi)+εi,i=1,…,n,30 via iteratively reweighted schemes, are questions the paper leaves unanswered.
Conclusion
The paper establishes a complete minimax theory for yi=A(f†)(xi)+εi,i=1,…,n,31-regularized nonlinear statistical inverse learning in vv-RKHSs: almost-sure consistency, high-probability rates yi=A(f†)(xi)+εi,i=1,…,n,32 in interpolated norms, and matching information-theoretic lower bounds over the class yi=A(f†)(xi)+εi,i=1,…,n,33. Its distinctive contribution is the equivalence linking best yi=A(f†)(xi)+εi,i=1,…,n,34-term approximation decay, the spaces yi=A(f†)(xi)+εi,i=1,…,n,35, and variational source conditions, which makes the abstract assumptions verifiable for concrete sparsity models. Applications to elliptic coefficient identification and filtered Radon transforms demonstrate that the theory produces explicit, filter-dependent convergence guarantees for standard tomographic imaging models.