- The paper presents a full asymptotic formula for learning curves, delineating over-regularized, under-regularized, and interpolation regimes in spectral algorithms.
- It establishes precise conditions under which benign overfitting occurs, showing that interpolation can yield minimax-optimal risk in high-dimensional settings.
- The analysis bridges kernel methods with Gaussian sequence models, offering practical insights into algorithm tuning and model selection in overparameterized regimes.
Learning Curves and Benign Overfitting of Spectral Algorithms in Large Dimensions
Introduction
This paper presents a rigorous asymptotic analysis of learning curves for spectral algorithms—particularly kernel ridge regression (KRR) and related analytic filters—in high-dimensional nonparametric regression settings where the sample size n and dimension d are coupled as n≍dγ for some γ>0. The analysis is primarily focused on inner-product kernels defined over the high-dimensional sphere Sd−1, leveraging their tractable spectral properties, but is later extended to more general kernel families under mild spectral and hypercontractivity conditions.
The core innovation is a full characterization of the learning curve as a function of the regularization parameter λ, unifying and extending previous results that were restricted to only the optimally tuned or interpolation (λ→0) regimes. Central to the analysis is the identification of three statistically distinct regimes—over-regularized, under-regularized, and interpolation—for the excess risk, and an explicit sharp threshold for the emergence of benign overfitting. The implications touch both practical model selection as well as the broader theory of generalization in overparameterized regimes, with immediate connections to the benign overfitting observed in "lazy" (NTK) neural training.
Theoretical Framework and Notation
The standard nonparametric regression model is considered, with n i.i.d. observations (xi,yi) where xi∈Sd−1, and d0, with d1 zero-mean noise. The target is the regression function d2, assumed to belong to a power space d3 associated with the reproducing kernel Hilbert space (RKHS) d4 of the kernel. The parameter d5 quantifies function smoothness relative to d6.
Spectral algorithms are parameterized by analytic filter functions with qualification d7, covering KRR (d8), kernel gradient flow (d9), and higher-qualification methods. Regularization is set as n≍dγ0 for n≍dγ1, with the limit n≍dγ2 corresponding to interpolation. Inner-product kernels benefit from a block-diagonal spectral decomposition into spherical harmonics, facilitating explicit control over the learning dynamics.
Main Results: Learning Curve Trichotomy
A central contribution is the derivation of a full asymptotic formula for the excess risk along the regularization path:
n≍dγ3
where n≍dγ4 (with n≍dγ5, n≍dγ6), and n≍dγ7.
This encapsulates three distinct regimes as the regularization is tuned:
- Over-regularized: n≍dγ8 is large (n≍dγ9); bias dominates, and excess risk decreases with decreasing γ>00 (increasing γ>01).
- Under-regularized: γ>02; variance increases and begins to dominate beyond a critical γ>03, leading to plateaus.
- Interpolation: γ>04 (i.e., γ>05); risk stabilizes, no further substantial benefit from lowering regularization.



Figure 1: A graphical representation of the learning curves of large dimensional spectral algorithms with γ>06 (blue) and γ>07 (orange) obtained in Theorem 3.1.
The theory reveals that the learning curve does not exhibit a simple classical U-shape, but instead is composed of plateaus, flat regions, and sharp transitions.
Benign Overfitting: Sharp Threshold and Regime Persistence
A principal finding is the identification of the exact conditions for benign overfitting—where interpolation or near-interpolation yields minimax-optimal excess risk. Define the minimax rate as
γ>08
Benign overfitting occurs if and only if γ>09, where
Sd−10
Crucially, benign overfitting is not restricted to the interpolation point (Sd−11) but persists throughout the entire under-regularized regime (Sd−12). Thus, once the regularization becomes sufficiently weak, further reductions do not degrade generalization so long as Sd−13.

Figure 2: Another graphical representation of the learning curve of large dimensional spectral algorithms with Sd−14 (blue), Sd−15 (green), and Sd−16 (orange) obtained in Theorem 3.1.
If Sd−17, excess risk increases for too-small Sd−18, reflecting the classical overfitting phenomenon. This boundary provides a rigorous formalization of the phase transition for benign vs. non-benign overfitting in high dimensions.
Saturation and Algorithm Qualification
The paper also recovers and geometrically explains the "saturation effect," whereby KRR and analogous qualification-1 spectral methods become suboptimal for learning highly smooth signals (Sd−19), while higher-qualification methods (e.g., kernel gradient flow) can achieve minimax rates up to the information-theoretic limit. The trichotomy of the learning curves reveals that, for algorithms with finite qualification, the minimum attainable risk is constrained by a slower descent early in the regularization path, and the global learning minimum occurs for greater λ0 than in infinite-qualification methods.
Experimental Validation
Experiments are performed with synthetic data and both NTK and RBF kernels, verifying the predicted exponents and transitions. Detailed convergence rates as functions of λ1, λ2, λ3, and λ4 match the theoretical learning curves derived.

Figure 3: Type 1 experiments with parameters λ5 and regularization λ6; left: NTK, right: RBF kernels.
Figure 4: Type 1 experiments with λ7; left: NTK, right: RBF kernels.
Figure 5: Comparison of the experimental and theoretical convergence rates for Type 2 experiments with λ8; NTK (left), RBF (right).
Extensions to General Kernels and Domains
The analysis is extended beyond the spherical case to much broader classes of kernels and domains, under spectral scaling (blockwise decay) and eigenfunction hypercontractivity. This encompasses kernels such as RBFs on λ9, random feature kernels on the hypercube, and others, supporting the universality of the identified phenomena in large-scale kernel learning.
Equivalence with Sequence Model
Another significant implication is the formal equivalence between sufficiently regularized kernel regression and an associated Gaussian sequence model. For λ→00, the sequence model yields identical learning rates to kernel regression, reinforcing the value of sequence model as a theoretically tractable surrogate in high-dimensional analysis.
Practical and Theoretical Implications
- Model Selection: The work resolves the behavior of kernel methods across the entire regularization path, providing clear guidance for selection and tuning in high-dimensional regimes. Practitioners can identify minimax-optimal risk by targeting the appropriate regime rather than solely focusing on interpolation.
- Theoretical Generalization: The findings clarify the circumstances under which benign overfitting is theoretically justified, revealing the limits of overparameterization and guiding expectations for wider classes of learning algorithms, including lazy neural networks.
- Algorithm Design: The insights into qualification and saturation inform the design of better regularized and early-stopped spectral algorithms, especially for signals of varying smoothness.
Conclusion
This work offers an exhaustive, sharp characterization of the learning curves and benign overfitting phenomena for spectral algorithms in modern high-dimensional, large-sample regimes. The explicitly determined thresholds and rates unify and extend prior results, clarify the scope and limits of benign overfitting, and align with empirical behaviors observed in neural network models. The extension to broader kernel and domain families, as well as the equivalence to sequence models, positions the theory as a foundational tool for understanding generalization and algorithm selection in overparameterized learning.