Papers
Topics
Authors
Recent
Search
2000 character limit reached

Learning Curves and Benign Overfitting of Spectral Algorithms in Large Dimensions

Published 25 Apr 2026 in stat.ML, cs.LG, and math.ST | (2604.23212v1)

Abstract: Existing large-dimensional theory for spectral algorithms resolves either the optimally tuned point or the interpolation limit, but leaves the under-regularized regime unexplored. We study the learning curve and benign overfitting of spectral algorithms in the large-dimensional setting where the sample size and dimension are of comparable order, i.e., nd<sup>γn \asymp d<sup>γ for some $γ&gt;0$. We first consider inner-product kernels on the sphere S<sup>d1\mathbb{S}<sup>{d-1} and establish a sharp asymptotic characterization of the excess risk across the full regularization path under various source conditions s0s \geq 0, where ss measures the relative smoothness of the regression function. Our results reveal that the learning curve is not simply U-shaped but instead consists of three distinct regimes: over-regularized, under-regularized, and interpolation regimes. This characterization allows us to fully capture the benign overfitting phenomenon, demonstrating that benign overfitting arises consistently across both the under-regularized and interpolation regimes whenever ss is positive but no larger than a critical threshold. We further show that, in the sufficiently regularized regime, the kernel learning curve is recovered by an associated sequence model. Finally, we extend the learning-curve analysis to large-dimensional KRR for a class of kernels on general domains in R<sup>d\mathbb{R}<sup>d whose low-degree eigenspaces satisfy spectral-scaling and hyper-contractivity conditions.

Summary

  • The paper presents a full asymptotic formula for learning curves, delineating over-regularized, under-regularized, and interpolation regimes in spectral algorithms.
  • It establishes precise conditions under which benign overfitting occurs, showing that interpolation can yield minimax-optimal risk in high-dimensional settings.
  • The analysis bridges kernel methods with Gaussian sequence models, offering practical insights into algorithm tuning and model selection in overparameterized regimes.

Learning Curves and Benign Overfitting of Spectral Algorithms in Large Dimensions

Introduction

This paper presents a rigorous asymptotic analysis of learning curves for spectral algorithms—particularly kernel ridge regression (KRR) and related analytic filters—in high-dimensional nonparametric regression settings where the sample size nn and dimension dd are coupled as ndγn \asymp d^\gamma for some γ>0\gamma>0. The analysis is primarily focused on inner-product kernels defined over the high-dimensional sphere Sd1\mathbb{S}^{d-1}, leveraging their tractable spectral properties, but is later extended to more general kernel families under mild spectral and hypercontractivity conditions.

The core innovation is a full characterization of the learning curve as a function of the regularization parameter λ\lambda, unifying and extending previous results that were restricted to only the optimally tuned or interpolation (λ0\lambda\to 0) regimes. Central to the analysis is the identification of three statistically distinct regimes—over-regularized, under-regularized, and interpolation—for the excess risk, and an explicit sharp threshold for the emergence of benign overfitting. The implications touch both practical model selection as well as the broader theory of generalization in overparameterized regimes, with immediate connections to the benign overfitting observed in "lazy" (NTK) neural training.

Theoretical Framework and Notation

The standard nonparametric regression model is considered, with nn i.i.d. observations (xi,yi)(x_i, y_i) where xiSd1x_i \in \mathbb{S}^{d-1}, and dd0, with dd1 zero-mean noise. The target is the regression function dd2, assumed to belong to a power space dd3 associated with the reproducing kernel Hilbert space (RKHS) dd4 of the kernel. The parameter dd5 quantifies function smoothness relative to dd6.

Spectral algorithms are parameterized by analytic filter functions with qualification dd7, covering KRR (dd8), kernel gradient flow (dd9), and higher-qualification methods. Regularization is set as ndγn \asymp d^\gamma0 for ndγn \asymp d^\gamma1, with the limit ndγn \asymp d^\gamma2 corresponding to interpolation. Inner-product kernels benefit from a block-diagonal spectral decomposition into spherical harmonics, facilitating explicit control over the learning dynamics.

Main Results: Learning Curve Trichotomy

A central contribution is the derivation of a full asymptotic formula for the excess risk along the regularization path:

ndγn \asymp d^\gamma3

where ndγn \asymp d^\gamma4 (with ndγn \asymp d^\gamma5, ndγn \asymp d^\gamma6), and ndγn \asymp d^\gamma7.

This encapsulates three distinct regimes as the regularization is tuned:

  • Over-regularized: ndγn \asymp d^\gamma8 is large (ndγn \asymp d^\gamma9); bias dominates, and excess risk decreases with decreasing γ>0\gamma>00 (increasing γ>0\gamma>01).
  • Under-regularized: γ>0\gamma>02; variance increases and begins to dominate beyond a critical γ>0\gamma>03, leading to plateaus.
  • Interpolation: γ>0\gamma>04 (i.e., γ>0\gamma>05); risk stabilizes, no further substantial benefit from lowering regularization. Figure 1

Figure 1

Figure 1

Figure 1

Figure 1: A graphical representation of the learning curves of large dimensional spectral algorithms with γ>0\gamma>06 (blue) and γ>0\gamma>07 (orange) obtained in Theorem 3.1.

The theory reveals that the learning curve does not exhibit a simple classical U-shape, but instead is composed of plateaus, flat regions, and sharp transitions.

Benign Overfitting: Sharp Threshold and Regime Persistence

A principal finding is the identification of the exact conditions for benign overfitting—where interpolation or near-interpolation yields minimax-optimal excess risk. Define the minimax rate as

γ>0\gamma>08

Benign overfitting occurs if and only if γ>0\gamma>09, where

Sd1\mathbb{S}^{d-1}0

Crucially, benign overfitting is not restricted to the interpolation point (Sd1\mathbb{S}^{d-1}1) but persists throughout the entire under-regularized regime (Sd1\mathbb{S}^{d-1}2). Thus, once the regularization becomes sufficiently weak, further reductions do not degrade generalization so long as Sd1\mathbb{S}^{d-1}3. Figure 2

Figure 2

Figure 2: Another graphical representation of the learning curve of large dimensional spectral algorithms with Sd1\mathbb{S}^{d-1}4 (blue), Sd1\mathbb{S}^{d-1}5 (green), and Sd1\mathbb{S}^{d-1}6 (orange) obtained in Theorem 3.1.

If Sd1\mathbb{S}^{d-1}7, excess risk increases for too-small Sd1\mathbb{S}^{d-1}8, reflecting the classical overfitting phenomenon. This boundary provides a rigorous formalization of the phase transition for benign vs. non-benign overfitting in high dimensions.

Saturation and Algorithm Qualification

The paper also recovers and geometrically explains the "saturation effect," whereby KRR and analogous qualification-1 spectral methods become suboptimal for learning highly smooth signals (Sd1\mathbb{S}^{d-1}9), while higher-qualification methods (e.g., kernel gradient flow) can achieve minimax rates up to the information-theoretic limit. The trichotomy of the learning curves reveals that, for algorithms with finite qualification, the minimum attainable risk is constrained by a slower descent early in the regularization path, and the global learning minimum occurs for greater λ\lambda0 than in infinite-qualification methods.

Experimental Validation

Experiments are performed with synthetic data and both NTK and RBF kernels, verifying the predicted exponents and transitions. Detailed convergence rates as functions of λ\lambda1, λ\lambda2, λ\lambda3, and λ\lambda4 match the theoretical learning curves derived. Figure 3

Figure 3

Figure 3: Type 1 experiments with parameters λ\lambda5 and regularization λ\lambda6; left: NTK, right: RBF kernels.

Figure 4

Figure 4

Figure 4: Type 1 experiments with λ\lambda7; left: NTK, right: RBF kernels.

Figure 5

Figure 5

Figure 5: Comparison of the experimental and theoretical convergence rates for Type 2 experiments with λ\lambda8; NTK (left), RBF (right).

Extensions to General Kernels and Domains

The analysis is extended beyond the spherical case to much broader classes of kernels and domains, under spectral scaling (blockwise decay) and eigenfunction hypercontractivity. This encompasses kernels such as RBFs on λ\lambda9, random feature kernels on the hypercube, and others, supporting the universality of the identified phenomena in large-scale kernel learning.

Equivalence with Sequence Model

Another significant implication is the formal equivalence between sufficiently regularized kernel regression and an associated Gaussian sequence model. For λ0\lambda\to 00, the sequence model yields identical learning rates to kernel regression, reinforcing the value of sequence model as a theoretically tractable surrogate in high-dimensional analysis.

Practical and Theoretical Implications

  • Model Selection: The work resolves the behavior of kernel methods across the entire regularization path, providing clear guidance for selection and tuning in high-dimensional regimes. Practitioners can identify minimax-optimal risk by targeting the appropriate regime rather than solely focusing on interpolation.
  • Theoretical Generalization: The findings clarify the circumstances under which benign overfitting is theoretically justified, revealing the limits of overparameterization and guiding expectations for wider classes of learning algorithms, including lazy neural networks.
  • Algorithm Design: The insights into qualification and saturation inform the design of better regularized and early-stopped spectral algorithms, especially for signals of varying smoothness.

Conclusion

This work offers an exhaustive, sharp characterization of the learning curves and benign overfitting phenomena for spectral algorithms in modern high-dimensional, large-sample regimes. The explicitly determined thresholds and rates unify and extend prior results, clarify the scope and limits of benign overfitting, and align with empirical behaviors observed in neural network models. The extension to broader kernel and domain families, as well as the equivalence to sequence models, positions the theory as a foundational tool for understanding generalization and algorithm selection in overparameterized learning.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.