Papers
Topics
Authors
Recent
Search
2000 character limit reached

Polynomial Convergence for Gaussian KRR

Updated 18 August 2025
  • The paper establishes explicit polynomial convergence rates for Gaussian KRR by linking bias–variance decompositions with polynomial eigenvalue decay.
  • It details both L2 and uniform error bounds, emphasizing the roles of smoothness, source conditions, and saturation effects in estimation accuracy.
  • Empirical validations and scalable algorithms illustrate that fixed-width Gaussian KRR can efficiently handle high-dimensional nonparametric regression.

Polynomial convergence rates for Gaussian kernel ridge regression (KRR) describe the rate at which the KRR estimator’s prediction error decays as a function of the sample size under polynomial eigenvalue decay scenarios. For the Gaussian kernel, which is infinitely smooth, recent results have precisely quantified these rates for both L2L^{2} and uniform norms. This encompasses classical bias-variance decompositions, saturation effects, alignment phenomena, and the interplay with spectral and statistical characteristics. The topic is central to theoretical nonparametric regression, distributed algorithms, scalable solvers, and the statistical learning theory of kernel methods.

1. Theoretical Foundations and Frameworks

Kernel ridge regression estimates a target function f0f_{0} by minimizing the regularized empirical risk functional in a reproducing kernel Hilbert space (RKHS) HK\mathcal{H}_{K} generated by a positive-definite kernel KK. For data {(xi,yi)}i=1n\{(x_{i}, y_{i})\}_{i=1}^{n}, the estimator

f^λ=argminfHK1ni=1n(yif(xi))2+λfHK2\hat{f}_{\lambda} = \arg\min_{f \in \mathcal{H}_{K}} \frac{1}{n} \sum_{i=1}^{n} (y_{i} - f(x_{i}))^{2} + \lambda \|f\|_{\mathcal{H}_{K}}^{2}

is characterized by first-order optimality conditions and Mercer decompositions. In the context of Gaussian kernels (and more generally radial kernels), the eigenvalues {μj}\{\mu_{j}\} of the associated integral operator play a fundamental role, with rates often depending on their polynomial decay μjjt\mu_{j} \sim j^{-t} (Zhang et al., 2013).

Error analysis is typically based on a bias–variance decomposition, with the squared L2L^{2} error controlled by the regularization parameter λ\lambda, the spectral decay, and the smoothness of the target function: f0f_{0}0 where the effective dimension is f0f_{0}1 (Zhang et al., 2013).

2. Polynomial Rates for Gaussian Kernel Ridge Regression

Recent advances have established explicit polynomial convergence rates under fixed Gaussian kernel hyperparameters, closing historical gaps in both f0f_{0}2 and uniform norms (Dommel et al., 15 Aug 2025). Under the assumption that f0f_{0}3 on f0f_{0}4 with f0f_{0}5, the estimator obeys: f0f_{0}6 for a regularization sequence f0f_{0}7 and with probability at least f0f_{0}8 (Dommel et al., 15 Aug 2025). This quantifies the polynomial rate f0f_{0}9 for smooth HK\mathcal{H}_{K}0.

For uniform convergence: HK\mathcal{H}_{K}1 with HK\mathcal{H}_{K}2, under stronger smoothness conditions (HK\mathcal{H}_{K}3) and additional decay assumptions on the expansion coefficients (Dommel et al., 15 Aug 2025).

These bounds provide theoretical justification for using fixed-width Gaussian KRR in nonparametric regression, correcting prior beliefs that only sub-polynomial or logarithmic rates were possible for fixed bandwidths.

3. Role of Smoothness, Source Condition, and Saturation

The convergence rate for Gaussian KRR is sensitive to the interplay between the kernel eigenvalue decay and the smoothness of the target function. The source condition is typically formulated as HK\mathcal{H}_{K}4, an interpolation space with smoothness HK\mathcal{H}_{K}5; the polynomial rate becomes HK\mathcal{H}_{K}6. For HK\mathcal{H}_{K}7, the estimator is minimax optimal (matches the lower bound); for HK\mathcal{H}_{K}8, the rate “saturates” at HK\mathcal{H}_{K}9, reflecting the fact that further smoothness does not yield better rates—a phenomenon known as saturation (Long et al., 2024).

In high-dimensional settings, with sample size KK0, one observes periodic plateau behavior and multiple descent phenomena: the error rate remains constant over intervals of KK1, then drops sharply as KK2 increases (Zhang et al., 2024), elucidating non-monotonic phases in the learning curve. This analysis unifies results from several previous works by allowing interpolation parameter KK3 to vary freely.

4. Connections to Gaussian Process Regression and Capacity-Dependent Analysis

The optimal convergence rates for Gaussian KRR align closely with those of Gaussian process (GP) regression, especially when the imposed kernel is smoother than the underlying true function (Wang et al., 2021). GP sample paths typically have smoothness KK4, and if the KRR kernel has smoothness KK5, then both regression procedures achieve the minimax rate KK6, with the Gaussian kernel’s effective dimension modulating the capacity and the learning rate.

Capacity-dependent analysis addresses the scenario when the true regression function does not lie in the RKHS. The rates then depend explicitly on a regularity source parameter KK7 and an effective dimension exponent KK8, yielding (Lin et al., 2018): KK9 where {(xi,yi)}i=1n\{(x_{i}, y_{i})\}_{i=1}^{n}0 optimally balances bias and variance.

5. Computational and Algorithmic Aspects

Polynomial rates interact with computational complexity in scalable KRR solvers. Partition-based approaches decompose the estimation error into approximation, bias, variance, and regularization components; distributed algorithms attain minimax optimal rates as long as partitioning preserves effective dimensionality (Zhang et al., 2013, Tandon et al., 2016). Sparse approximations (Nyström, SVGP) enable polynomial rates with dramatically reduced cost:

  • SE kernel: {(xi,yi)}i=1n\{(x_{i}, y_{i})\}_{i=1}^{n}1, rate {(xi,yi)}i=1n\{(x_{i}, y_{i})\}_{i=1}^{n}2
  • Matérn kernel: {(xi,yi)}i=1n\{(x_{i}, y_{i})\}_{i=1}^{n}3, rate {(xi,yi)}i=1n\{(x_{i}, y_{i})\}_{i=1}^{n}4 (Vakili et al., 2022)

Randomized preconditioners (RPCholesky, KRILL) decouple convergence rates from the size/condition of the kernel matrix, ensuring rapid, condition-number-independent CG convergence when the spectrum decays polynomially (Díaz et al., 2023). Linear convergence of full KRR with scalable solvers such as ASkotch is achieved via Nyström preconditioners of rank comparable to the effective dimension (Rathore et al., 2024).

6. Alignment, Truncation, and Transient Phenomena

Alignment between the target function and the kernel spectrum can induce faster polynomial rates, particularly under spectral truncation (TKRR). If the target’s expansion coefficients decay as {(xi,yi)}i=1n\{(x_{i}, y_{i})\}_{i=1}^{n}5 and the kernel eigenvalues as {(xi,yi)}i=1n\{(x_{i}, y_{i})\}_{i=1}^{n}6, alignment parameter {(xi,yi)}i=1n\{(x_{i}, y_{i})\}_{i=1}^{n}7 produces accelerated rate {(xi,yi)}i=1n\{(x_{i}, y_{i})\}_{i=1}^{n}8 for TKRR—surpassing the standard KRR rate {(xi,yi)}i=1n\{(x_{i}, y_{i})\}_{i=1}^{n}9 in “over-aligned” regimes (Amini et al., 2022). Truncation also induces multiple descent and non-monotonic learning curve phenomena, especially when the target’s spectrum is bandlimited.

7. Empirical Validation and Implications

Extensive numerical experiments verify theoretical polynomial rates for both noiseless and noisy KRR estimators under varying smoothness, dimension, and kernel choices. The empirical risk matches polynomial bounds, confirming minimax optimality for f^λ=argminfHK1ni=1n(yif(xi))2+λfHK2\hat{f}_{\lambda} = \arg\min_{f \in \mathcal{H}_{K}} \frac{1}{n} \sum_{i=1}^{n} (y_{i} - f(x_{i}))^{2} + \lambda \|f\|_{\mathcal{H}_{K}}^{2}0 and showing saturation for f^λ=argminfHK1ni=1n(yif(xi))2+λfHK2\hat{f}_{\lambda} = \arg\min_{f \in \mathcal{H}_{K}} \frac{1}{n} \sum_{i=1}^{n} (y_{i} - f(x_{i}))^{2} + \lambda \|f\|_{\mathcal{H}_{K}}^{2}1 (Long et al., 2024, Saber et al., 2023). Distributed and partitioned estimators have demonstrated computational superiority while retaining optimal rates (Tandon et al., 2016).

These results justify the use of fixed Gaussian kernel ridge regression in large-scale, high-dimensional regression, under mild smoothness and noise conditions, and provide concrete guidance for selecting regularization and approximation parameters to achieve predictable polynomial error decay.

References

Definition Search Book Streamline Icon: https://streamlinehq.com
References (16)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Polynomial Convergence Rates for Gaussian Kernel Ridge Regression (KRR).