Papers
Topics
Authors
Recent
Search
2000 character limit reached

Scale-Mixture Representations for Isotropic Kernels

Updated 10 January 2026
  • Scale-mixture representations for isotropic kernels express positive-definite kernels as integrals over scale parameters, unifying classical RBFs with multiscale approaches.
  • The framework leverages Bochner’s and Schoenberg’s theorems to provide explicit constructions of reproducing kernel Hilbert spaces with minimal-decomposition norms.
  • Applications range from efficient random Fourier feature sampling in machine learning to multiscale image registration and neural network kernel limits.

A scale-mixture representation for isotropic kernels expresses a positive-definite (PD) kernel as an integral (or sum) over a parametric family of isotropic kernels, typically controlled by a scale parameter. This framework unifies classical descriptions of radial basis functions (RBFs), allows for multiscale modeling, and provides explicit constructions for both the kernel and the associated reproducing kernel Hilbert space (RKHS). Scale mixtures are central to topics ranging from machine learning via random Fourier features to image registration, and connect directly to foundational characterizations by Bochner and Schoenberg.

1. Fundamental Representation and Theoretical Framework

A function k(∥x−y∥)k(\|x-y\|) on Rd×Rd\mathbb{R}^d \times \mathbb{R}^d is called an isotropic kernel if it depends only on the Euclidean distance between xx and yy. Classical results (Bochner, Schoenberg) establish that any continuous, shift-invariant, positive-definite isotropic kernel is a scale mixture of basic kernel “atoms.” Specifically, for a wide class of parameterized kernels φ(r,s)\varphi(r, s) and a finite nonnegative measure μ\mu on [0,∞)[0, \infty):

k(∥x−y∥)=∫0∞φ(∥x−y∥,s)  dμ(s).k(\|x-y\|) = \int_0^\infty \varphi(\|x-y\|, s)\; d\mu(s).

Typical choices include φ(r,s)=exp⁡(−sr2)\varphi(r, s) = \exp(-s r^2) (Gaussian), yielding mixtures of Gaussians, and other forms yielding Matérn or compactly supported kernels (Hotz et al., 2012).

These integrals produce kernels that are positive definite for any nonnegative measure μ\mu, as nonnegative linear combinations or integrals of PD kernels are themselves PD (Bruveris et al., 2011). The kernel Rd×Rd\mathbb{R}^d \times \mathbb{R}^d0 is then the reproducing kernel of the image of a direct-integral Hilbert space of functions parameterized by Rd×Rd\mathbb{R}^d \times \mathbb{R}^d1, with the RKHS norm given by a minimal decomposition property:

Rd×Rd\mathbb{R}^d \times \mathbb{R}^d2

2. Scale-Mixture Representations: Classical and Generalized Forms

Numerous kernel families admit scale-mixture representations as specific cases of the above framework:

  • Rational Quadratic: Rd×Rd\mathbb{R}^d \times \mathbb{R}^d3 is a scale mixture of Gaussians, with a mixing measure corresponding to an inverse-gamma distribution (Hotz et al., 2012).
  • Matérn Kernel: The Matérn family is expressible as

Rd×Rd\mathbb{R}^d \times \mathbb{R}^d4

with Rd×Rd\mathbb{R}^d \times \mathbb{R}^d5 derived from the Bessel function representation (Hotz et al., 2012, Langrené et al., 2024).

  • Generalized Cauchy and Exponential Power: For generalized Cauchy Rd×Rd\mathbb{R}^d \times \mathbb{R}^d6 and exponential-power Rd×Rd\mathbb{R}^d \times \mathbb{R}^d7, the spectral (Bochner) densities also admit scale-mixture forms as integrals over Gaussians with a properly chosen density Rd×Rd\mathbb{R}^d \times \mathbb{R}^d8 (Langrené et al., 2024).

Schoenberg’s theorem provides the most general characterization: any Rd×Rd\mathbb{R}^d \times \mathbb{R}^d9-invariant kernel on xx0 can be written as an infinite series of radial kernels weighted by normalized Gegenbauer polynomials (zonal polynomials), with strictly positive definite kernels characterized by conditions on the radial coefficients xx1 (Benning et al., 27 Jun 2025). The scale-mixture form in the stationary case recovers classical “Gaussian mixtures,” with the kernel written as

xx2

where xx3 is a dimension-dependent Bessel-type function.

3. Associated Hilbert Spaces and Minimal-Decomposition Norms

The scale-mixture construction induces a RKHS via a direct integral. For a discrete mixture with xx4 scales and kernels xx5, the corresponding RKHS xx6 is equipped with a norm defined by

xx7

where each xx8 is the RKHS associated to xx9 (Bruveris et al., 2011). The reproducing kernel of yy0 is then the sum of the individual kernels:

yy1

In the continuous case, the direct-integral space consists of functions yy2 such that yy3, and the kernel is given by the integral over yy4 (Hotz et al., 2012). The RKHS norm is again a minimal-decomposition norm.

This formalism generalizes to include Mercer expansions, integral-operator kernels, and various compactly supported kernels (e.g., Wendland kernels), providing a unified approach to many classical and modern kernel classes.

4. Applications in Learning and Geometry

Scale-mixture representations have significant practical and theoretical applications:

  • Random Fourier Features (RFF): For shift-invariant isotropic kernels, the spectral density can be written as a Gaussian-scale mixture, enabling efficient sampling for RFF construction. Instead of sampling from a fixed Gaussian, one samples a variance parameter yy5 from the mixing law and then samples yy6 from yy7. This enables RFF approximations for a wide range of kernels, including Matérn, generalized Cauchy, exponential power, Beta, Kummer, and Tricomi families (Langrené et al., 2024).
  • Kernel Ridge Regression, SVM, Gaussian Processes: Scale mixtures yield closed-form expressions for kernels suitable for low-rank approximation and efficient learning (Langrené et al., 2024).
  • Image Registration and LDDMM: In large-deformation diffeomorphic metric mapping (LDDMM), mixed-kernel RKHSs correspond to multiscale models for diffeomorphic flows. The equivalence between variational formulations using a single sum-kernel and joint multiscale optimization is established via Lagrange multipliers and relates to an iterated semidirect-product decomposition of diffeomorphism groups (Bruveris et al., 2011).
  • Inverse Problems and Integral Operators: Regularization strategies can be implemented in large direct-integral RKHSs, then pulled back to finite-rank expansions (Hotz et al., 2012).

5. Spectral and Structural Characterizations

The scale-mixture view is underpinned by spectral theory. By Bochner’s theorem, any continuous, shift-invariant, PD kernel yy8 is the Fourier transform of a finite nonnegative measure yy9. Schoenberg’s extension ensures that for isotropic kernels, φ(r,s)\varphi(r, s)0 is the Laplace transform of a positive measure over φ(r,s)\varphi(r, s)1 in φ(r,s)\varphi(r, s)2, implying complete monotonicity (Langrené et al., 2024, Hotz et al., 2012, Benning et al., 27 Jun 2025).

In spectral mixture representations:

φ(r,s)\varphi(r, s)3

where φ(r,s)\varphi(r, s)4 is the mixture density determined by the kernel, allowing constructive sampling and explicit feature map construction for a large set of RBF kernels (Langrené et al., 2024).

6. Connections to General Isotropic Kernels and Neural Network Limits

The most general φ(r,s)\varphi(r, s)5-invariant (isotropic) kernels are parametrized not only by the distance but also by the dot product, reducing to scale mixtures in the stationary case and to Taylor expansions in the dot-product case. Continuous, φ(r,s)\varphi(r, s)6-invariant, PD kernels admit expansions of the form:

φ(r,s)\varphi(r, s)7

where φ(r,s)\varphi(r, s)8 are scale-mixture coefficients and φ(r,s)\varphi(r, s)9 are normalized Gegenbauer polynomials. The stationary case corresponds to μ\mu0 as explicit scale mixtures over radial functions, while dot-product kernels have μ\mu1 (Benning et al., 27 Jun 2025).

Infinite-width limits of neural networks yield μ\mu2-invariant kernels in this class, with explicit Gegenbauer or Hermite expansions determined by the activation function (Benning et al., 27 Jun 2025).

7. Examples and Practical Construction

A comparative summary of prototypical isotropic kernel scale-mixtures:

Kernel Class Mixture Formulation Mixing Measure / Density
Gaussian μ\mu3 Dirac at μ\mu4
Rational Quadratic μ\mu5 Inverse-gamma over scale μ\mu6
Matérn μ\mu7 μ\mu8
Generalized Cauchy μ\mu9 [0,∞)[0, \infty)0
Wendland (compact support) [0,∞)[0, \infty)1 Any finite positive [0,∞)[0, \infty)2

Explicit algorithms for random feature sampling (RFF) involve drawing the scale parameter from the kernel’s mixture density, then sampling a Gaussian direction. No additional complexity is introduced compared to the basic RFF approach; the only change is in the law of the scale parameter (Langrené et al., 2024).

References

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Scale-Mixture Representations for Isotropic Kernels.