Papers
Topics
Authors
Recent
Search
2000 character limit reached

Multi-Bandwidth Gaussian Processes

Updated 10 July 2026
  • Multi-bandwidth Gaussian processes are models that leverage multiple characteristic scales—such as coordinate-specific rescaling, kernel mixtures, and spectral components—to capture complex anisotropic and local variations.
  • They overcome the rigidity of a single global bandwidth by enabling adaptive, nonstationary, and multi-output representations, which improve both flexibility and estimation accuracy.
  • These methods demonstrate practical benefits in achieving minimax optimal rates and computational scalability in high-dimensional, multi-resolution, and spatially varying contexts.

Multi-bandwidth Gaussian processes are Gaussian-process models in which dependence is governed by more than one characteristic scale. In the cited literature, this appears in several technically distinct forms: dimension-specific rescalings of a homogeneous GP, additive mixtures of kernels with distinct bandwidths, multi-output kernels in which latent covariance components carry separate hyperparameters, spectral constructions with multiple spectral widths or compact passbands, and nonstationary models in which the effective bandwidth varies across the input domain. This suggests that “multi-bandwidth Gaussian process” is not a single canonical construction, but a family resemblance among models that relax the single global bandwidth assumption (Bhattacharya et al., 2011, Feinberg et al., 2017, Parra et al., 2017, Beckman et al., 2022).

1. Meanings of bandwidth heterogeneity

In these works, “bandwidth” refers variously to a coordinate-specific rescaling, a latent-component length-scale, a spectral variance, a compact frequency passband, or a location-dependent anisotropy matrix. The common motivation is that a single stationary covariance with one global scale is often too rigid for anisotropic functions, multi-output dependence, spectral multi-scale structure, or local regime changes.

Interpretation Mechanism Representative papers
Coordinate-specific bandwidths WtA=W(A1t1,,Adtd)W_t^{\mathbf{A}} = W(A_1 t_1,\dots,A_d t_d) (Bhattacharya et al., 2011)
Additive scale mixing pβpKp\sum_p \beta_p K_p or qbij(q)kq(xz)\sum_q b_{ij}^{(q)} k_q(x-z) (Archambeau et al., 2011, Feinberg et al., 2017)
Spectral multi-bandwidth Gaussian or rectangular spectral components with distinct Σq\Sigma_q, Δ\Delta, or ξ0\xi_0 (Parra et al., 2017, Tobar, 2019, Simpson et al., 2021)
Input-dependent bandwidths location-dependent Σ(x)\Sigma(x) or local GP experts (Beckman et al., 2022, Zhang et al., 2019)
Multiresolution scale variation nested partitions or multiple-resolution basis expansions (Fox et al., 2012, Katzfuss et al., 2017)

A recurrent distinction is between globally coherent kernels and local-mixture constructions. Some models define one valid covariance K(x,x)K(x,x') over the whole domain, while others induce multi-bandwidth behavior through mixtures of local stationary GPs or through multiresolution approximations. This distinction matters because it changes the interpretation of uncertainty, cross-region dependence, and identifiability (Zhang et al., 2019, Katzfuss et al., 2017).

2. Dimension-specific and anisotropic bandwidths

The most direct use of the phrase appears in anisotropic function estimation, where a homogeneous Gaussian field WW is rescaled by a vector of positive bandwidths,

WtA=W(A1t1,,Adtd),W_t^{\mathbf{A}} = W(A_1 t_1,\dots,A_d t_d),

rather than by a single scalar bandwidth (Bhattacharya et al., 2011). In this formulation, the pβpKp\sum_p \beta_p K_p0-th coordinate has its own inverse bandwidth pβpKp\sum_p \beta_p K_p1, and the prior on pβpKp\sum_p \beta_p K_p2 is constructed hierarchically through a Dirichlet-distributed simplex variable pβpKp\sum_p \beta_p K_p3 and coordinatewise heavy-tailed scale priors. The statistical objective is simultaneous adaptation to unknown anisotropic smoothness and, in a related formulation, to unknown intrinsic dimension (Bhattacharya et al., 2011).

The anisotropic smoothness benchmark is governed by

pβpKp\sum_p \beta_p K_p4

or by

pβpKp\sum_p \beta_p K_p5

when only coordinates in an active subset pβpKp\sum_p \beta_p K_p6 matter (Bhattacharya et al., 2011). Under the multibandwidth prior, the posterior contracts at

pβpKp\sum_p \beta_p K_p7

or

pβpKp\sum_p \beta_p K_p8

which is minimax optimal up to logarithmic factors (Bhattacharya et al., 2011). The same work also shows that a homogeneous Gaussian process with a single bandwidth can be polynomially sub-optimal in anisotropic settings, so the move from one scale to coordinate-specific scales is not merely parametric refinement but a change in asymptotic adaptivity (Bhattacharya et al., 2011).

This anisotropic viewpoint is the cleanest setting in which “multi-bandwidth GP” means exactly what the phrase suggests: multiple bandwidths are attached directly to coordinate directions. Later constructions generalize the idea from coordinatewise anisotropy to latent components, output pairs, spectral bands, or spatial locations.

3. Additive scale mixtures and multi-output coregionalization

A second major route to multi-bandwidth behavior is additive mixing across kernels or latent processes. In Bayesian multiple kernel learning, the effective covariance is a positive linear combination

pβpKp\sum_p \beta_p K_p9

with Gaussian process priors on latent functions and generalized inverse Gaussian priors on the associated scales (Archambeau et al., 2011). If the dictionary qbij(q)kq(xz)\sum_q b_{ij}^{(q)} k_q(x-z)0 contains kernels with different bandwidths or length-scales, the model becomes a direct Bayesian multi-bandwidth GP: different scales are activated or suppressed by the learned weights. The function-space interpretation is a sum of independent latent GPs at different scales, with heavy-tailed process priors inducing adaptive sparsity among components (Archambeau et al., 2011).

In multi-output regression, the same additive principle appears in the linear model of coregionalization,

qbij(q)kq(xz)\sum_q b_{ij}^{(q)} k_q(x-z)1

Each latent kernel qbij(q)kq(xz)\sum_q b_{ij}^{(q)} k_q(x-z)2 can have its own hyperparameters, so if qbij(q)kq(xz)\sum_q b_{ij}^{(q)} k_q(x-z)3 is an RBF or Matérn kernel, each component may carry a distinct length-scale (Feinberg et al., 2017). The covariance between outputs qbij(q)kq(xz)\sum_q b_{ij}^{(q)} k_q(x-z)4 and qbij(q)kq(xz)\sum_q b_{ij}^{(q)} k_q(x-z)5 is then an additive mixture of latent components, each with its own smoothness scale and its own coregionalization weight qbij(q)kq(xz)\sum_q b_{ij}^{(q)} k_q(x-z)6. In this sense, large linear multi-output GP learning is multi-bandwidth in a component-wise LMC sense rather than through a bespoke bandwidth-sharing kernel (Feinberg et al., 2017).

The output coupling matrices are parameterized as

qbij(q)kq(xz)\sum_q b_{ij}^{(q)} k_q(x-z)7

and on a shared interpolation grid the covariance has the form

qbij(q)kq(xz)\sum_q b_{ij}^{(q)} k_q(x-z)8

This algebra is the basis for the LLGP scalability results: Toeplitz structure in one dimension, block-Toeplitz with Toeplitz blocks in low dimensions, and efficient matrix-vector multiplication without forming the dense covariance explicitly (Feinberg et al., 2017). A conceptual subtlety is that even when every qbij(q)kq(xz)\sum_q b_{ij}^{(q)} k_q(x-z)9 is stationary on the input space, the full multi-output kernel on Σq\Sigma_q0 is generally not stationary. Accordingly, LLGP’s main novelty is computational scalability for standard LMC models that already support multiple latent bandwidths (Feinberg et al., 2017).

4. Spectral mixtures, compact passbands, and frequency-domain multi-bandwidth models

A third and especially influential meaning of multi-bandwidth arises in the spectral domain. For stationary scalar kernels, Bochner’s theorem gives

Σq\Sigma_q1

so one can design the covariance through the power spectral density Σq\Sigma_q2 (Parra et al., 2017). The spectral mixture kernel chooses Σq\Sigma_q3 as a Gaussian mixture and yields

Σq\Sigma_q4

Here Σq\Sigma_q5 controls frequency and Σq\Sigma_q6 controls spectral width, which by Fourier duality corresponds to characteristic length-scale or bandwidth in the covariance domain (Parra et al., 2017). Multi-bandwidth behavior is therefore intrinsic: each mixture component contributes a different spectral band and a different effective scale.

The multi-output extension in MOSM uses Cramér’s theorem and factorizes the matrix-valued cross-spectral density as Σq\Sigma_q7, with output-specific amplitudes, frequencies, spectral covariances, delays, and phases (Parra et al., 2017). The resulting closed-form covariance for output pair Σq\Sigma_q8,

Σq\Sigma_q9

shows explicitly how cross-bandwidth, cross-frequency, delay, and phase interact (Parra et al., 2017). This construction is directly multi-frequency, multi-bandwidth, and multi-output.

MOCSM modifies this spectral-mixture program by using the convolution theorem to design cross-channel dependencies through cross convolution of time- and phase-delayed components in the spectral domain (Chen et al., 2018). It is noteworthy for two reasons. First, it reduces to the standard SM kernel in the single-channel case. Second, it is presented as avoiding the undesirable scale effects that MOSM can exhibit when spectral densities of different channels are either very close or very far from each other in the frequency domain (Chen et al., 2018).

A distinct line replaces Gaussian spectral components with compactly supported bands. The sinc-kernel construction takes a symmetric rectangular PSD,

Δ\Delta0

and yields

Δ\Delta1

The centered case gives Δ\Delta2, and the generalized sinc kernel admits a frequency-varying weighting Δ\Delta3 as well as an approximation by a sum of narrow sinc components (Tobar, 2019). The paper explicitly supports single-band compact spectra and suggests multi-band constructions through this generalized frequency-domain viewpoint (Tobar, 2019).

The Minecraft kernel pushes the compact-support idea into the multi-output setting by replacing Gaussian spectral atoms with finite-bandwidth rectangular step functions (Simpson et al., 2021). Its stated motivation is that Gaussian multi-output spectral mixtures have a critical blind spot: when marginal spectral densities differ in shape, the cross-covariance is not reproducible across the full range of permitted correlations except in the special case of identical spectral shapes. By using block components of finite bandwidth, the proposed family is presented as the first multi-output generalisation of the spectral mixture kernel that can approximate any stationary multi-output kernel to arbitrary precision (Simpson et al., 2021). In this literature, “multi-bandwidth” therefore sometimes means not many Gaussian widths, but many finite passbands.

Finally, MOHSM extends the spectral-mixture program to the nonstationary case by replacing stationary power spectral densities with harmonizable generalized cross-spectral densities. Its kernel combines a stationary-like oscillatory term in Δ\Delta4 with a localization term in Δ\Delta5, so different spectral widths and characteristic scales can be active in different regions of the input space (Altamirano et al., 2022).

5. Local bandwidth variation and nonstationary constructions

A fourth interpretation of multi-bandwidth Gaussian processes is explicitly nonstationary: the bandwidth itself varies across the domain. In the anisotropic Paciorek–Schervish framework, this role is played by a location-dependent positive definite matrix Δ\Delta6, which controls local range, orientation, and anisotropy (Beckman et al., 2022). In the scalable implementation studied for sea surface temperature anomalies, Δ\Delta7 is represented by a normalized RBF expansion,

Δ\Delta8

with positive definiteness enforced through Cholesky factors (Beckman et al., 2022). This is an explicit spatially adaptive bandwidth GP: the effective length-scale is a function of location rather than a global hyperparameter.

Not all nonstationary multi-bandwidth models take the form of one global kernel. Sequential GP-MOE replaces a single stationary GP by a Dirichlet-process mixture of local GPs, each with its own kernel hyperparameters Δ\Delta9, where for the RBF expert kernel larger ξ0\xi_00 means shorter effective length-scale (Zhang et al., 2019). The model is therefore multi-bandwidth in a local-mixture sense: different regions of the input space can be governed by different local stationary kernels. At the same time, the work explicitly distinguishes this from a single closed-form nonstationary covariance with smoothly varying length-scale; nonstationarity is induced through latent expert assignments and local kernel learning (Zhang et al., 2019).

The multiresolution GP offers another nonstationary route. It hierarchically couples smooth GPs over a nested binary partition tree, with top levels capturing long-range dependence and lower levels modeling local deviations and abrupt changes (Fox et al., 2012). The level-specific kernels are squared exponential,

ξ0\xi_01

with the typical “fractal” choice

ξ0\xi_02

As partition cells shrink, ξ0\xi_03 grows, so finer levels operate at finer effective bandwidths (Fox et al., 2012). The induced covariance is additive across scales but gated by shared partition membership, which makes the model both multiscale and nonstationary (Fox et al., 2012).

Functional Gaussian processes provide a spectral analogue of this local-regime idea. In NS-FGP, component-specific spectral densities ξ0\xi_04 are selected by a location-dependent stick-breaking process, but all components share the same spectral Gaussian vector ξ0\xi_05, so the model preserves a jointly Gaussian structure instead of breaking cross-component dependence (Duan et al., 2015). This produces local stationarity clusters with different effective spectral widths and anisotropic range parameters, especially in spatio-temporal settings (Duan et al., 2015).

6. Approximation, scalability, and recurring limitations

Because multi-bandwidth models often introduce many kernel parameters or many latent components, scalable linear algebra is a central theme. LLGP adapts structured kernel interpolation to multi-output LMC models by placing a common interpolation grid across outputs and exploiting Toeplitz or BTTB structure in the input-space matrices ξ0\xi_06. The approximation

ξ0\xi_07

preserves the LMC sum over latent kernels while enabling fast matrix-vector multiplication and gradient-based hyperparameter learning without explicitly forming the dense covariance (Feinberg et al., 2017). Its main contribution is therefore computational rather than conceptual: it scales standard multi-output component-wise multi-bandwidth models (Feinberg et al., 2017).

A different approximation strategy is the multi-resolution approximation framework, which writes the approximate process as

ξ0\xi_08

with basis functions across resolutions. The induced covariance

ξ0\xi_09

captures coarse-to-fine dependence, and the framework includes both taper and block versions (Katzfuss et al., 2017). This is multi-bandwidth in an approximation-centric sense: low resolutions capture broad structure and higher resolutions capture localized residual structure. It is not primarily an explicit multi-bandwidth kernel parameterization, but it can approximately represent any spatial covariance structure while achieving quasilinear scaling in Σ(x)\Sigma(x)0 under suitable choices of Σ(x)\Sigma(x)1 and Σ(x)\Sigma(x)2 (Katzfuss et al., 2017).

Adaptive finite-element-type decompositions sharpen the distinction between bandwidth-adaptive and true multi-bandwidth models. In the GPI construction, a parent squared-exponential GP with global inverse bandwidth Σ(x)\Sigma(x)3 is interpolated on a regular grid, and a prior on both Σ(x)\Sigma(x)4 and the number of basis functions Σ(x)\Sigma(x)5 yields posterior contraction

Σ(x)\Sigma(x)6

over Hölder classes (Kim et al., 29 May 2025). By contrast, the SPDE/Matérn finite-element approximation with fixed smoothness is shown to be suboptimal for sufficiently smooth truths (Kim et al., 29 May 2025). The paper is therefore relevant to multi-bandwidth GPs chiefly as a study of global bandwidth randomization plus random basis resolution, not as a model with several concurrent bandwidths (Kim et al., 29 May 2025).

Fast multi-group GP factor models add another nuance. mDLAG gives each latent factor Σ(x)\Sigma(x)7 its own GP timescale Σ(x)\Sigma(x)8, so different latent dimensions can have different bandwidths, but a given latent uses the same Σ(x)\Sigma(x)9 across all groups and groups differ through delays K(x,x)K(x,x')0 rather than through separate group-specific bandwidths (Gokcen et al., 2024). The paper’s two approximations reduce runtime from cubic to linear in both trial length and group number, with the frequency-domain approach consistently providing the greatest runtime benefits with the fewest trade-offs in statistical performance (Gokcen et al., 2024). At the same time, the inducing-variable approximation can oversmooth fast latents when the inducing grid is too coarse, and the frequency-domain approximation can underestimate timescales and delays for short trials (Gokcen et al., 2024).

A recurring misconception is to treat all of these constructions as interchangeable. They are not. Anisotropic rescaling attaches bandwidths to coordinates (Bhattacharya et al., 2011); additive kernel mixtures attach them to latent components or basis kernels (Archambeau et al., 2011, Feinberg et al., 2017); spectral models attach them to Gaussian or rectangular frequency bands (Parra et al., 2017, Tobar, 2019, Simpson et al., 2021); and nonstationary models attach them to locations, partitions, or local experts (Beckman et al., 2022, Fox et al., 2012, Zhang et al., 2019). A plausible implication is that model choice should be matched to the locus of heterogeneity—coordinates, latent components, output pairs, frequencies, or regions of the input space—rather than to the phrase “multi-bandwidth” alone.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Multi-Bandwidth Gaussian Processes.