Papers
Topics
Authors
Recent
Search
2000 character limit reached

Gaussian Sieve Priors

Updated 18 November 2025
  • Gaussian sieve priors are hierarchical Bayesian priors that express infinite-dimensional functions using truncated orthonormal basis expansions.
  • They enable adaptive nonparametric inference by selecting a variable truncation level, achieving near-minimax global L2 contraction rates.
  • However, they exhibit suboptimal performance for pointwise and semi-parametric loss functions due to insufficient regularization of intermediate-frequency components.

Gaussian sieve priors are hierarchical Bayesian priors designed for adaptive nonparametric inference, especially in settings where the underlying signal or function admits a sparse or truncated orthonormal expansion. In such models, the key feature is to express the infinite-dimensional parameter (such as a function or spectral density) in a suitable basis, and then encode prior information via a variable truncation level and, conditionally, independent Gaussian priors on the expansion coefficients. The formulation enables dimension adaptation and allows contraction rates that (up to logarithmic factors) closely track minimax optimality for certain global loss functions. Gaussian sieve priors have been rigorously analyzed for models including the Gaussian white noise model and semi-parametric Gaussian time series, revealing both their strengths in global L2L_2 adaptation and their limitations under pointwise or semi-parametric loss functions (Arbel et al., 2012, Kruijer et al., 2012).

1. Construction of Gaussian Sieve Priors

In the Gaussian white noise model

dXn(t)=f0(t)dt+n1/2dW(t),t[0,1],dX^n(t)=f_0(t)dt+n^{-1/2}dW(t), \quad t\in[0,1],

where f0f_0 is the unknown function and WW is standard Brownian motion, the function is expressed in an orthonormal basis {φj}j1\{\varphi_j\}_{j\geq 1} as f0(t)=j=1θ0jφj(t)f_0(t)=\sum_{j=1}^\infty \theta_{0j}\varphi_j(t). The observations in the basis are

Xjn=01φj(t)dXn(t)=θ0j+n1/2ξj,ξjN(0,1).X_j^n = \int_0^1 \varphi_j(t) dX^n(t) = \theta_{0j} + n^{-1/2}\xi_j, \quad \xi_j\sim N(0,1).

The Gaussian sieve prior places a hierarchical structure on the coefficients θ=(θj)j1\theta=(\theta_j)_{j\geq 1}: Π(dθ)=k=1π(k)Πk(dθ),\Pi(d\theta) = \sum_{k=1}^\infty \pi(k)\Pi_k(d\theta), where

  • π(k)\pi(k) is a prior over truncation level dXn(t)=f0(t)dt+n1/2dW(t),t[0,1],dX^n(t)=f_0(t)dt+n^{-1/2}dW(t), \quad t\in[0,1],0,
  • Conditionally on dXn(t)=f0(t)dt+n1/2dW(t),t[0,1],dX^n(t)=f_0(t)dt+n^{-1/2}dW(t), \quad t\in[0,1],1, dXn(t)=f0(t)dt+n1/2dW(t),t[0,1],dX^n(t)=f_0(t)dt+n^{-1/2}dW(t), \quad t\in[0,1],2 for dXn(t)=f0(t)dt+n1/2dW(t),t[0,1],dX^n(t)=f_0(t)dt+n^{-1/2}dW(t), \quad t\in[0,1],3, and dXn(t)=f0(t)dt+n1/2dW(t),t[0,1],dX^n(t)=f_0(t)dt+n^{-1/2}dW(t), \quad t\in[0,1],4 for dXn(t)=f0(t)dt+n1/2dW(t),t[0,1],dX^n(t)=f_0(t)dt+n^{-1/2}dW(t), \quad t\in[0,1],5.

Canonical choices are dXn(t)=f0(t)dt+n1/2dW(t),t[0,1],dX^n(t)=f_0(t)dt+n^{-1/2}dW(t), \quad t\in[0,1],6 a Poisson distribution with parameter dXn(t)=f0(t)dt+n1/2dW(t),t[0,1],dX^n(t)=f_0(t)dt+n^{-1/2}dW(t), \quad t\in[0,1],7, and dXn(t)=f0(t)dt+n1/2dW(t),t[0,1],dX^n(t)=f_0(t)dt+n^{-1/2}dW(t), \quad t\in[0,1],8 for dXn(t)=f0(t)dt+n1/2dW(t),t[0,1],dX^n(t)=f_0(t)dt+n^{-1/2}dW(t), \quad t\in[0,1],9, f0f_00 (Arbel et al., 2012). The prior is thus a random mixture over finite-dimensional Gaussians, which induces dimension reduction and penalizes complexity via the decay of f0f_01 (typically f0f_02 for constants f0f_03).

In semi-parametric time series models, e.g., the FEXP model for Gaussian long-memory series,

f0f_04

a sieve prior is formulated by independently assigning:

  • f0f_05 a (fixed) density supported in f0f_06,
  • f0f_07 either fixed at a deterministic rate in f0f_08 or random with a Poisson/geometric prior,
  • f0f_09 a distribution supported on a Sobolev ball of smoothness WW0 (Kruijer et al., 2012).

2. Posterior Contraction Rates: WW1 and Global Loss

Under mild regularity assumptions, Gaussian sieve priors yield adaptive minimax-optimal posterior contraction rates (up to log-factors) for global WW2 (or WW3) loss over appropriately defined Sobolev-type parameter spaces.

Specifically, for the Sobolev ball

WW4

the minimax WW5-estimation rate is WW6. The Gaussian sieve prior achieves

WW7

in the sense that, for any WW8 and WW9 sufficiently large,

{φj}j1\{\varphi_j\}_{j\geq 1}0

as {φj}j1\{\varphi_j\}_{j\geq 1}1 [(Arbel et al., 2012), Theorem 3.4, Proposition 4.1, Section 5.1]. The Bayes risk associated with the posterior mean also achieves {φj}j1\{\varphi_j\}_{j\geq 1}2.

The key mechanism underlying these results is a balance of approximation error (controlled by the truncation {φj}j1\{\varphi_j\}_{j\geq 1}3 and the prior scales {φj}j1\{\varphi_j\}_{j\geq 1}4) and stochastic error (arising from the noise level and the prior’s effective sample size). The proof utilizes decisive prior-mass bounds over Kullback–Leibler neighborhoods, metric entropy estimates, and non-asymptotic testing inequalities.

3. Adaptation and Suboptimality for Other Losses

While Gaussian sieve priors provide sharp global {φj}j1\{\varphi_j\}_{j\geq 1}5 adaptation, their behavior under other loss functions can be markedly different. For pointwise (local) risk

{φj}j1\{\varphi_j\}_{j\geq 1}6

where {φj}j1\{\varphi_j\}_{j\geq 1}7, the minimax rate over {φj}j1\{\varphi_j\}_{j\geq 1}8 is {φj}j1\{\varphi_j\}_{j\geq 1}9. Under the Gaussian sieve prior, the attained pointwise risk decays only as

f0(t)=j=1θ0jφj(t)f_0(t)=\sum_{j=1}^\infty \theta_{0j}\varphi_j(t)0

which is slower than the minimax rate by a polynomial factor in f0(t)=j=1θ0jφj(t)f_0(t)=\sum_{j=1}^\infty \theta_{0j}\varphi_j(t)1 [(Arbel et al., 2012), Proposition 5.3]. The dominant error arises from intermediate-frequency coefficients that are insufficiently regularized by the sieve prior, causing excess local variance.

A similar phenomenon appears in semi-parametric estimation of long-memory parameters in time series. With random truncation priors (Poisson or geometric) on the expansion length f0(t)=j=1θ0jφj(t)f_0(t)=\sum_{j=1}^\infty \theta_{0j}\varphi_j(t)2, the contraction rate of the long-memory parameter f0(t)=j=1θ0jφj(t)f_0(t)=\sum_{j=1}^\infty \theta_{0j}\varphi_j(t)3 is

f0(t)=j=1θ0jφj(t)f_0(t)=\sum_{j=1}^\infty \theta_{0j}\varphi_j(t)4

whereas the minimax rate is f0(t)=j=1θ0jφj(t)f_0(t)=\sum_{j=1}^\infty \theta_{0j}\varphi_j(t)5. Only when f0(t)=j=1θ0jφj(t)f_0(t)=\sum_{j=1}^\infty \theta_{0j}\varphi_j(t)6 is tuned deterministically to scale as f0(t)=j=1θ0jφj(t)f_0(t)=\sum_{j=1}^\infty \theta_{0j}\varphi_j(t)7 (rather than assigned a prior) does the sieve prior attain the nearly optimal rate [(Kruijer et al., 2012), Theorems 3.2–3.4].

4. Key Technical Conditions and Proof Structure

Contraction theorems for sieve priors rest on an overview of several analytic conditions:

  • KL-approximation (A1): Existence of low-dimensional truncations that approximate the true model in Kullback–Leibler divergence.
  • Reverse KL–f0(t)=j=1θ0jφj(t)f_0(t)=\sum_{j=1}^\infty \theta_{0j}\varphi_j(t)8 control (A2): Uniform control of KL divergence in terms of f0(t)=j=1θ0jφj(t)f_0(t)=\sum_{j=1}^\infty \theta_{0j}\varphi_j(t)9-distance around the truncation.
  • Metric entropy and covering (A3): Ability to cover sieved parameter sets in Xjn=01φj(t)dXn(t)=θ0j+n1/2ξj,ξjN(0,1).X_j^n = \int_0^1 \varphi_j(t) dX^n(t) = \theta_{0j} + n^{-1/2}\xi_j, \quad \xi_j\sim N(0,1).0-distance by Xjn=01φj(t)dXn(t)=θ0j+n1/2ξj,ξjN(0,1).X_j^n = \int_0^1 \varphi_j(t) dX^n(t) = \theta_{0j} + n^{-1/2}\xi_j, \quad \xi_j\sim N(0,1).1-balls, which facilitates the construction of exponentially powerful tests.
  • Testing (A4): Construction of tests with exponentially decaying type I and II errors for hypotheses separated in Xjn=01φj(t)dXn(t)=θ0j+n1/2ξj,ξjN(0,1).X_j^n = \int_0^1 \varphi_j(t) dX^n(t) = \theta_{0j} + n^{-1/2}\xi_j, \quad \xi_j\sim N(0,1).2.
  • Prior tails and scales (A5): Appropriate decay of Xjn=01φj(t)dXn(t)=θ0j+n1/2ξj,ξjN(0,1).X_j^n = \int_0^1 \varphi_j(t) dX^n(t) = \theta_{0j} + n^{-1/2}\xi_j, \quad \xi_j\sim N(0,1).3, sufficient mass for scales Xjn=01φj(t)dXn(t)=θ0j+n1/2ξj,ξjN(0,1).X_j^n = \int_0^1 \varphi_j(t) dX^n(t) = \theta_{0j} + n^{-1/2}\xi_j, \quad \xi_j\sim N(0,1).4, and tail regularity for the conditional prior densities.

The proof proceeds via a testing–prior-mass approach. The numerator of the posterior probability for "bad" sets (where the contraction fails) is controlled by a union bound over tests, while the denominator is lower-bounded by the prior mass of suitable KL neighborhoods. Contributions from large and small Xjn=01φj(t)dXn(t)=θ0j+n1/2ξj,ξjN(0,1).X_j^n = \int_0^1 \varphi_j(t) dX^n(t) = \theta_{0j} + n^{-1/2}\xi_j, \quad \xi_j\sim N(0,1).5 are handled via tail bounds on Xjn=01φj(t)dXn(t)=θ0j+n1/2ξj,ξjN(0,1).X_j^n = \int_0^1 \varphi_j(t) dX^n(t) = \theta_{0j} + n^{-1/2}\xi_j, \quad \xi_j\sim N(0,1).6 (Arbel et al., 2012).

5. Implications for Adaptive Bayesian Estimation

Gaussian sieve priors illustrate both the strengths and limitations of hierarchical Bayesian adaptivity. For global loss functions such as integrated squared error, the prior achieves adaptive minimax rates over a large class of smoothness spaces (e.g., Sobolev balls). The adaptation occurs via the random selection (or deterministic choice) of the truncation level Xjn=01φj(t)dXn(t)=θ0j+n1/2ξj,ξjN(0,1).X_j^n = \int_0^1 \varphi_j(t) dX^n(t) = \theta_{0j} + n^{-1/2}\xi_j, \quad \xi_j\sim N(0,1).7, which balances bias and variance automatically.

However, for more localized or semi-parametric functionals (e.g., pointwise function estimation, long-memory parameter estimation), full Bayesian adaptation via a sieve prior leads to a trade-off in contraction rate. The failure to attain optimal local rates is due to a fundamental mismatch between the global regularization induced by the sieve structure and the localized risk structure of the problem (Arbel et al., 2012, Kruijer et al., 2012).

This behavior underscores the necessity of careful prior design or tuning for objectives that go beyond global estimation: for instance, fixing the sieve truncation to match the minimax-optimal dimension achieves nearly optimal rates for long-memory parameters in time series, whereas data-driven or fully random truncation does not.

6. Comparison with Frequentist and Other Bayesian Approaches

Analysis reveals that the convergence properties of Gaussian sieve priors often parallel those of frequentist sieve estimators, particularly in nonparametric and semi-parametric models. For example, periodogram-based estimators for long-memory time series achieve the minimax rate Xjn=01φj(t)dXn(t)=θ0j+n1/2ξj,ξjN(0,1).X_j^n = \int_0^1 \varphi_j(t) dX^n(t) = \theta_{0j} + n^{-1/2}\xi_j, \quad \xi_j\sim N(0,1).8, which matches the contraction rate of the sieve prior with deterministic Xjn=01φj(t)dXn(t)=θ0j+n1/2ξj,ξjN(0,1).X_j^n = \int_0^1 \varphi_j(t) dX^n(t) = \theta_{0j} + n^{-1/2}\xi_j, \quad \xi_j\sim N(0,1).9 (up to log-factors) (Kruijer et al., 2012).

In contrast to fully Bayesian methods that randomize θ=(θj)j1\theta=(\theta_j)_{j\geq 1}0 with heavy-tailed priors (facilitating automatic adaptation over the entire parameter space), deterministic or empirically tuned sieves can deliver sharper convergence rates for specific functionals. This reflects an inherent trade-off: full Bayesian adaptation excels for function estimation in global metrics, but sacrifices efficiency for certain semi-parametric objectives.

7. Summary Table: Sieve Prior Contraction Rates

Problem/Class Sieve Prior Type Achieved Rate (up to logs) Minimax/Optimal Rate
Global θ=(θj)j1\theta=(\theta_j)_{j\geq 1}1 (Sobolev, white noise) Poisson θ=(θj)j1\theta=(\theta_j)_{j\geq 1}2, θ=(θj)j1\theta=(\theta_j)_{j\geq 1}3 θ=(θj)j1\theta=(\theta_j)_{j\geq 1}4 θ=(θj)j1\theta=(\theta_j)_{j\geq 1}5
Pointwise (Sobolev, white noise) Same θ=(θj)j1\theta=(\theta_j)_{j\geq 1}6 θ=(θj)j1\theta=(\theta_j)_{j\geq 1}7
Long-memory θ=(θj)j1\theta=(\theta_j)_{j\geq 1}8, FEXP (semi-parametric) θ=(θj)j1\theta=(\theta_j)_{j\geq 1}9 (deterministic) Π(dθ)=k=1π(k)Πk(dθ),\Pi(d\theta) = \sum_{k=1}^\infty \pi(k)\Pi_k(d\theta),0 Π(dθ)=k=1π(k)Πk(dθ),\Pi(d\theta) = \sum_{k=1}^\infty \pi(k)\Pi_k(d\theta),1
Long-memory Π(dθ)=k=1π(k)Πk(dθ),\Pi(d\theta) = \sum_{k=1}^\infty \pi(k)\Pi_k(d\theta),2, FEXP Poisson/geometric Π(dθ)=k=1π(k)Πk(dθ),\Pi(d\theta) = \sum_{k=1}^\infty \pi(k)\Pi_k(d\theta),3 Π(dθ)=k=1π(k)Πk(dθ),\Pi(d\theta) = \sum_{k=1}^\infty \pi(k)\Pi_k(d\theta),4 Π(dθ)=k=1π(k)Πk(dθ),\Pi(d\theta) = \sum_{k=1}^\infty \pi(k)\Pi_k(d\theta),5

The contraction properties confirm that Gaussian sieve priors are robust tools for adaptive nonparametric Bayes inference in high-dimensional and infinite-dimensional settings, but their performance must be evaluated in light of the specific estimation criterion of interest.

References: (Arbel et al., 2012, Kruijer et al., 2012).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (2)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Gaussian Sieve Priors.