---
title: Posterior Contraction Rates
url: https://www.emergentmind.com/topics/posterior-contraction-rates
type: topic
---

# Posterior Contraction Rates

Posterior contraction rates quantify the asymptotic speed at which the Bayesian posterior distribution concentrates around the true data-generating parameter or function as sample size increases. They are central to Bayesian nonparametric theory, providing a frequentist benchmarking of Bayesian learning and establishing certainty quantification for Bayesian inference in both parametric and high- or infinite-dimensional models. The mathematical machinery for posterior contraction rates ("PCRs"—*Editor's term*) has evolved to include information- and transportation-based distances, particularly the Wasserstein metrics, and adapts to dominated, non-dominated, parametric, nonparametric, linear, nonlinear, and even computationally discrete settings.

## 1. Formal Definition and Conceptual Framework

Given a statistical model $\{P_\theta : \theta \in \Theta\}$ with parameter space $\Theta$ (which may be infinite-dimensional), observations $X^{(n)} = (X_1, ..., X_n)$, a prior $\Pi$ on $\Theta$, and sample size $n$, the posterior $\Pi_n(\cdot|X^{(n)})$ quantifies the conditional belief on $\theta$ after observing data. A sequence $\varepsilon_n \to 0$ is a posterior contraction rate at $\theta_0$ if, for every $M_n \to \infty$,
\[
\Pi_n(\{\theta: d(\theta, \theta_0) > M_n \varepsilon_n \} \mid X^{(n)}) \to 0 \quad \text{in $P_{\theta_0}$-probability,}
\]
where $d$ is an appropriate metric, such as $L^2$, Hellinger, or Wasserstein [2011.14425][2203.10754][2201.12225]. This quantifies that posterior mass contracts around $\theta_0$ at rate $\varepsilon_n$.

In nonparametric and high-dimensional models, $d$ is often taken as a norm on function spaces or a probability metric (e.g., Wasserstein-$p$ distance):
\[
W_p(\mu, \nu) = \left( \inf_{\gamma \in \Gamma(\mu, \nu)} \int \|x-y\|^p \, d\gamma(x,y) \right)^{1/p}
\]
with $\Gamma(\mu, \nu)$ the set of couplings of $\mu,\nu$ [2203.10754].

## 2. Posterior Contraction in Dominated and Non-dominated Models

- **Dominated models:** Classical PCRs exploit Bayes' formula; the posterior can be written as
  \[
  \Pi_n(d\theta\,|\,X^{(n)}) \propto \prod_{i=1}^n f(X_i|\theta) \Pi(d\theta)
  \]
  where $f(\cdot|\theta)$ is a density with respect to a dominating measure [2011.14425].
  
- **Non-dominated models:** In many Bayesian nonparametric constructions (e.g., Dirichlet process mixtures, normalized random measures), no single dominating measure $f(\cdot|\theta)$ exists. In this case, one works with posterior kernels $T_n$ arising from disintegration (de Finetti decomposition), and PCRs are defined using, for example, the $W_p$ metric on $\mathscr P(\mathscr P(\mathcal X))$ [2201.12225]. The formal PCR is then:
  \[
  \varepsilon_n := \mathbb{E}_{p_0}\left[ W_p\left( T_n(\cdot|X^{(n)}), \delta_{p_0} \right) \right]
  \]
  which quantifies the expected $p$-Wasserstein distance of the random posterior to the Dirac mass at $p_0$.

## 3. Methodologies for Establishing Posterior Contraction Rates

The general theoretical strategy for establishing PCRs involves verifying a combination of:
- **Entropy (complexity) conditions:** Covering numbers or metric entropy of the effective parameter space at scale $\varepsilon_n$, e.g. $\log N(\varepsilon, \Theta, d) \lesssim n \varepsilon^2$ [1507.07412][2105.07410][2605.11652].
- **Prior-mass (small ball) conditions:** Sufficient prior mass near the truth, i.e. $\Pi(\theta: K(P_{\theta_0}, P_{\theta}) < \varepsilon_n^2 ) \ge e^{-Cn\varepsilon_n^2}$ for Kullback–Leibler neighborhood [1507.07412][2201.12225].
- **Testing and concentration inequalities:** Existence of exponentially powerful tests for testing $H_0 : \theta = \theta_0$ against alternatives at distance $>\varepsilon_n$ [1703.08358][2601.17805][2411.06981].
- **Sieve construction:** In nonparametrics or under weak regularity, contractive "sieves" (compact/entropy-controlled sets) inside which entropy and testing conditions hold, with negligible prior mass outside [2201.12225].

Recent developments leverage the **Wasserstein metric** and the dynamic Benamou–Brenier formulation [2011.14425][2203.10754], providing two novelties:
- **Avoidance of sieves** in certain contexts — PCRs can be established in strong metrics via the dynamic formulation [2011.14425][2203.10754][2411.06981].
- **Direct connection to empirical process rates and Poincaré inequalities:** PCRs are linked to rates of convergence in the empirical measure (Glivenko–Cantelli), Sanov's large deviation principle in Wasserstein distance, and the estimation of weighted Poincaré–Wirtinger constants [2203.10754].

## 4. Quantitative Examples in Parametric, Nonparametric, and Infinite-dimensional Models

### Regular parametric models
For regular finite-dimensional models, the posterior contracts at the optimal parametric rate:
\[
\varepsilon_n = O(n^{-1/2})
\]
for any prior with positive, continuous density at $\theta_0$ [2203.10754].

### Nonparametric Dirichlet-Laplace mixtures
For the model $p_G = f * G$ with Laplace mixing and a Dirichlet process prior, Gao and van der Vaart [1507.07412] show:
\[
W_1(G, G_0) = O(n^{-1/8} (\log n)^{5/8})
\]
for the mixing distribution, and
\[
h(p_G, p_{G_0}) = O((\log n / n)^{3/8})
\]
for the density, matching minimax lower bounds up to log factors.

### Deep Gaussian process priors and compositional classes
Under deep GP priors for functions expressed as compositions $f = h_q \circ \ldots \circ h_0$ with layerwise Hölder or Besov regularity, the contraction rate in $L^2$ is
\[
\varepsilon_n(\eta^*) = \max_{i=0,\ldots,q^*} n^{-\frac{\beta_i^*\alpha_i^*}{2 \beta_i^*\alpha_i^* + t_i^*}}
\]
achieving minimax adaptivity to unknown compositional structure [2105.07410].

### Besov-Laplace priors in white noise
For functions $f_0 \in B^{\beta}_{1,1}$ ($\beta > d/2$) and smoothness-matching Besov-Laplace priors, the strong posterior contraction rate in the Sobolev norm is
\[
\|\cdot\|_{H^s}\text{ at rate }\xi_n = n^{-(\beta-s)/(2\beta+d)}, \qquad 0 \le s < \beta - d/2
\]
matching minimax lower bounds [2411.06981].

### High-dimensional and nonparametric sparsity
For regression with spike-and-slab or shrinkage priors, the rate is
\[
\epsilon_n = \sqrt{ s_0 \log p / n }
\]
where $s_0$ is true sparsity, and $p \gg n$ [1904.04417][2405.01206].

### Non-dominated nonparametric models
For Dirichlet process ({\small DP}) or normalized Gamma process priors in non-dominated settings (i.e., not absolutely continuous in the parameter), the contraction rate in $W_p$ is [2201.12225]:
\[
\varepsilon_n \lesssim E_{n, GC} + O(n^{-(a-ds)/(d+p)})
\]
where $E_{n,GC}$ is the expected Wasserstein convergence of the empirical measure, and the second term depends on the metric entropy and prior concentration.

## 5. Analytical Structure: Wasserstein Dynamics, Glivenko–Cantelli, and Poincaré Constants

A central methodological innovation is the combination of:
- **Local Lipschitz-continuity of the posterior:** If $W_p(\Pi^*(\cdot|b), \Pi^*(\cdot|b')) \leq L_n \|b-b'\|_B$ for sufficient statistics or empirical measures $b, b'$, then PCRs can be controlled by data fluctuation rates [2011.14425][2203.10754].
- **Dynamic Benamou–Brenier formulation of $W_p$:** The infimum over transport plans is reframed as an infimum over absolutely continuous probability curve flows $(\mu_t)_{t\in[0,1]}$ solving the continuity equation with minimal kinetic energy [2203.10754].
- **Laplace method and weighted Poincaré–Wirtinger constants:** Asymptotics in both finite and infinite dimension rely on Laplace expansions and spectral gap estimates for

Source: https://www.emergentmind.com/topics/posterior-contraction-rates