Papers
Topics
Authors
Recent
Search
2000 character limit reached

Bayesian Spike-and-Slab Framework

Updated 1 February 2026
  • Bayesian Spike-and-Slab Framework is a hierarchical model that combines a sharp 'spike' at zero with a diffuse 'slab' to represent coefficients for sparse estimation.
  • It enables effective variable selection and uncertainty quantification in high-dimensional regression by inducing a multimodal posterior over sparse supports.
  • The framework leverages restricted isometry properties and efficient rejection sampling techniques to ensure accurate, scalable posterior sampling with formal computational guarantees.

A Bayesian Spike-and-Slab Framework is a hierarchical probabilistic modeling approach for sparse estimation, variable selection, and uncertainty quantification in high-dimensional inference problems. The framework combines a "spike" component (typically a point mass at zero or a sharply peaked continuous density) with a "slab" component (a diffuse or heavy-tailed density) as a prior distribution over model coefficients, supports, or structural parameters. This yields a multimodal posterior that encodes combinatorial uncertainty over sparsity patterns and continuous uncertainty over effect sizes, with formal guarantees and efficient computational algorithms now available for exact posterior sampling in regimes previously inaccessible to traditional methods (Kumar et al., 4 Mar 2025).

1. Model Specification and Prior Construction

Bayesian spike-and-slab regression is defined via observations y=Xθ+ϵy = X\,\theta + \epsilon, with ϵN(0,σ2In)\epsilon \sim N(0,\sigma^2 I_n), XRn×dX \in \mathbb{R}^{n \times d}, and an unknown sparse θRd\theta \in \mathbb{R}^d. The canonical spike-and-slab prior is

π(θ)=i=1d[(1αi)δ0+αiμ]\pi(\theta) = \bigotimes_{i=1}^d \Big[(1-\alpha_i)\,\delta_0 + \alpha_i\,\mu\Big]

where each coordinate is zero ("spike") with probability 1αi1-\alpha_i or drawn i.i.d. from the slab density μ\mu (e.g., N(0,1)N(0,1) or Laplace) with probability αi\alpha_i (Kumar et al., 4 Mar 2025). This induces a prior over supports S=supp(θ)S = \mathrm{supp}(\theta): ϵN(0,σ2In)\epsilon \sim N(0,\sigma^2 I_n)0 The prior is thus both discrete (over possible supports) and continuous (over effect sizes in the slab).

2. Posterior Characterization and Analytic Structure

Under Gaussian noise and a Gaussian slab, the posterior takes the form: ϵN(0,σ2In)\epsilon \sim N(0,\sigma^2 I_n)1 In the Gaussian-slab case, one obtains an explicit mixture representation: ϵN(0,σ2In)\epsilon \sim N(0,\sigma^2 I_n)2 with explicit formulas for the mixture means, covariances, and weights (Kumar et al., 4 Mar 2025): ϵN(0,σ2In)\epsilon \sim N(0,\sigma^2 I_n)3

ϵN(0,σ2In)\epsilon \sim N(0,\sigma^2 I_n)4

For Laplace slab, closed-form Gaussian integrals are lost, but the same underlying mixture structure persists, subject to numerical integration (Kumar et al., 4 Mar 2025).

3. Sampling-Complexity, Restricted Isometry and Statistical Regimes

A key insight is that posterior sampling, in order to be accurate and tractable in high dimensions, demands that ϵN(0,σ2In)\epsilon \sim N(0,\sigma^2 I_n)5 satisfy a restricted isometry property (RIP) up to sparsity ϵN(0,σ2In)\epsilon \sim N(0,\sigma^2 I_n)6, where ϵN(0,σ2In)\epsilon \sim N(0,\sigma^2 I_n)7 is the expected sparsity and ϵN(0,σ2In)\epsilon \sim N(0,\sigma^2 I_n)8 is total variation error tolerance (Kumar et al., 4 Mar 2025). For Gaussian ϵN(0,σ2In)\epsilon \sim N(0,\sigma^2 I_n)9,

  • Polynomial-time, high-accuracy sampler: XRn×dX \in \mathbb{R}^{n \times d}0 achieves TV error XRn×dX \in \mathbb{R}^{n \times d}1 in XRn×dX \in \mathbb{R}^{n \times d}2 operations.
  • Near-linear-time sampler: XRn×dX \in \mathbb{R}^{n \times d}3 achieves the same accuracy in XRn×dX \in \mathbb{R}^{n \times d}4 (Kumar et al., 4 Mar 2025).

RIP suffices for standard random matrix ensembles (sub-Gaussian, subsampled Fourier, etc.), requiring only a sublinear scaling of samples in the dimension XRn×dX \in \mathbb{R}^{n \times d}5. These bounds break past barriers of strictly linear sample regimes or strong SNR assumptions common in previous literature.

4. Algorithmic Advances: Hint-Vector Estimation and Product Proposals

The sampling framework proceeds in two stages:

  1. Hint-vector estimation: Fast sparse-recovery (XRn×dX \in \mathbb{R}^{n \times d}6 or XRn×dX \in \mathbb{R}^{n \times d}7-based) yields an initial estimate XRn×dX \in \mathbb{R}^{n \times d}8 with support XRn×dX \in \mathbb{R}^{n \times d}9. With high posterior probability, θRd\theta \in \mathbb{R}^d0 for sample support θRd\theta \in \mathbb{R}^d1 and θRd\theta \in \mathbb{R}^d2 small (Proposition 2.8).
  2. Coordinatewise product-proposal + rejection sampling: Using "recentering" (Lemma 3.3), weights over supports θRd\theta \in \mathbb{R}^d3 become simple coordinatewise products. A conditional Poisson product distribution θRd\theta \in \mathbb{R}^d4 over θRd\theta \in \mathbb{R}^d5 can be sampled in θRd\theta \in \mathbb{R}^d6 time and is provably close (within constant ratios, Lemma 3.7) to the true posterior mass, allowing TV-accurate rejection sampling to the true posterior support θRd\theta \in \mathbb{R}^d7 (Kumar et al., 4 Mar 2025).

Once support θRd\theta \in \mathbb{R}^d8 is drawn, sampling θRd\theta \in \mathbb{R}^d9 is exact (Lemma 3.11).

5. Provable Posterior Guarantees, Sparsity, and Estimation-to-Sampling Results

The main theorem (Kumar et al., 4 Mar 2025): Under stipulated RIP and sample size,

π(θ)=i=1d[(1αi)δ0+αiμ]\pi(\theta) = \bigotimes_{i=1}^d \Big[(1-\alpha_i)\,\delta_0 + \alpha_i\,\mu\Big]0

The computational cost is as above, and rejection sampling mixes efficiently over the posterior support (mixing cost π(θ)=i=1d[(1αi)δ0+αiμ]\pi(\theta) = \bigotimes_{i=1}^d \Big[(1-\alpha_i)\,\delta_0 + \alpha_i\,\mu\Big]1). The near-linear-time sampler requires π(θ)=i=1d[(1αi)δ0+αiμ]\pi(\theta) = \bigotimes_{i=1}^d \Big[(1-\alpha_i)\,\delta_0 + \alpha_i\,\mu\Big]2 samples and achieves comparable accuracy.

Structural lemmas include:

  • Support-sparsity (Corollary 2.10): For any product prior, π(θ)=i=1d[(1αi)δ0+αiμ]\pi(\theta) = \bigotimes_{i=1}^d \Big[(1-\alpha_i)\,\delta_0 + \alpha_i\,\mu\Big]3.
  • Posterior-to-estimation (Proposition 2.8): Any estimator π(θ)=i=1d[(1αi)δ0+αiμ]\pi(\theta) = \bigotimes_{i=1}^d \Big[(1-\alpha_i)\,\delta_0 + \alpha_i\,\mu\Big]4 that is good in metric π(θ)=i=1d[(1αi)δ0+αiμ]\pi(\theta) = \bigotimes_{i=1}^d \Big[(1-\alpha_i)\,\delta_0 + \alpha_i\,\mu\Big]5 with high probability induces a sampling procedure that, with probability π(θ)=i=1d[(1αi)δ0+αiμ]\pi(\theta) = \bigotimes_{i=1}^d \Big[(1-\alpha_i)\,\delta_0 + \alpha_i\,\mu\Big]6, draws π(θ)=i=1d[(1αi)δ0+αiμ]\pi(\theta) = \bigotimes_{i=1}^d \Big[(1-\alpha_i)\,\delta_0 + \alpha_i\,\mu\Big]7-samples within double the estimation error in π(θ)=i=1d[(1αi)δ0+αiμ]\pi(\theta) = \bigotimes_{i=1}^d \Big[(1-\alpha_i)\,\delta_0 + \alpha_i\,\mu\Big]8.

6. Extension to Laplace Slabs

For slab π(θ)=i=1d[(1αi)δ0+αiμ]\pi(\theta) = \bigotimes_{i=1}^d \Big[(1-\alpha_i)\,\delta_0 + \alpha_i\,\mu\Big]9, the Gaussian mixture structure lacks closed-form integrals. The algorithm adapts by:

  • Using Monte Carlo or annealing-based normalizer estimation for each mixture component (Prop 4.1, Cor 4.6) to accuracy 1αi1-\alpha_i0 in 1αi1-\alpha_i1.
  • Restricting to 1αi1-\alpha_i2 to control Laplace tail errors (Lemma 4.2).

Theorem 1.3 (Cor 4.14): For 1αi1-\alpha_i3, 1αi1-\alpha_i4, 1αi1-\alpha_i5, sample in 1αi1-\alpha_i6 time with TV1αi1-\alpha_i7. Near-linear-time algorithms analogous to the Gaussian case hold for 1αi1-\alpha_i8 (Kumar et al., 4 Mar 2025).

7. Technical and Applied Impact

This framework supplies the first polynomial-time, sublinear-measurement, provably exact samplers for spike-and-slab posteriors in high-dimensional sparse linear regression, valid for all SNRs, with flexible extension to Laplace diffuse priors (Kumar et al., 4 Mar 2025). The approach unifies continuous algorithms, precise RIP-based concentration, and rejection-sampling with conditional-Poisson tractable supports for rigorous total variation-accuracy and running-time guarantees.

The framework establishes spike-and-slab posterior sampling—with certifiable accuracy and scalable computation—as the theoretical and practical gold standard for Bayesian sparse regression in high dimensions.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Bayesian Spike-and-Slab Framework.