Papers
Topics
Authors
Recent
Search
2000 character limit reached

Bayesian Group Global-Local Shrinkage Prior

Updated 20 November 2025
  • Bayesian group global-local shrinkage prior is a flexible hierarchical model that uses group-specific local scales and a global parameter for adaptive variable selection in high dimensions.
  • It employs a polynomial-tailed modification to strongly shrink noise while retaining large signals, ensuring robust group selection and optimal estimation.
  • The approach features an efficient half-thresholding rule that outperforms traditional spike-and-slab and group LASSO methods in both empirical and theoretical studies.

The Bayesian group global-local shrinkage prior is a flexible class of hierarchical priors developed to address high-dimensional variable selection and estimation problems where covariates or coefficients are structured in groups. Building on the success of continuous global-local shrinkage approaches such as the horseshoe, these priors enable simultaneous adaptation to group-level sparsity and signal strength by assigning each group a local scale parameter that interacts multiplicatively with a global shrinkage parameter. This construction yields strong shrinkage for groups with negligible effects, while preserving estimation accuracy for groups that contain genuine signal, and admits polynomial tails for detection of large effects. A salient variant of this framework uses a “modified global-local” structure that induces polynomially decaying tails for the group coefficients, optimally balancing the need to shrink noise while retaining prominence for large signals. Theoretical and empirical investigations document favorable selection and estimation properties, with performance rivaling or exceeding canonical two-group spike-and-slab approaches in group selection under high-dimensional scaling (Paul et al., 2023).

1. Hierarchical Model Formulation

Let yRny \in \mathbb{R}^n be the response and X=[X1,,XG]X = [X_1, \dots, X_G] the design matrix concatenated from GG groups (XgX_g of size n×mgn \times m_g). The target coefficients are partitioned as β=(β1,,βG)\beta = (\beta_1, \ldots, \beta_G) with βgRmg\beta_g \in \mathbb{R}^{m_g}, gmg=p\sum_g m_g = p. The Gaussian linear model is

yX,β,σ2N(Xβ,σ2In).y \mid X, \beta, \sigma^2 \sim N(X\beta, \sigma^2 I_n).

The group global-local shrinkage prior adopts the following form (‘global-local g-prior’): βgλg,τ,σ2Nmg(0,σ2τ2λg2(XgXg)1), λg2π(λg2),where π(λg2)(λg2)a1L(λg2), a>0,\begin{aligned} \beta_g \mid \lambda_g, \tau, \sigma^2 &\sim N_{m_g}\left(0,\, \sigma^2 \tau^2 \lambda_g^2 (X_g^\top X_g)^{-1}\right), \ \lambda_g^2 &\sim \pi(\lambda_g^2), \quad \text{where} \ \pi(\lambda_g^2) \propto (\lambda_g^2)^{-a-1} L(\lambda_g^2), \ a>0, \end{aligned} with X=[X1,,XG]X = [X_1, \dots, X_G]0 slowly varying. The global scale X=[X1,,XG]X = [X_1, \dots, X_G]1 is either set as a tuning parameter when the group sparsity level is known, or assigned a prior (full or empirical Bayes estimation) such as a truncated half-Cauchy. The variance X=[X1,,XG]X = [X_1, \dots, X_G]2 is given a Jeffreys’ prior (X=[X1,,XG]X = [X_1, \dots, X_G]3) in practice or sometimes fixed for theoretical analysis.

The joint prior density is thus explicitly

X=[X1,,XG]X = [X_1, \dots, X_G]4

2. Polynomial-tailed Modification and Tail Properties

The critical feature distinguishing the Bayesian group global-local shrinkage prior is its polynomial-tailed structure on the group coefficient vector. Specifically, the local scales X=[X1,,XG]X = [X_1, \dots, X_G]5 are described by

X=[X1,,XG]X = [X_1, \dots, X_G]6

with X=[X1,,XG]X = [X_1, \dots, X_G]7 Karamata slowly varying. The resulting marginal prior on X=[X1,,XG]X = [X_1, \dots, X_G]8 is

X=[X1,,XG]X = [X_1, \dots, X_G]9

inducing heavy (polynomial) tails. The exponent GG0 directly modulates the tail decay: small GG1 yields heavier tails, which encourages concentration of mass at zero, but does not overly penalize large signals. Special cases include the horseshoe prior and other 'one-group' polynomial-tailed forms as in Tang et al. (2018) (Paul et al., 2023).

3. Selection via the Half-Thresholding Rule

A distinguishing feature is the explicit computationally tractable selection rule. For a block-orthogonal design (GG2 for GG3), the posterior mean of each group factors as

GG4

Let GG5 denote the shrinkage factor. The half-thresholding rule declares group GG6 active if

GG7

This threshold rule is fully specified by the posterior mean, requiring no marginal likelihood computation or combinatorial search, and is adaptive to signal strength and group size (Paul et al., 2023).

4. Global Scale (GG8) Selection Strategies

The choice of global shrinkage parameter GG9 is pivotal for controlling the trade-off between bias and variance:

  • Known sparsity: If the proportion of active groups, XgX_g0, is known, a near-optimal choice is XgX_g1 for small XgX_g2.
  • Empirical Bayes: When sparsity is unknown, an empirical Bayes estimator (after van der Pas et al.) is used:

XgX_g3

with XgX_g4, XgX_g5, XgX_g6.

  • Full Bayes: A truncated half-Cauchy prior, XgX_g7 on XgX_g8, ensures that XgX_g9 concentrates in the oracle regime.

Adaptation to unknown sparsity is thus achieved without combinatorial model enumeration (Paul et al., 2023).

5. Theoretical Guarantees

Let n×mgn \times m_g0, n×mgn \times m_g1, and total number of groups n×mgn \times m_g2 with n×mgn \times m_g3.

  • Variable selection consistency: Under standard regularity (group designs with bounded eigenvalues, signals not vanishing, bounded group size), and suitable choices of n×mgn \times m_g4 (e.g., n×mgn \times m_g5), the half-thresholding rule is selection consistent:

n×mgn \times m_g6

  • Oracle estimation rates: For any unit vector n×mgn \times m_g7 with support in n×mgn \times m_g8 and under further eigenvalue and signal bounds, the estimator achieves asymptotic normality at the minimax-optimal rate:

n×mgn \times m_g9

  • These properties extend to empirical Bayes and full Bayes strategies, requiring only mild technical modifications for β=(β1,,βG)\beta = (\beta_1, \ldots, \beta_G)0 or alternative empirical β=(β1,,βG)\beta = (\beta_1, \ldots, \beta_G)1 selection (Paul et al., 2023).

6. Empirical Performance and Method Comparisons

Extensive simulations were conducted across nine regimes (varying β=(β1,,βG)\beta = (\beta_1, \ldots, \beta_G)2, signal strength, group sizes, orthogonality of design). Principal comparators include:

Performance metrics: Misclassification Probability (MP), False Positive Rate (FPR), True Positive Rate (TPR).

Findings:

  • MGH (and GH) priors yield the lowest MP and FPR and highest TPR, especially under weak or moderate signal regimes and smaller β=(β1,,βG)\beta = (\beta_1, \ldots, \beta_G)3.
  • Empirical Bayes and Full Bayes variants match nearly the oracle-tuned half-thresholding rule.
  • Two-group priors (GSD-SSS, BGL-SS) require stronger signal or larger β=(β1,,βG)\beta = (\beta_1, \ldots, \beta_G)4 to achieve similar performance.
  • Group LASSO tends to overselection (high FPR) except under strong signals or large β=(β1,,βG)\beta = (\beta_1, \ldots, \beta_G)5.

This demonstrates that one-group, polynomial-tailed global-local priors with the half-thresholding rule match or outperform classical two-group spike-and-slab or penalized likelihood group selection methods while simultaneously offering substantial computational and inferential simplicity (Paul et al., 2023).

7. Broader Context and Extensions

The group global-local paradigm extends naturally to multilevel and network-structured problems (e.g., multivariate responses (Kundu et al., 2019), multilevel models with joint control via Dirichlet or Beta-P distributions (Aguilar et al., 2022), gene network estimation (Leday et al., 2015), and network-based classification (Guha et al., 2020)). Each variant tailors the local scales to correspond to natural groupings and adapts the thresholding or selection scheme appropriately. Notably, the polynomial-tailed forms enable robust signal recovery in ultra-high-dimensional or weak-signal settings and facilitate practical model selection via continuous shrinkage without the need for discrete model search. In summary, the Bayesian group global-local shrinkage prior furnishes a unified, theoretically rigorous, and empirically validated approach to sparse estimation and group selection across a broad array of high-dimensional settings.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Bayesian Group Global-Local Shrinkage Prior.