Empirical Partially Bayes Methods Review
- Empirical partially Bayes methods are defined by using Bayesian hierarchies only partially, estimating hyperparameters from data and then treating them as fixed in subsequent inference.
- They balance computational efficiency and uncertainty quantification by replacing full integration with plug-in steps based on marginal likelihoods.
- Multiple variants, including nuisance-only and population empirical Bayes, demonstrate these methods’ practical applications in high-dimensional inference, genomics, and systems medicine.
Empirical partially Bayes methods are procedures in which a Bayesian hierarchy is used only in part: a prior, hyperprior, or nuisance-parameter distribution is estimated from data and then treated as fixed for subsequent posterior calculation, decision making, or testing. In the canonical form, one starts from a hierarchical model
but replaces full posterior inference on with a plug-in step based on the marginal likelihood . This places empirical Bayes, and the broader class of “partially Bayes” procedures, between classical frequentist methods and fully hierarchical Bayes: the final inference is posterior-shaped, but uncertainty in the learned prior component is not fully propagated (Polson et al., 20 May 2026). Several recent works use “partially Bayes” to describe both classical type-II maximum-likelihood hyperparameter estimation and nuisance-only hierarchies in which parameters of interest are treated as fixed while nuisance distributions are learned across many related problems (Klebanov et al., 2016, Ignatiadis et al., 9 Dec 2025).
1. Canonical formulation and the meaning of partial Bayes
In the standard empirical-Bayes construction, the joint density is
and the fully Bayesian posterior is
Empirical Bayes replaces integration over by a type-II maximum-likelihood estimate,
and then forms
This plug-in posterior ignores posterior uncertainty in and uses the data twice (Polson et al., 20 May 2026).
The same structural idea appears in broader partial-Bayes settings. One formulation assumes a partially specified prior family with unknown hyperparameter 0, together with sampling model 1, and seeks exact, frequentist-calibrated inference on 2 even though 3 is unspecified (Qiu et al., 2018). In another formulation, the prior itself is treated as an unknown object estimated from a population 4, after which Bayesian updating is carried out for a new individual. This is explicitly described as “empirical Bayes” or “partially Bayes,” and as a bridge between purely frequentist and purely Bayesian paradigms (Klebanov et al., 2016).
A related compound-decision perspective replaces the unknown prior by the empirical distribution 5. Under exchangeability, the compound risk of a separable rule equals the Bayes risk under 6, so empirical partially Bayes methods can be viewed as attempts to estimate an oracle mixing distribution and then apply the usual Bayes rule (Koenker et al., 2024).
2. Inferential targets, double use of data, and coherence
A central criticism of empirical Bayes is that it does not target the same inferential object as a fully hierarchical model. Dennis Lindley’s quip that “there is only one thing worse than a frequentist, and that is an empirical Bayesian” is presented not merely as caricature but as a technical objection: empirical Bayes uses the same data twice, conflates levels of a hierarchy, and produces posterior-shaped summaries whose uncertainty quantification differs from what a fully hierarchical model delivers (Polson et al., 20 May 2026).
This distinction has become more salient as empirical-Bayes ideas have been extended. Recent variants include empirical Bayes via probabilistic symmetries, empirical Bayes with implicit likelihoods through simulation-based inference, and empirical Bayes for combining experimental and observational data through calibration studies. The criticism advanced against these developments is that they often target inferential objects distinct from the posterior conditional on the realized data. A fully hierarchical alternative retains
7
so credible intervals for 8 propagate directly into inference for 9, preserving coherence (Polson et al., 20 May 2026).
The computational argument that historically favored empirical Bayes has also weakened. Modern hierarchical Bayes routinely uses MCMC, including Gibbs sampling, Metropolis within Gibbs, and HMC; variational inference, including mean-field, structured VI, and black-box VI; neural amortization with normalizing flows and invertible networks; and sequential Monte Carlo or particle methods for streaming data. In this setting, the extra cost of sampling a hyperparameter or adding an extra variational dimension is described as modest relative to the benefit of avoiding data-twice use, obtaining correct coverage, and preserving coherence (Polson et al., 20 May 2026).
A closely related point appears in nuisance-only partial-Bayes testing. Partially Bayes 0-values condition on ancillary nuisance-parameter statistics and pool information across units only through those ancillaries, thereby avoiding the posterior-predictive reuse of the same test statistic in both fitting and checking. This is explicitly contrasted with posterior predictive 1-values, which double-use the data and are not uniform even in a fully Bayesian frame (Ignatiadis et al., 9 Dec 2025).
3. Tweedie identities, f-modeling, and coherent shrinkage
Much of empirical partially Bayes methodology is organized around identities that express posterior summaries in terms of marginal densities. In the normal-means model 2 with marginal density
3
Tweedie’s formula states
4
Efron’s 5-modeling estimates 6 by smoothing the observed 7 with splines, kernels, or histograms, differentiates 8, and plugs the estimated score into Tweedie’s identity. The important caveat is that a generic smoothed 9 need not have the form 0 for any probability density 1; consequently, the resulting shrinkage estimate need not equal the posterior mean under any joint model. The horseshoe Tweedie formula is presented as a coherent alternative because it does arise from an explicit hierarchical prior (Polson et al., 20 May 2026).
For the hierarchical horseshoe prior,
2
the shrinkage weight
3
yields the conditional posterior-mean identity
4
The exact score function satisfies
5
The half-Cauchy mixture produces aggressive shrinkage near zero and near-zero shrinkage for large 6, a combination emphasized as one reason to prefer coherent global-local priors over ad hoc score smoothing (Polson et al., 20 May 2026).
The same logic has been generalized to heteroscedastic normal means with unequal and unknown variances. If
7
and 8 denotes the joint marginal density under an unspecified prior 9, then under weighted loss 0 the Bayes rule is
1
and can be written as
2
That framework also introduces a moment-generating-function representation of the full posterior, so that once 3 is known the posterior law is determined without explicitly specifying or estimating the prior 4. The method is proposed as a unified 5-modeling framework for point estimation, uncertainty quantification, and hypothesis testing in heterogeneous data environments (Zhao et al., 23 Apr 2026).
4. Major methodological variants
The contemporary literature contains several distinct empirical partially Bayes constructions, differing mainly in what is estimated and what inferential target is retained.
| Variant | Estimated object | Resulting target |
|---|---|---|
| Population empirical Bayes | Empirical population distribution or latent dataset prior | Population posterior or POP predictive density |
| Bayesian empirical Bayes via probabilistic symmetries | Directing measure 6 from an ergodic decomposition | Plug-in posterior conditional on 7 |
| Implicit-likelihood empirical Bayes | Simulator-induced posterior approximation 8 | Amortized “EB posterior” or “population posterior” |
| Nuisance-only partially Bayes | Mixing law of nuisance parameters across units | Conditional tail-area 9-values and moderated tests |
Population empirical Bayes (POP-EB) introduces a latent dataset 0 with prior 1, the empirical distribution of the observed sample. The POP predictive density is
2
with bootstrap approximations yielding a MAP version based on the single best bootstrap and a full-Bayes approximation based on a weighted mixture over bootstraps. For complex models, POP-EB uses bumping variational inference (BUMP-VI), which interleaves bootstrap-specific ELBO gradients with predictive-score selection on the original data. Reported applications include linear regression on health data, a Gaussian mixture model of natural images, and latent Dirichlet allocation on scientific documents, where POP-EB improves held-out predictive scores relative to classical Bayesian inference (Kucukelbir et al., 2014).
A second family of methods estimates a directing measure directly. Under probabilistic symmetries, exchangeable arrays, spatial fields, and covariate-indexed hierarchies admit an ergodic decomposition over a directing measure 3, which is then estimated by maximizing
4
Variational approximations with neural networks are used for 5. Closely related simulation-based methods replace explicit likelihood evaluation by a simulator 6, then use neural posterior estimation, neural likelihood estimation, or neural ratio estimation to produce 7, sometimes treating it directly as an empirical-Bayes posterior even when no joint prior on 8 is written down (Polson et al., 20 May 2026).
In nuisance-only partially Bayes inference, the primary parameters remain fixed while only nuisance parameters are hierarchically modeled. One formulation observes summaries 9 with 0 a test statistic and 1 ancillary for the primary parameter, posits
2
and defines the partially Bayes 3-value by conditioning on all ancillaries but leaving out 4 when learning the nuisance law for unit 5. Under the fully Bayes frame this yields exact uniformity, and under large-6 or large-7 asymptotics it yields frequentist uniformity as well (Ignatiadis et al., 9 Dec 2025).
This nuisance-only perspective has become prominent in high-throughput biology. In the normal-means problem with unknown variances, conditional 8-values based on a prior 9 over variances satisfy an Eddington/Tweedie-type formula, and nonparametric maximum-likelihood estimation of 0 yields plug-in 1-values that can be combined with Benjamini–Hochberg for asymptotic FDR control (Ignatiadis et al., 2023). Limma-trend is analyzed as computing approximate partially Bayes 2-values that condition on residual sample variance and a unit-level summary, while nonparametric generalizations estimate the residual variance prior by NPMLE and retain asymptotic FDR control even when the trend is misspecified or inconsistently estimated (Nandy et al., 20 May 2026). Two-sample extensions with unequal variances use either a prior on the variance ratio 3 or a joint prior on 4, both estimated by NPMLE, yielding asymptotically uniform 5-values as the number of features grows with fixed replicate counts (Ling et al., 1 Oct 2025).
5. Frequentist calibration, asymptotics, and computation
A major line of work studies whether empirical Bayes and Bayes eventually agree. For a prior family 6 and plug-in posterior 7, Bayesian weak merging with every fixed-8 posterior is equivalent to weak consistency of the empirical-Bayes posterior under exchangeability and injectivity of 9. Under regularity conditions, marginal-likelihood empirical Bayes asymptotically selects the hyperparameter value whose prior density most favors the truth, and in regular parametric settings the plug-in posterior can merge strongly in total variation with the Bayesian posterior at the oracle hyperparameter. In nonparametric Dirichlet-process mixture settings, by contrast, only weak merging is generally possible because priors with different concentration parameters are mutually singular (Petrone et al., 2012).
A different route to calibration is to avoid plug-in posteriors altogether. For partial-Bayes problems with prior family 0 and unspecified 1, the inferential-model approach builds an association 2, adds a prior mapping 3, and then propagates a predictive random set through a reduced association for 4. The resulting plausibility interval
5
satisfies
6
under mild regularity. When the missing part of the prior is known, the construction recovers the Bayesian 7 credible interval; in large samples it converges to the optimal one (Qiu et al., 2018).
For compound decisions and denoising, nonparametric maximum-likelihood estimation gives sharper oracle comparisons. In exchangeable models 8, the Bayes rule under squared-error loss uses the posterior mean under the unknown mixing distribution 9. NPMLE estimates 0 by convex optimization without an explicit tuning parameter, and the resulting empirical-Bayes posterior mean enjoys oracle inequalities: for sub-Gaussian 1, regret relative to the oracle Bayes rule is
2
while heavier-tailed classes achieve the stated nonparametric rates up to logarithmic factors (Koenker et al., 2024). In multivariate heteroscedastic Gaussian mixtures,
3
finite-dimensional approximations to the NPMLE still yield average Hellinger bounds for the fitted marginals and regret bounds showing that the empirical-Bayes posterior means nearly match the oracle posterior means (Soloff et al., 2021).
Empirical partially Bayes computation has also developed into a distinct technical area. One strategy runs a single MCMC chain under a reference prior 4, then reweights to estimate the entire curves
5
for all 6. Under geometric ergodicity and empirical-process conditions, this yields uniform strong consistency, functional central limit theorems, argmax consistency for empirical-Bayes hyperparameter estimates, and simultaneous confidence bands (Doss et al., 2018). A related multi-chain method chooses skeleton priors, estimates normalizing-constant ratios by reverse logistic regression, and then uses importance sampling with control variates to estimate Bayes factors and posterior expectations across a continuum of hyperparameters (Buta et al., 2012).
6. Domains of application and contemporary practice
In systems medicine, empirical Bayes is used to construct informative priors from many patients before carrying out patient-specific inference. One comparison considered four priors: a non-informative prior, a nonparametric maximum-likelihood prior, a maximum penalized likelihood prior, and a doubly-smoothed MLE prior. In a harmonic-oscillator example with 7, the reported posterior success-rate error standard deviations were 8 for the uniform prior, 9 for NPMLE, 00 for DS-MLE, and 01 for MPLE. In a 02 ODE model of the human menstrual cycle, estimated priors produced much sharper marginal posteriors than the non-informative baseline, while DS-MLE and MPLE avoided the atomic collapse of NPMLE (Klebanov et al., 2016).
In high-dimensional prediction, empirical Bayes is used to learn from many variables and from external “co-data.” In Gaussian ridge models, hyperparameters can be estimated by marginal likelihood or moment matching; in group-regularized logistic regression, hybrid empirical Bayes–full Bayes models retain EB-estimated group multipliers while placing a hyperprior on a global shrinkage parameter. In spike-and-slab variable selection, co-data enter through
03
with 04 estimated by Gibbs-EB or moment-based EB. These constructions are described as especially useful when the prior has multiple parameters that encode a priori information on variables (Wiel et al., 2017).
In matrix completion, empirical Bayes has been adapted to binary observation models. For a partially observed binary matrix 05, the method places a row-wise Gaussian prior 06, estimates 07 by MCEM from the marginal likelihood, and samples from the posterior using Albert–Chib augmentation and Gibbs updates. In the continuous Gaussian limit, the construction connects to the Efron–Morris singular-value shrinker; in the 1-bit setting the Gibbs sampler performs adaptive singular-value shrinkage through 08. Reported real-data results show that the heterogeneous-column model EB2 attained classification accuracy 09, cross-entropy 10, and 11 on Jester, and accuracy 12, cross-entropy 13, and 14 on MovieLens 100K (Matsuda, 10 May 2026).
Large-scale inference in genomics and related fields has become a principal domain for empirical partially Bayes methods. Limma-trend can be interpreted as a parametric partially Bayes procedure that places a trend-dependent prior on residual variances and then computes moderated-15 16-values conditional on the residual sample variance and a unit-level summary. Nonparametric generalizations estimate either the residual multiplier distribution or the full conditional variance distribution and obtain asymptotic FDR control under dense signals (Nandy et al., 20 May 2026). In unequal-variance two-sample testing, VREPB and DVEPB estimate respectively a prior on 17 or a joint prior on 18, again by NPMLE, to produce partially Bayes 19-values with asymptotic type-I control (Ling et al., 1 Oct 2025).
Pharmacovigilance provides another compound-decision setting. In spontaneous reporting systems, counts satisfy
20
with signal strengths 21 drawn from an unknown prior 22. Nonparametric empirical Bayes methods estimate 23 by discrete-support convex optimization, spline-based exponential-family deconvolution, 24-gamma mixtures, or sparse overfitted general-gamma mixtures. Posterior summaries include the empirical Bayes geometric mean
25
and the lower-bound 26, the 27th posterior percentile. The package pvEBayes implements these methods and graphical summaries; in one statin example, approximately 28 AE-drug signals with 29 were reported among 30 statins (Tan et al., 30 Nov 2025).
Across these applications, a consistent contemporary recommendation is to redeploy modern variational, neural, and simulation-based machinery in service of fully hierarchical Bayes rather than to bypass the hierarchy. In that view, one should write down a hyperprior for 31 or for the directing measure 32, use scalable inference to approximate the full posterior, replace incoherent marginal-score smoothing by g-modeling or continuous global-local priors, and prefer online VB or sequential Monte Carlo for streaming data. The broad implication is not that empirical partially Bayes methods are obsolete, but that they are best understood as approximations with distinct inferential targets, useful theoretical guarantees in some regimes, and increasingly explicit alternatives when coherence is the primary objective (Polson et al., 20 May 2026).