Empirical Bayes Factor (EBF) Methods
- Empirical Bayes Factor (EBF) is a statistical method that uses data-derived prior distributions to correct bias in Bayes factors for hypothesis tests.
- EBF addresses two key applications by providing bias-corrected posterior Bayes factors for composite hypotheses and flexible testing for variance components in mixed models.
- Computational approaches for EBF include Laplace approximation, importance sampling, and the Savage–Dickey density ratio, making it versatile for diverse model specifications.
to=arxiv_search anasiyana 微信公众号天天中彩票json {"query":"Empirical Bayes factor variance components Dudbridge common hypothesis tests", "max_results": 10} to=arxiv_search Empirical Bayes Factor (EBF) denotes a Bayes-factor construction in which a prior-like distribution is estimated from the data rather than fixed entirely a priori. In the recent literature, the term has two closely related uses. One use defines an adjusted posterior Bayes factor for common hypothesis tests, correcting the bias created when the posterior from the data at hand is reused in the Bayes factor (Dudbridge, 2023). Another use defines a flexible Bayes factor for variance-component testing by replacing the boundary null with the equivalent point-null on random effects and estimating their lower-level distribution empirically from the fitted mixed model (Vieira et al., 2024). A related formulation emphasizes a Savage–Dickey density ratio computed from a single full-model fit and recommends plug-in posterior means for variance components in the empirical prior (Vieira et al., 2 Aug 2025).
1. Conceptual scope and notation
In the cited literature, EBF is not a single universally standardized object. For composite hypothesis tests, it is defined as a bias-corrected posterior Bayes factor built from posterior marginal likelihoods under the competing hypotheses. For variance-component testing, it is defined through an empirical lower-level distribution for random effects and typically evaluated by a Savage–Dickey density ratio. Both constructions are empirical Bayes in the sense that the data are used to estimate a distribution that plays the role of a prior or prior-like ingredient in the Bayes factor (Dudbridge, 2023).
A recurrent source of confusion is orientation. In the mixed-effects formulation, one paper defines
so values greater than $1$ support the presence of random effects, while its Savage–Dickey quantity is written as and is the reciprocal of that EBF. A related formulation instead writes directly for the Savage–Dickey ratio
so positive supports exclusion of the tested random effect (Vieira et al., 2024). This suggests that empirical Bayes factor results must always be read together with their stated orientation, rather than inferred from the label “EBF” alone (Vieira et al., 2 Aug 2025).
The methodological motivations also differ. In common hypothesis tests, the central issue is how to obtain a Bayes-factor-like quantity without committing to subjective proper priors for composite hypotheses. In variance-component testing, the central issue is the boundary problem created by testing nonnegative variance parameters at zero. The two strands are therefore linked by empirical-Bayes calibration, but they address different inferential bottlenecks.
2. Bias-corrected posterior Bayes factors for common tests
For composite hypotheses and , the proper Bayes factor is
0
The difficulty is that marginal likelihoods require proper priors, whereas improper priors cannot be used and objective priors may be subjectively unreasonable. Dudbridge revisits the posterior Bayes factor by truncating a baseline posterior to the hypothesis space and then evaluating the posterior predictive density of the same data: 1 Because the same data are used both to form the posterior and to evaluate its predictive density, the resulting quantity is biased relative to a Bayes factor that would use a prior learned from independent replicate data (Dudbridge, 2023).
The key calibration result is an expected log-bias term 2. In the regular normal model 3 with known full-rank 4 and a flat prior on 5, the expected bias is 6 for an unrestricted mean. For one-sided components the bias is halved, and for finite interval or point hypotheses the expected bias is zero. This yields the adjusted empirical Bayes factor
7
with the practical approximation
8
where 9 is an effective dimension: 0 for point or finite interval hypotheses, 1 per one-sided parameter, and 2 per two-sided parameter.
This framework yields closed-form or test-based formulas for standard procedures. For a two-sided scalar normal test,
3
For a multivariate normal or 4-type test with dimension 5,
6
The same paper develops corresponding constructions for 7-tests and 8-tests, with bias terms depending on the degrees of freedom, maps likelihood-ratio tests to the 9 formula, and gives the rule of thumb $1$0 when only a $1$1-value is available for small-to-moderate $1$2 (Dudbridge, 2023).
A multiple-testing extension replaces single-test posteriors with a mixture involving the ensemble of tests,
$1$3
and constructs adjusted posterior marginal likelihoods per test. The paper relates this extension to the Optimal Discovery Procedure by replacing unknown densities with posterior densities and bias-correcting only the self-referential term.
3. Variance-component testing via random effects
In mixed-effects models,
$1$4
testing the presence of a random effect is difficult because a variance component must be nonnegative. A null such as $1$5 is therefore a boundary hypothesis. The cited discussion notes that the likelihood-ratio statistic then does not have the standard chi-square distribution; asymptotic results involve mixture chi-square distributions, score tests face similar complications, and parametric bootstrapping is computationally intensive because it requires refitting models many times (Vieira et al., 2024).
The proposed remedy is to test the equivalent point-null on the random effects themselves: $1$6 where $1$7 is a parametric lower-level distribution whose hyperparameters $1$8 are estimated from the data. A common choice is $1$9. In marginal-likelihood form,
0
with
1
The quantity is empirical because the lower-level distribution of the random effects is estimated from the fitted model rather than chosen externally (Vieira et al., 2024).
Operationally, the method is implemented through a Savage–Dickey density ratio. Under 2, let
3
where 4 are empirical estimates of the variance components, and approximate the posterior of 5 under the full model by
6
Then
7
In the notation of the 2024 paper, 8; in the related 2025 formulation, the same Savage–Dickey quantity is itself written as 9 and interpreted directly as evidence for exclusion (Vieira et al., 2 Aug 2025).
The later formulation makes the empirical-prior choice explicit. If 0 parameterizes the covariance of 1, one may either average over the posterior of 2 or use a plug-in estimate 3. The recommended form is
4
where 5 and 6 are the posterior mean and covariance of the tested random effects. The paper states that using the full posterior over 7 in the denominator can bias the EBF toward the full model when the posterior of 8 is spiked near zero, whereas the posterior mean avoids 9 and yields better calibration (Vieira et al., 2 Aug 2025).
4. Computation and model classes
The variance-component EBF admits three computational routes. In an MLE/REML route, one fits the full mixed model, computes 0 by Laplace approximation, adaptive Gauss–Hermite quadrature, or importance sampling, fits the corresponding fixed-effects-only model to obtain 1, and forms 2. In an MCMC route, one estimates
3
and combines this with 4 from the restricted fit. The fastest route is the Savage–Dickey shortcut, which requires only the full model fit and standard outputs such as posterior means, BLUP-like estimates, and covariance matrices (Vieira et al., 2024).
The random-effects distribution 5 need not be a simple i.i.d. Gaussian. The cited formulations allow Gaussian random effects with correlated or block-diagonal covariance, spatial CAR/ICAR effects with structured precision or covariance, state-space priors for latent dynamic states, and Gaussian process priors in nonlinear regression. For large random-effect dimension 6, log-determinants are computed via Cholesky decompositions, and when software does not supply the full covariance 7, a diagonal approximation using squared standard errors is described as conservative (Vieira et al., 2024).
| Model class | Tested random effects | Computation note |
|---|---|---|
| Logistic crossed mixed-effects models | crossed random intercepts and slopes | Laplace approximation, adaptive Gauss–Hermite quadrature, or Savage–Dickey |
| Spatial random effects models (CAR/ICAR) | exchangeable and spatial effects | structured Gaussian covariance or precision in the determinant |
| Dynamic structural equation models / state-space | latent states or time-level random effects | Kalman filtering/smoothing in the Gaussian case |
| Random intercept cross-lagged panel models | subject-specific random intercepts | joint or separate tests via posterior covariance blocks |
| Nonlinear regression and Gaussian processes | nonlinear random effects or GP components | GP priors treated as structured random effects |
Specific model formulations appear throughout the cited work. For logistic crossed mixed-effects models,
8
For spatial random effects,
9
with the corresponding joint Gaussian covariance implied by the neighborhood matrix. For state-space and dynamic SEM models,
0
For Gaussian processes,
1
The method is therefore presented as a general random-effects testing device rather than a procedure tied to a single likelihood family (Vieira et al., 2024).
Recommended software workflows in the cited material include lme4, glmmTMB, mgcv, spaMM, INLA, Stan/brms, JAGS, TMB, blavaan, and lavaan; the later paper also mentions an R package “EBF” for automating computations (Vieira et al., 2 Aug 2025).
5. Calibration, interpretation, and methodological relations
In Dudbridge’s formulation, the EBF is closely related to the widely applicable information criterion. In the regular normal model,
2
which equals WAIC under a vague prior. Consequently,
3
The stated penalty per parameter is 4, smaller than the AIC penalty of 5, while the method does not impose the 6 penalty characteristic of BIC (Dudbridge, 2023).
The same paper proposes interpreting Bayes factors on a logarithmic scale with base 7. One unit of evidence is a Bayes factor of 8, with approximate landmarks of 9 for two units, 0 for three units, and 1 for four units. This gives an evidence scale explicitly tied to Bayes factors rather than 2-values, even though the same paper also provides the pragmatic approximation 3 when only a 4-value is available (Dudbridge, 2023).
For variance-component EBFs, the cited guidance is Jeffreys/Kass–Raftery style. One formulation states that 5 gives weak evidence for random effects, 6 moderate evidence, 7 strong evidence, and values above 8 very strong evidence. The same source recommends reporting 9 for interpretability and numerical stability. The related formulation writes the evidence in the opposite orientation, with 0 indicating no evidence, 1–2 moderate evidence for exclusion, 3–4 strong evidence, and values below 5 supporting inclusion. A plausible implication is that interpretation is straightforward only after the direction of the ratio has been specified (Vieira et al., 2024).
The comparison with competing methods is explicit. Likelihood-ratio tests, restricted likelihood-ratio tests, and score tests still face the boundary hypothesis; mixture chi-square calibrations can be inaccurate in finite samples, and parametric bootstrap is costly. Fully Bayesian Bayes factors with proper priors are principled but can be highly sensitive to prior choices and expensive to compute. Information criteria and cross-validation may require fitting many model combinations. The EBF is presented as attractive because it tests 6 rather than 7, removes the need to hand-specify priors for difficult variance components, and often requires only the full model fit (Vieira et al., 2024).
The limitations are equally explicit. The variance-component EBF relies on a Gaussian posterior approximation for the tested random effects; accuracy may degrade in small samples, highly non-Gaussian posteriors, or multimodal settings. Its value depends on how 8, 9, 00, or 01 are estimated; classical fits may be slightly liberal; misspecification of the likelihood or lower-level distribution can bias conclusions; and high-dimensional determinant calculations may dominate the evidence. The papers also note the “double use of data” critique of empirical Bayes constructions, though they argue that the empirical plug-in substantially reduces prior sensitivity relative to subjective variance-component priors (Vieira et al., 2 Aug 2025).
6. Empirical behavior, applications, and reporting practice
The synthetic studies summarized in the variance-component papers vary the number of clusters, cluster sizes, and the magnitude of true heterogeneity. The 2024 paper reports that Bayesian fits produced positive 02 more clearly when the tested variance was essentially zero, while classical fits often yielded near-zero evidence; when other random effects truly varied, the EBFs decisively supported their inclusion. The 2025 formulation, using a cross-classified random-effects design with 03 replications per design cell, reports that the posterior-mean plug-in 04 yields 05 near 06 and 07 for larger 08, with evidence strengthening as 09 and 10 increase (Vieira et al., 2024).
The later paper also provides detailed empirical applications computed from a single full-model fit. In a crossed GLMM for basketball foul calls, random slopes for foul differential had log EBFs 11, 12, and 13, and random intercepts had log EBFs 14, 15, and 16, all in the orientation where positive values support exclusion; the joint test “only random intercepts vs full model” gave 17. In a spatial Poisson model for Scottish lip cancer, the exchangeable random intercept had log EBF 18 and the spatial random effect had log EBF 19, leading to inclusion of the spatial effect only. In a DSEM for binge eating disorder, person-level and time-level intercept random effects and person-level loading random effects were strongly supported, whereas time-level loading random effects had log EBFs 20, 21, and 22, supporting exclusion. In an RI-CLPM with two constructs across five days, the joint inclusion of both random intercepts gave 23, and separate tests also strongly supported both intercepts. In a nonlinear mixed-effects model for Loblolly trees, the three independent random effects ‘Asym’, ‘R0’, and ‘lrc’ had log EBFs 24, 25, and 26, and the joint test excluding ‘R0’ and ‘lrc’ versus the full model gave 27 (Vieira et al., 2 Aug 2025).
Reporting guidance is correspondingly technical. The recommended elements are the fitted model specification, the tested variance component or random-effect block, the plug-in estimates 28 or 29, the method used to obtain them (MLE/REML versus MCMC), convergence and posterior-approximation diagnostics, and sensitivity checks such as alternative 30, diagonal versus full covariance approximations, or Laplace versus MCMC computation. For multiple random-effect terms, the papers distinguish per-block testing from joint testing by stacking blocks into a single vector and using the corresponding joint covariance matrix (Vieira et al., 2024).
A common misconception is that empirical Bayes factors are simply generic Bayes factors with estimated hyperparameters. The cited literature is narrower and more structured. In one strand, the EBF is a bias-corrected posterior Bayes factor calibrated to the scale of a proper Bayes factor; in the other, it is a variance-component testing device that replaces a boundary null on a variance by a point-null on the corresponding random effects and uses the model’s lower-level distribution as an empirical prior. Their shared theme is empirical calibration, but their inferential targets, notation, and computational formulas differ materially (Dudbridge, 2023).