- The paper develops a formal framework comparing Bayesian and quasi-Bayes estimators for Poisson decisions, focusing on oracle regret minimization.
- It establishes that quasi-Bayes estimates from Newton’s algorithm asymptotically merge with full Bayesian procedures in risk performance.
- Empirical studies confirm that quasi-Bayes methods dramatically reduce computational cost while preserving statistical optimality in both one- and multidimensional settings.
Merging Bayes and Quasi-Bayes Empirical Bayes Procedures for Poisson Compound Decisions
Problem Setting and Methodological Framework
The Poisson compound decision model considers estimation of θ1,…,θn where observations Y1,…,Yn are independent Poisson random variables with means θi and loss is measured by squared error. Under a hierarchical mixture, the unknown means θi are i.i.d. draws from an unknown mixing distribution G on R+. The Bayes estimator then takes the standard form (Robbins’ formula):
θ^G(y)=(y+1)pG(y)pG(y+1),pG(y)=∫ΘPoisson(y∣θ)G(dθ).
Empirical Bayes approaches require replacing G (or pG) with an estimate G^, producing a plug-in estimator. Two nonparametric Y1,…,Yn0-modeling strategies for estimating Y1,…,Yn1 are compared:
- Bayesian empirical Bayes: Y1,…,Yn2 is equipped with a Dirichlet Process (DP) prior, posterior inference is used, and Y1,…,Yn3 is taken as the posterior mean measure.
- Quasi-Bayes empirical Bayes: Y1,…,Yn4 is estimated via Newton’s recursive algorithm (originally a stochastic approximation procedure).
While the Bayesian approach is known to be computationally costly in moderate/high-dimensions, the quasi-Bayes variant offers computational efficiency but lacked a rigorous justification as an approximation to the full Bayes procedure, especially in terms of frequentist risk.
Main Theoretical Results: Frequentist Merging of Bayes and Quasi-Bayes
The paper develops a formal framework to compare Bayes and quasi-Bayes plug-in estimators asymptotically under a frequentist oracle model where the data are generated by a Poisson mixture with oracle mixing distribution Y1,…,Yn5. The key metric is regret, i.e. excess Bayes risk:
Y1,…,Yn6
where Y1,…,Yn7 is the Bayes rule under the oracle Y1,…,Yn8. Regret is analyzed for both approaches, as well as for the difference between the Bayes and quasi-Bayes estimators.
Posterior contraction rates for marginal PMFs: For the Bayesian approach, the marginal PMF derived from the posterior mean mixing measure contracts to Y1,…,Yn9 at θi0. For Newton’s algorithm (quasi-Bayes), the rate θi1 depends on the learning rate exponent θi2, with typical rates slower than the DP Bayes rate.
Regret convergence: Both procedures yield plug-in estimators whose regret with respect to the oracle vanishes, but with different rates. In both the 1D and θi3D cases, the difference between Bayes and quasi-Bayes estimators (in θi4) vanishes at the same rate as the quasi-Bayes estimator’s own regret—notably, this rate can be made as close as desired to the Bayesian rate by tuning θi5.
Of particular interest is the strong theoretical assertion that the plug-in quasi-Bayes estimator merges with the fully Bayesian estimator in oracle regret:
θi6
at a rate comparable to the regret of the quasi-Bayes estimator itself.
Numerical Study: Accuracy vs Computational Cost
Comprehensive synthetic benchmarking is performed in both one- and two-dimensional settings, across a range of priors (Weibull, Uniform, Half-Gaussian, square-root Half-Cauchy) and sample sizes.
Figure 1
Figure 1: Weibull prior, θi7: data points plotted against the "true" parameters (grey), together with the corresponding oracle Bayes (black), Bayes (red), and quasi-Bayes (blue) estimates.

The figure demonstrates close agreement between the Bayes and quasi-Bayes estimates even at moderate θi8 for the Weibull prior.
Figure 2
Figure 2: Weibull prior, θi9: data points plotted against the "true" parameters (grey), together with the corresponding oracle Bayes (black), Bayes (red), and quasi-Bayes (blue) estimates.

As θi0 increases, the two procedures become indistinguishable both in estimation accuracy and visual trace relative to the oracle Bayes estimate.
Figure 3
Figure 3: Weibull prior: quasi-Bayes (blue) and Bayes (red) estimates compared by E-regret (top panels), computational units (middle panels), and CPU time (bottom panels).

This figure provides a direct tradeoff: E-regret of quasi-Bayes is not statistically inferior to Bayes, but quasi-Bayes incurs orders-of-magnitude less computational burden, especially in CPU time.
Figure 4
Figure 4: Weibull prior: E-regret incurred by using the quasi-Bayes estimate in place of the Bayes estimate.

The empirical regret for quasi-Bayes versus Bayes diminishes rapidly with θi1, supporting the theoretical merging results.
Multidimensional Extension
All principal contraction and merging results extend to θi2-dimensional Poisson mixture problems, with similar structural results for the multi-index Robbins formula. While DP mixture posterior computation becomes drastically more intensive with increasing θi3, quasi-Bayes (Newton) remains feasible. The regret bounds for the difference between Bayes and quasi-Bayes estimators in θi4 dimensions are consistent with the one-dimensional results, with only a logarithmic penalty in the contraction rate.
Discussion and Implications
The theoretical guarantees for merging, combined with the empirical evidence, provide a rigorous rationale for adopting quasi-Bayes (Newton-style recursive) methods in large- or high-dimensional Poisson compound decision settings. Notably:
- Statistical optimality is preserved: The quasi-Bayes estimator matches the Bayes estimator in oracle risk asymptotically.
- Computational efficiency: The difference in CPU time is dramatic—Newton’s method exploits a light sequential structure, while DP posterior samplers require expensive MCMC for every batch of data.
- Robustness to tuning: Empirically, the performance of Newton’s method is robust to grid resolution and the choice of learning rate; regret convergence is not sensitive to fine hyperparameter adjustment.
There remain open questions on the optimality and sharpness of the regret rates for Newton’s method: the statistical rates follow from general stochastic approximation theory, not from minimax lower bounds. It remains to be seen whether the currently proved rate dependence on θi5 is intrinsic or an artifact; sharper non-asymptotic analyses may yield better rates or prescriptions for tuning.
The merging phenomenon is not Poisson-specific: analogous arguments may extend to other exponential family compound decision settings (e.g., Gaussian via Tweedie’s formula). Extensions to estimation of more complex functionals—such as sums or predictive aggregates—are natural directions for further work.
Conclusion
This work establishes that quasi-Bayes empirical Bayes procedures for the Poisson compound decision problem, based on Newton's algorithm, are not only computationally superior, but statistically equivalent to fully Bayesian procedures in both theoretical regret and practical performance. The approach resolves a key open question regarding the legitimacy of sequential approximations to Bayes empirical Bayes under frequentist risk, and points to quasi-Bayes recursion as a default choice for nonparametric θi6-modeling in large-scale or high-dimensional settings (2607.02340).