Papers
Topics
Authors
Recent
Search
2000 character limit reached

Geometric ergodicity of Gibbs samplers for linear latent models with GIG variance mixtures

Published 8 Feb 2026 in math.ST and math.PR | (2602.07944v1)

Abstract: We study geometric ergodicity of the Gibbs sampler for linear latent non-Gaussian models (LLnGMs), a class of hierarchical models in which conditional Gaussian structure is preserved through generalized inverse Gaussian (GIG) variance-mixture augmentation. Two complementary routes to geometric ergodicity are developed for the marginal chain on the mixing variables. First, we show that the associated Markov operator is trace-class, and hence admits a spectral gap, over a large portion of the GIG parameter space. Second, for the remaining boundary and heavy-tail regimes, we establish geometric ergodicity via drift and minorization, subject to an explicit null-smallness condition that quantifies how the drift interacts with the null space of the observation operator. Together, these results cover the full GIG parameter space, including the normal-inverse Gaussian, generalized asymmetric Laplace, and Student-tt special cases. The geometric ergodicity of this chain underpins the consistency of Gibbs-based stochastic-gradient estimators for maximum likelihood estimation, and we provide conditions that make the required integrability checks transparent. Numerical experiments illustrate the theoretical findings, contrasting mixing efficiency across parameter regimes and probing the role of the null-smallness constant.

Summary

  • The paper establishes geometric ergodicity across the full GIG parameter space by combining trace-class spectral-gap results with drift–minorization arguments, while identifying boundary cases requiring a null-smallness condition.
  • The analysis shows that NIG models are trace-class and geometrically ergodic for every drift value, whereas some boundary regimes have slower mixing or fail to be trace-class despite retaining geometric ergodicity.
  • The paper demonstrates that non-centered parameterization resolves a score-integrability failure when the Gamma shape satisfies α ≤ 1/2, enabling theoretically justified Rao–Blackwellized stochastic-gradient maximum-likelihood estimation.

This paper establishes geometric ergodicity of the Gibbs sampler for linear latent non-Gaussian models (LLnGMs), a class of hierarchical models in which non-Gaussianity is introduced through generalized inverse Gaussian (GIG) variance-mixture augmentation while conditional Gaussian structure is retained. The analysis proceeds along two complementary routes: an operator-theoretic trace-class argument, and a drift–minorization argument for boundary and heavy-tail regimes. Together these results cover the full GIG parameter space, and they are used to justify stochastic-gradient maximum-likelihood estimation via Rao–Blackwellized gradient estimators.

Model class and Gibbs sampler

An LLnGM specifies a Gaussian observation layer YW,VN(Xβ+AW,σϵ2I)Y \mid W,V \sim N(X\beta + AW, \sigma_\epsilon^2 I), a latent field whose precision is modulated by independent mixing variables ViGIG(p,a,b)V_i \sim \mathrm{GIG}(p,a,b) through an invertible structure matrix K(ζ)K(\zeta), and a drift/skewness parameter μ\mu. The GIG family spans the interior ΨI\Psi_I, the inverse-Gamma boundary ΨIG\Psi_{IG}, and the Gamma boundary ΨΓ\Psi_{\Gamma}; as mixing distributions these induce the generalized hyperbolic family, including Student-tt, normal-inverse Gaussian (NIG), and generalized asymmetric Laplace (GAL) marginals. Because the full conditionals of both the latent field and the mixing variables are available in closed form—Gaussian for WW (or MM) and GIG for ViGIG(p,a,b)V_i \sim \mathrm{GIG}(p,a,b)0—a two-block Gibbs sampler is available.

The paper introduces a non-centered parameterization via ViGIG(p,a,b)V_i \sim \mathrm{GIG}(p,a,b)1, which induces the same marginal likelihood and, crucially, the same ViGIG(p,a,b)V_i \sim \mathrm{GIG}(p,a,b)2-marginal transition kernel as the centered parameterization (proved via the bijective change of variables). This equivalence allows the ergodicity analysis to be carried out in whichever parameterization is convenient, while the non-centered form later resolves an integrability obstruction in the score functions.

Operator-theoretic route: trace-class and spectral gap

The ViGIG(p,a,b)V_i \sim \mathrm{GIG}(p,a,b)3-marginal chain has transition density ViGIG(p,a,b)V_i \sim \mathrm{GIG}(p,a,b)4, inducing a Markov operator ViGIG(p,a,b)V_i \sim \mathrm{GIG}(p,a,b)5 on ViGIG(p,a,b)V_i \sim \mathrm{GIG}(p,a,b)6. The paper shows ViGIG(p,a,b)V_i \sim \mathrm{GIG}(p,a,b)7 is a self-adjoint, positive semidefinite contraction with spectrum in ViGIG(p,a,b)V_i \sim \mathrm{GIG}(p,a,b)8 and simple eigenvalue at ViGIG(p,a,b)V_i \sim \mathrm{GIG}(p,a,b)9 (using Harris recurrence and strict positivity of K(ζ)K(\zeta)0), so geometric ergodicity is equivalent to a spectral gap at K(ζ)K(\zeta)1. By a criterion of Qin–Ročková-type, K(ζ)K(\zeta)2 is trace-class if and only if K(ζ)K(\zeta)3.

The main trace-class theorem establishes this integrability in two regimes: (i) K(ζ)K(\zeta)4 with arbitrary K(ζ)K(\zeta)5, and (ii) K(ζ)K(\zeta)6 with K(ζ)K(\zeta)7. The proof bounds the joint density by controlling Bessel-function ratios via explicit lower bounds on K(ζ)K(\zeta)8, applying a Cauchy–Schwarz split and a Gaussian moment lemma to integrate out K(ζ)K(\zeta)9, and then establishing an exponential exponent gap μ\mu0 uniformly over sign vectors μ\mu1. The exponential decay in both μ\mu2 and μ\mu3 dominates the polynomial prefactors, yielding integrability. As a corollary, the NIG model (μ\mu4) is trace-class and geometrically ergodic for all drift values μ\mu5—a clean, practically relevant guarantee since NIG is one of the two distributions closed under convolution limits used in the ngme2 software.

Drift–minorization route for boundary regimes

For regimes outside the trace-class theorem, the paper verifies a geometric drift condition with Lyapunov functions of the form μ\mu6 together with a minorization condition on sub-level sets, invoking Rosenthal's drift–minorization theorem. Three cases are treated:

  • Case I (μ\mu7, μ\mu8, μ\mu9): the ΨI\Psi_I0-update is Inverse-Gamma, and a fractional-moment computation with ΨI\Psi_I1 yields a contractive drift coefficient ΨI\Psi_I2 via a convexity argument on Gamma-function ratios. Notably, the paper proves (Lemma on failure of trace-class) that in this regime, with ΨI\Psi_I3, the operator is not trace-class—so the drift route is not redundant but covers genuinely different chains.
  • Cases II and III (ΨI\Psi_I4 and ΨI\Psi_I5): these require a null-smallness assumption: with ΨI\Psi_I6, either ΨI\Psi_I7 or ΨI\Psi_I8, where ΨI\Psi_I9 and ΨIG\Psi_{IG}0 encode the projection of the drift direction onto the unidentifiable subspace. The condition is automatically satisfied when ΨIG\Psi_{IG}1, and the paper gives structural sufficient conditions (orthogonality of the constant mode, structure preservation by ΨIG\Psi_{IG}2, uniform non-degeneracy on the null space) under which the constant remains bounded as ΨIG\Psi_{IG}3.

The drift proofs decompose the conditional mean ΨIG\Psi_{IG}4 into null-space and range components; the null component contributes exactly the null-smallness ratio as the coefficient of ΨIG\Psi_{IG}5, while the range component and Gaussian fluctuation terms are absorbed into constants. Negative moments of the GIG updates are controlled uniformly in ΨIG\Psi_{IG}6 via Bessel-ratio bounds with a carefully chosen exponent ΨIG\Psi_{IG}7 ensuring ΨIG\Psi_{IG}8.

Combining both routes yields a complete map: geometric ergodicity holds across the full GIG parameter space, with trace-class known in the two regimes above, unknown in two boundary regimes (where geometric ergodicity still holds under the null-smallness condition), and provably absent in the ΨIG\Psi_{IG}9 case. The null-smallness assumption alone is shown to be insufficient for trace-class but sufficient for geometric ergodicity.

Consequences for stochastic-gradient maximum likelihood

Geometric ergodicity is motivated by SGD-based maximum likelihood, as implemented in ngme2. Fisher's identity expresses the log-likelihood gradient as a posterior expectation of the complete-data score, approximated by ergodic averages of Gibbs samples. A Rao–Blackwellized estimator integrates out the Gaussian layer analytically given ΨΓ\Psi_{\Gamma}0, so the estimator depends only on the ΨΓ\Psi_{\Gamma}1-chain—aligning with the ergodicity results—and has reduced variance. The paper is careful to note that finite-length Gibbs runs introduce bias rather than exact unbiasedness; geometric ergodicity gives geometric bias decay ΨΓ\Psi_{\Gamma}2, which is summable under standard step-size schedules ΨΓ\Psi_{\Gamma}3, ΨΓ\Psi_{\Gamma}4, preserving almost-sure convergence of stochastic approximation.

A sharp integrability result identifies a genuine obstruction in the centered parameterization: the posterior expectation ΨΓ\Psi_{\Gamma}5 is finite if and only if ΨΓ\Psi_{\Gamma}6, where ΨΓ\Psi_{\Gamma}7 is the Gamma shape of the mixing variable. For GAL noise with ΨΓ\Psi_{\Gamma}8—and, critically, for SPDE-based models where mesh refinement drives the effective shape parameter ΨΓ\Psi_{\Gamma}9—the complete-data score fails to be integrable under the posterior, so ergodic theorems cannot be applied directly and SGD convergence cannot be established. This explains the instability observed previously in MCEM algorithms for such models. The non-centered parameterization resolves this structurally: its score functions depend on tt0 only polynomially (the tt1-score contains tt2 and self-normalized quadratic forms rather than tt3 terms), so integrability follows directly from the polynomial moment bounds supplied by the drift analysis. The non-centered parameterization is thus not merely a computational device but a necessary theoretical ingredient for rigorous SGD convergence.

Numerical evidence

Two simulation studies (tt4, AR(1)-type tt5, four overdispersed chains of tt6 iterations) support the theory. The first compares mixing across six representative parameter points spanning all regimes. Trace-class interior regimes exhibit integrated autocorrelation times near 1 (ESS/sec up to roughly 105–120), while the tt7, tt8 regime shows the strongest degradation, with IACT rising to about 4.7 and 6.8 for the negative-moment and log summaries respectively—consistent with the boundary-sensitive Lyapunov functions in the drift analysis. Split-tt9 values are approximately 1.0001 throughout. The second study fixes the design geometry with a one-dimensional null space (WW0) and scans WW1, varying the null-smallness constant WW2 while holding WW3 fixed. Increasing WW4 produces systematic IACT inflation, most pronounced for the null-direction statistic and the log summary, providing empirical evidence that the null-smallness constant captures a genuine mixing mechanism rather than a proof artefact.

Limitations and open questions

The paper is explicit about several restrictions. Trace-class remains unresolved in the regimes WW5 and WW6; only geometric ergodicity (conditional on the null-smallness assumption) is established there. The null-smallness condition itself is a constraint on the interaction of WW7, WW8, and WW9; the paper provides dimension-free sufficient conditions but does not characterize necessity, and the drift–minorization constants are conservative and do not yield sharp non-asymptotic Monte Carlo error bounds. The Rao–Blackwellized estimator is argued to behave better under the centered parameterization, but the paper favors changing the parameterization and does not fully analyze integrability of the Rao–Blackwellized centered score. Finally, the framework is tailored to conditionally Gaussian variance-mixture structures; nonlinear observation models or likelihoods without auxiliary-variable representations fall outside its scope, and the authors note that the proof strategy extends to other normal mean–variance mixtures only when comparable envelope bounds and moment control on the mixing density near MM0 and MM1 can be verified.

Conclusion

The paper delivers a complete characterization of geometric ergodicity for Gibbs samplers in GIG variance-mixture latent models, combining a trace-class/spectral-gap argument with a drift–minorization analysis governed by an interpretable null-smallness condition. These stability results, together with the non-centered parameterization's resolution of the score-integrability obstruction at shape parameters below MM2, provide the theoretical foundation for Rao–Blackwellized stochastic-gradient maximum-likelihood estimation in this model class, with numerical experiments corroborating the predicted mixing behavior across parameter regimes and the practical relevance of the null-smallness constant.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.