- The paper establishes geometric ergodicity across the full GIG parameter space by combining trace-class spectral-gap results with drift–minorization arguments, while identifying boundary cases requiring a null-smallness condition.
- The analysis shows that NIG models are trace-class and geometrically ergodic for every drift value, whereas some boundary regimes have slower mixing or fail to be trace-class despite retaining geometric ergodicity.
- The paper demonstrates that non-centered parameterization resolves a score-integrability failure when the Gamma shape satisfies α ≤ 1/2, enabling theoretically justified Rao–Blackwellized stochastic-gradient maximum-likelihood estimation.
This paper establishes geometric ergodicity of the Gibbs sampler for linear latent non-Gaussian models (LLnGMs), a class of hierarchical models in which non-Gaussianity is introduced through generalized inverse Gaussian (GIG) variance-mixture augmentation while conditional Gaussian structure is retained. The analysis proceeds along two complementary routes: an operator-theoretic trace-class argument, and a drift–minorization argument for boundary and heavy-tail regimes. Together these results cover the full GIG parameter space, and they are used to justify stochastic-gradient maximum-likelihood estimation via Rao–Blackwellized gradient estimators.
Model class and Gibbs sampler
An LLnGM specifies a Gaussian observation layer Y∣W,V∼N(Xβ+AW,σϵ2I), a latent field whose precision is modulated by independent mixing variables Vi∼GIG(p,a,b) through an invertible structure matrix K(ζ), and a drift/skewness parameter μ. The GIG family spans the interior ΨI, the inverse-Gamma boundary ΨIG, and the Gamma boundary ΨΓ; as mixing distributions these induce the generalized hyperbolic family, including Student-t, normal-inverse Gaussian (NIG), and generalized asymmetric Laplace (GAL) marginals. Because the full conditionals of both the latent field and the mixing variables are available in closed form—Gaussian for W (or M) and GIG for Vi∼GIG(p,a,b)0—a two-block Gibbs sampler is available.
The paper introduces a non-centered parameterization via Vi∼GIG(p,a,b)1, which induces the same marginal likelihood and, crucially, the same Vi∼GIG(p,a,b)2-marginal transition kernel as the centered parameterization (proved via the bijective change of variables). This equivalence allows the ergodicity analysis to be carried out in whichever parameterization is convenient, while the non-centered form later resolves an integrability obstruction in the score functions.
Operator-theoretic route: trace-class and spectral gap
The Vi∼GIG(p,a,b)3-marginal chain has transition density Vi∼GIG(p,a,b)4, inducing a Markov operator Vi∼GIG(p,a,b)5 on Vi∼GIG(p,a,b)6. The paper shows Vi∼GIG(p,a,b)7 is a self-adjoint, positive semidefinite contraction with spectrum in Vi∼GIG(p,a,b)8 and simple eigenvalue at Vi∼GIG(p,a,b)9 (using Harris recurrence and strict positivity of K(ζ)0), so geometric ergodicity is equivalent to a spectral gap at K(ζ)1. By a criterion of Qin–Ročková-type, K(ζ)2 is trace-class if and only if K(ζ)3.
The main trace-class theorem establishes this integrability in two regimes: (i) K(ζ)4 with arbitrary K(ζ)5, and (ii) K(ζ)6 with K(ζ)7. The proof bounds the joint density by controlling Bessel-function ratios via explicit lower bounds on K(ζ)8, applying a Cauchy–Schwarz split and a Gaussian moment lemma to integrate out K(ζ)9, and then establishing an exponential exponent gap μ0 uniformly over sign vectors μ1. The exponential decay in both μ2 and μ3 dominates the polynomial prefactors, yielding integrability. As a corollary, the NIG model (μ4) is trace-class and geometrically ergodic for all drift values μ5—a clean, practically relevant guarantee since NIG is one of the two distributions closed under convolution limits used in the ngme2 software.
Drift–minorization route for boundary regimes
For regimes outside the trace-class theorem, the paper verifies a geometric drift condition with Lyapunov functions of the form μ6 together with a minorization condition on sub-level sets, invoking Rosenthal's drift–minorization theorem. Three cases are treated:
- Case I (μ7, μ8, μ9): the ΨI0-update is Inverse-Gamma, and a fractional-moment computation with ΨI1 yields a contractive drift coefficient ΨI2 via a convexity argument on Gamma-function ratios. Notably, the paper proves (Lemma on failure of trace-class) that in this regime, with ΨI3, the operator is not trace-class—so the drift route is not redundant but covers genuinely different chains.
- Cases II and III (ΨI4 and ΨI5): these require a null-smallness assumption: with ΨI6, either ΨI7 or ΨI8, where ΨI9 and ΨIG0 encode the projection of the drift direction onto the unidentifiable subspace. The condition is automatically satisfied when ΨIG1, and the paper gives structural sufficient conditions (orthogonality of the constant mode, structure preservation by ΨIG2, uniform non-degeneracy on the null space) under which the constant remains bounded as ΨIG3.
The drift proofs decompose the conditional mean ΨIG4 into null-space and range components; the null component contributes exactly the null-smallness ratio as the coefficient of ΨIG5, while the range component and Gaussian fluctuation terms are absorbed into constants. Negative moments of the GIG updates are controlled uniformly in ΨIG6 via Bessel-ratio bounds with a carefully chosen exponent ΨIG7 ensuring ΨIG8.
Combining both routes yields a complete map: geometric ergodicity holds across the full GIG parameter space, with trace-class known in the two regimes above, unknown in two boundary regimes (where geometric ergodicity still holds under the null-smallness condition), and provably absent in the ΨIG9 case. The null-smallness assumption alone is shown to be insufficient for trace-class but sufficient for geometric ergodicity.
Consequences for stochastic-gradient maximum likelihood
Geometric ergodicity is motivated by SGD-based maximum likelihood, as implemented in ngme2. Fisher's identity expresses the log-likelihood gradient as a posterior expectation of the complete-data score, approximated by ergodic averages of Gibbs samples. A Rao–Blackwellized estimator integrates out the Gaussian layer analytically given ΨΓ0, so the estimator depends only on the ΨΓ1-chain—aligning with the ergodicity results—and has reduced variance. The paper is careful to note that finite-length Gibbs runs introduce bias rather than exact unbiasedness; geometric ergodicity gives geometric bias decay ΨΓ2, which is summable under standard step-size schedules ΨΓ3, ΨΓ4, preserving almost-sure convergence of stochastic approximation.
A sharp integrability result identifies a genuine obstruction in the centered parameterization: the posterior expectation ΨΓ5 is finite if and only if ΨΓ6, where ΨΓ7 is the Gamma shape of the mixing variable. For GAL noise with ΨΓ8—and, critically, for SPDE-based models where mesh refinement drives the effective shape parameter ΨΓ9—the complete-data score fails to be integrable under the posterior, so ergodic theorems cannot be applied directly and SGD convergence cannot be established. This explains the instability observed previously in MCEM algorithms for such models. The non-centered parameterization resolves this structurally: its score functions depend on t0 only polynomially (the t1-score contains t2 and self-normalized quadratic forms rather than t3 terms), so integrability follows directly from the polynomial moment bounds supplied by the drift analysis. The non-centered parameterization is thus not merely a computational device but a necessary theoretical ingredient for rigorous SGD convergence.
Numerical evidence
Two simulation studies (t4, AR(1)-type t5, four overdispersed chains of t6 iterations) support the theory. The first compares mixing across six representative parameter points spanning all regimes. Trace-class interior regimes exhibit integrated autocorrelation times near 1 (ESS/sec up to roughly 105–120), while the t7, t8 regime shows the strongest degradation, with IACT rising to about 4.7 and 6.8 for the negative-moment and log summaries respectively—consistent with the boundary-sensitive Lyapunov functions in the drift analysis. Split-t9 values are approximately 1.0001 throughout. The second study fixes the design geometry with a one-dimensional null space (W0) and scans W1, varying the null-smallness constant W2 while holding W3 fixed. Increasing W4 produces systematic IACT inflation, most pronounced for the null-direction statistic and the log summary, providing empirical evidence that the null-smallness constant captures a genuine mixing mechanism rather than a proof artefact.
Limitations and open questions
The paper is explicit about several restrictions. Trace-class remains unresolved in the regimes W5 and W6; only geometric ergodicity (conditional on the null-smallness assumption) is established there. The null-smallness condition itself is a constraint on the interaction of W7, W8, and W9; the paper provides dimension-free sufficient conditions but does not characterize necessity, and the drift–minorization constants are conservative and do not yield sharp non-asymptotic Monte Carlo error bounds. The Rao–Blackwellized estimator is argued to behave better under the centered parameterization, but the paper favors changing the parameterization and does not fully analyze integrability of the Rao–Blackwellized centered score. Finally, the framework is tailored to conditionally Gaussian variance-mixture structures; nonlinear observation models or likelihoods without auxiliary-variable representations fall outside its scope, and the authors note that the proof strategy extends to other normal mean–variance mixtures only when comparable envelope bounds and moment control on the mixing density near M0 and M1 can be verified.
Conclusion
The paper delivers a complete characterization of geometric ergodicity for Gibbs samplers in GIG variance-mixture latent models, combining a trace-class/spectral-gap argument with a drift–minorization analysis governed by an interpretable null-smallness condition. These stability results, together with the non-centered parameterization's resolution of the score-integrability obstruction at shape parameters below M2, provide the theoretical foundation for Rao–Blackwellized stochastic-gradient maximum-likelihood estimation in this model class, with numerical experiments corroborating the predicted mixing behavior across parameter regimes and the practical relevance of the null-smallness constant.