- The paper develops a quantitative comparison framework that transfers weak Poincaré inequalities from exact Gibbs and conditional kernels to one- and two-block Metropolis-within-Gibbs samplers, with explicit composed rate functions.
- The paper proves that joint and marginal deterministic-scan Gibbs chains have matching weak Poincaré rate functions, enabling comparable subgeometric convergence guarantees and sharper marginal bounds for some samplers.
- The paper applies the theory to hierarchical models, regression, and discretely observed diffusions, demonstrating exponential, polynomial, and log-squared convergence rates driven by kernel properties and tuning parameters.
This paper develops a quantitative convergence theory for two-component deterministic-scan Metropolis-within-Gibbs (MwG) samplers via weak Poincaré inequalities (WPIs), extending the Dirichlet-form comparison framework to nonreversible chains and to subgeometric regimes where spectral gaps do not exist (2602.14692).
Setting and motivation
The deterministic-scan Gibbs (DG) kernel P=G1G2 on a product space X×Y is invariant but nonreversible with respect to the joint target Π, which obstructs spectral methods. The MwG variants replace one or both exact conditional updates by reversible kernels H1∣x, H2∣y targeting the full conditionals. Prior work on such samplers largely required geometric ergodicity or explicit spectral gaps (2602.14692); the authors instead assume only that (i) P∗P satisfies a WPI with function β0, and (ii) each conditional kernel satisfies a conditional WPI with functions β1(⋅,x), β2(⋅,y) — conditions described as extremely mild, holding under irreducibility.
Comparison machinery on the joint space
The core technical device is an optimized K∗-formulation of WPIs together with a decomposition of Dirichlet forms for compositions of self-adjoint kernels: X×Y0. A four-step comparison path — swap components, replace one conditional update, swap back, replace the other — yields explicit X×Y1-WPIs for the intermediate 1-MG (X×Y2), 2-MG (X×Y3), and full MG (X×Y4) samplers. The central result is that under the three assumptions above,
X×Y5
i.e., the rate function of the MwG sampler is a composition of those of the exact Gibbs kernel and the conditional kernels, up to constant factors absorbed into the final bound via scaling lemmas. A key enabling lemma shows that if X×Y6 satisfies a X×Y7-WPI then X×Y8 does with X×Y9; this is what permits comparison between a kernel and its adjoint in the subgeometric setting, addressing a gap the authors identify in prior spectral-gap decomposition work where adjoint compositions have unequal Dirichlet forms. In the geometric case, when Π0 has spectral gap Π1, the sharper bound Π2 follows.
Joint versus marginal chains
For the two-component DG kernel, the paper proves that the Π3-marginal chain Π4 and the joint chain Π5 satisfy WPIs with identical Π6-functions in both directions: a WPI for Π7 transfers verbatim to Π8, while a WPI for Π9 yields H1∣x0 for the joint chain. Consequently joint and marginal chains converge at comparable rates even in subgeometric regimes, extending classical operator-norm equivalences valid in the geometric case. Analogous transfer results hold for the 2-MG sampler and its marginal H1∣x1, and a marginal-level comparison between H1∣x2 and H1∣x3 produces a strictly better bound for H1∣x4: H1∣x5 versus H1∣x6 from the joint-space route. The authors note explicitly that this marginal route does not yield a WPI for H1∣x7 itself, so it cannot be chained further toward the full MG sampler.
Tensorization
Viewing independent chains as MwG samplers under conditional independence, the paper establishes a tensorization theorem: if H1∣x8 and H1∣x9 are independent, reversible, positive chains whose squared kernels satisfy WPIs with H2∣y0, then H2∣y1 satisfies a WPI with H2∣y2, generalizing to H2∣y3 chains by induction and recovering, in continuous-time analogy and in the spectral-gap case, the familiar minimum-of-gaps bound. This extends known discrete-time spectral product results to chains without spectral gaps.
Applications
Three concrete settings illustrate the theory:
Normal–inverse-gamma model with RWM inner kernels. With conditionally scaled step sizes (H2∣y4, H2∣y5), all spectral gaps are uniformly bounded and the MG chain converges exponentially at rate at least H2∣y6. With fixed step sizes H2∣y7, only polynomial rates are obtained, e.g. H2∣y8 when H2∣y9. The contrast makes precise how step-size scaling relative to the conditional scale determines geometric versus subgeometric behavior.
Bayesian hierarchical linear regression. For a Gaussian likelihood with flat prior on P∗P0 and P∗P1, the block Gibbs kernel admits a SPI (via an P∗P2-geometric drift argument on the P∗P3-marginal), while RWM updating of P∗P4 gives a 2-MG polynomial bound P∗P5, where the exponents depend explicitly on the data dimension P∗P6, parameter dimension P∗P7, and the RWM step size through P∗P8.
Discretely-observed diffusions. For a data augmentation scheme with IMH path updates, the Girsanov functional's upper bound yields indicator-type WPI functions P∗P9, and the framework produces explicit bounds; notably these extend prior scalability results by removing a restrictive compact-support assumption on the prior. Under uniform boundedness of β00, a SPI holds for the exact Gibbs kernel. For the Ornstein–Uhlenbeck process — whose linear drift violates the boundedness condition, so the authors supply a separate argument via a β01-independent minorization — the resulting bound is of log-squared type: β02.
Limitations and open questions
The paper concedes several limitations. The joint-space route introduces constant-factor losses (the factors 2 and 4 in β03), which may yield loose bounds, though asymptotically the tail behavior of the rate is preserved. The authors state it is not clear which of the two admissible compositions of β04 performs better across scenarios. The marginal-based improvement applies only to the 2-MG sampler and provides no WPI for β05, hence no direct chaining to β06; whether a sharper joint-space construction exists remains open. All bounds require β07, restricting the class of test functions. Finally, the diffusion SPI requires uniform boundedness conditions on the drift that exclude models such as the OU process, for which case-specific arguments were needed.
Conclusion
The paper delivers a complete, explicit framework transferring WPI-based convergence rates from exact Gibbs samplers and conditional kernels to their MwG counterparts, handles nonreversibility through a new adjoint-comparison lemma, establishes equivalence of joint and marginal convergence in subgeometric regimes, and proves tensorization of WPIs for independent chains. The applications demonstrate that the resulting bounds capture both exponential and genuinely subgeometric behavior with dependence on tuning parameters made explicit.