Papers
Topics
Authors
Recent
Search
2000 character limit reached

Weak Poincaré inequalities for Deterministic-scan Metropolis-within-Gibbs samplers

Published 16 Feb 2026 in stat.CO and math.PR | (2602.14692v1)

Abstract: Using the framework of weak Poincaré inequalities, we analyze the convergence properties of deterministic-scan Metropolis-within-Gibbs samplers, an important class of Markov chain Monte Carlo algorithms. Our analysis applies to nonreversible Markov chains and yields explicit (subgeometric) convergence bounds through novel comparison techniques based on Dirichlet forms. We show that the joint chain inherits the convergence behavior of the marginal chain and conversely. In addition, we establish several fundamental results for weak Poincaré inequalities for discrete-time Markov chains, such as a tensorization property for independent chains. We apply our theoretical results through applications to algorithms for Bayesian inference for a hierarchical regression model and a diffusion model under discretely-observed data.

Summary

  • The paper develops a quantitative comparison framework that transfers weak Poincaré inequalities from exact Gibbs and conditional kernels to one- and two-block Metropolis-within-Gibbs samplers, with explicit composed rate functions.
  • The paper proves that joint and marginal deterministic-scan Gibbs chains have matching weak Poincaré rate functions, enabling comparable subgeometric convergence guarantees and sharper marginal bounds for some samplers.
  • The paper applies the theory to hierarchical models, regression, and discretely observed diffusions, demonstrating exponential, polynomial, and log-squared convergence rates driven by kernel properties and tuning parameters.

This paper develops a quantitative convergence theory for two-component deterministic-scan Metropolis-within-Gibbs (MwG) samplers via weak Poincaré inequalities (WPIs), extending the Dirichlet-form comparison framework to nonreversible chains and to subgeometric regimes where spectral gaps do not exist (2602.14692).

Setting and motivation

The deterministic-scan Gibbs (DG) kernel P=G1G2P = G_1G_2 on a product space X×Y\mathcal{X}\times\mathcal{Y} is invariant but nonreversible with respect to the joint target Π\Pi, which obstructs spectral methods. The MwG variants replace one or both exact conditional updates by reversible kernels H1xH_{1|x}, H2yH_{2|y} targeting the full conditionals. Prior work on such samplers largely required geometric ergodicity or explicit spectral gaps (2602.14692); the authors instead assume only that (i) PPP^*P satisfies a WPI with function β0\beta_0, and (ii) each conditional kernel satisfies a conditional WPI with functions β1(,x)\beta_1(\cdot,x), β2(,y)\beta_2(\cdot,y) — conditions described as extremely mild, holding under irreducibility.

Comparison machinery on the joint space

The core technical device is an optimized KK^*-formulation of WPIs together with a decomposition of Dirichlet forms for compositions of self-adjoint kernels: X×Y\mathcal{X}\times\mathcal{Y}0. A four-step comparison path — swap components, replace one conditional update, swap back, replace the other — yields explicit X×Y\mathcal{X}\times\mathcal{Y}1-WPIs for the intermediate 1-MG (X×Y\mathcal{X}\times\mathcal{Y}2), 2-MG (X×Y\mathcal{X}\times\mathcal{Y}3), and full MG (X×Y\mathcal{X}\times\mathcal{Y}4) samplers. The central result is that under the three assumptions above,

X×Y\mathcal{X}\times\mathcal{Y}5

i.e., the rate function of the MwG sampler is a composition of those of the exact Gibbs kernel and the conditional kernels, up to constant factors absorbed into the final bound via scaling lemmas. A key enabling lemma shows that if X×Y\mathcal{X}\times\mathcal{Y}6 satisfies a X×Y\mathcal{X}\times\mathcal{Y}7-WPI then X×Y\mathcal{X}\times\mathcal{Y}8 does with X×Y\mathcal{X}\times\mathcal{Y}9; this is what permits comparison between a kernel and its adjoint in the subgeometric setting, addressing a gap the authors identify in prior spectral-gap decomposition work where adjoint compositions have unequal Dirichlet forms. In the geometric case, when Π\Pi0 has spectral gap Π\Pi1, the sharper bound Π\Pi2 follows.

Joint versus marginal chains

For the two-component DG kernel, the paper proves that the Π\Pi3-marginal chain Π\Pi4 and the joint chain Π\Pi5 satisfy WPIs with identical Π\Pi6-functions in both directions: a WPI for Π\Pi7 transfers verbatim to Π\Pi8, while a WPI for Π\Pi9 yields H1xH_{1|x}0 for the joint chain. Consequently joint and marginal chains converge at comparable rates even in subgeometric regimes, extending classical operator-norm equivalences valid in the geometric case. Analogous transfer results hold for the 2-MG sampler and its marginal H1xH_{1|x}1, and a marginal-level comparison between H1xH_{1|x}2 and H1xH_{1|x}3 produces a strictly better bound for H1xH_{1|x}4: H1xH_{1|x}5 versus H1xH_{1|x}6 from the joint-space route. The authors note explicitly that this marginal route does not yield a WPI for H1xH_{1|x}7 itself, so it cannot be chained further toward the full MG sampler.

Tensorization

Viewing independent chains as MwG samplers under conditional independence, the paper establishes a tensorization theorem: if H1xH_{1|x}8 and H1xH_{1|x}9 are independent, reversible, positive chains whose squared kernels satisfy WPIs with H2yH_{2|y}0, then H2yH_{2|y}1 satisfies a WPI with H2yH_{2|y}2, generalizing to H2yH_{2|y}3 chains by induction and recovering, in continuous-time analogy and in the spectral-gap case, the familiar minimum-of-gaps bound. This extends known discrete-time spectral product results to chains without spectral gaps.

Applications

Three concrete settings illustrate the theory:

Normal–inverse-gamma model with RWM inner kernels. With conditionally scaled step sizes (H2yH_{2|y}4, H2yH_{2|y}5), all spectral gaps are uniformly bounded and the MG chain converges exponentially at rate at least H2yH_{2|y}6. With fixed step sizes H2yH_{2|y}7, only polynomial rates are obtained, e.g. H2yH_{2|y}8 when H2yH_{2|y}9. The contrast makes precise how step-size scaling relative to the conditional scale determines geometric versus subgeometric behavior.

Bayesian hierarchical linear regression. For a Gaussian likelihood with flat prior on PPP^*P0 and PPP^*P1, the block Gibbs kernel admits a SPI (via an PPP^*P2-geometric drift argument on the PPP^*P3-marginal), while RWM updating of PPP^*P4 gives a 2-MG polynomial bound PPP^*P5, where the exponents depend explicitly on the data dimension PPP^*P6, parameter dimension PPP^*P7, and the RWM step size through PPP^*P8.

Discretely-observed diffusions. For a data augmentation scheme with IMH path updates, the Girsanov functional's upper bound yields indicator-type WPI functions PPP^*P9, and the framework produces explicit bounds; notably these extend prior scalability results by removing a restrictive compact-support assumption on the prior. Under uniform boundedness of β0\beta_00, a SPI holds for the exact Gibbs kernel. For the Ornstein–Uhlenbeck process — whose linear drift violates the boundedness condition, so the authors supply a separate argument via a β0\beta_01-independent minorization — the resulting bound is of log-squared type: β0\beta_02.

Limitations and open questions

The paper concedes several limitations. The joint-space route introduces constant-factor losses (the factors 2 and 4 in β0\beta_03), which may yield loose bounds, though asymptotically the tail behavior of the rate is preserved. The authors state it is not clear which of the two admissible compositions of β0\beta_04 performs better across scenarios. The marginal-based improvement applies only to the 2-MG sampler and provides no WPI for β0\beta_05, hence no direct chaining to β0\beta_06; whether a sharper joint-space construction exists remains open. All bounds require β0\beta_07, restricting the class of test functions. Finally, the diffusion SPI requires uniform boundedness conditions on the drift that exclude models such as the OU process, for which case-specific arguments were needed.

Conclusion

The paper delivers a complete, explicit framework transferring WPI-based convergence rates from exact Gibbs samplers and conditional kernels to their MwG counterparts, handles nonreversibility through a new adjoint-comparison lemma, establishes equivalence of joint and marginal convergence in subgeometric regimes, and proves tensorization of WPIs for independent chains. The applications demonstrate that the resulting bounds capture both exponential and genuinely subgeometric behavior with dependence on tuning parameters made explicit.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.