Papers
Topics
Authors
Recent
Search
2000 character limit reached

Hyper-V uniform ergodicity of Markov chains

Published 13 Aug 2026 in math.ST and stat.CO | (2608.12738v1)

Abstract: We develop a new uniform drift condition and local minorization that implies a stronger weighted form of uniform ergodicity for Markov chains we call hyper-V uniform ergodicity. The convergence guarantees geometric decay of the bias towards the invariant measure independently of the initialization for all functions controlled by a dominating function V. A key advantage of the approach is that it bypasses the need to establish a global minorization condition, which is often substantially more difficult to verify in practice, while yielding stronger convergence guarantees than global minorization. Optimal convergence bounds in a minimax sense of the framework are established. The utility of the framework is demonstrated through applications to the P'olya-Gamma and Kolmogorov-Gamma Gibbs samplers. We also show qualitative hyper-V uniform ergodicity convergence for two-variable Gibbs samplers can be inferred by the form of the invariant measure, bypassing convergence analysis entirely.

Authors (2)

Summary

  • The paper proves that a uniform drift condition combined with local minorization produces hyper-V uniform ergodicity, with geometric convergence for total variation and functions controlled by a Lyapunov function.
  • The paper derives explicit and minimax-optimal convergence rates within the drift–minorization framework, including Wasserstein bounds that require no kernel contractivity or Lipschitz assumptions.
  • The paper applies the theory to Pólya–Gamma and Kolmogorov–Gamma Gibbs samplers, establishing stronger weighted and Wasserstein guarantees while identifying limitations for independence Metropolis–Hastings chains.

This paper develops a framework for establishing a strengthened form of uniform ergodicity for Markov chains, termed hyper-VV uniform ergodicity, from two structurally simple assumptions: a uniform drift condition and a local minorization condition (2608.12738). The central contribution is that these assumptions, which require only local control of the transition kernel on sublevel sets of a Lyapunov function VV, suffice to recover—and strengthen—the guarantees classically obtained from a global (Doeblin) minorization condition. The framework yields explicit convergence rates, minimax optimality results within the proposed class of kernels, Wasserstein extensions, and applications to the Pólya–Gamma and Kolmogorov–Gamma Gibbs samplers.

From global minorization to drift plus local minorization

Classical uniform ergodicity requires a global minorization condition infxXP(x,)αν()\inf_{x \in X} P(x,\cdot) \ge \alpha \nu(\cdot), which enforces uniform mixing across the entire state space. Verifying such a condition demands uniform control of the kernel and is often difficult on unbounded or high-dimensional spaces; moreover, the resulting bounds primarily control bounded test functions.

The paper replaces this with:

  • Uniform drift: PV(x)KPV(x) \le K for all xXx \in X, for some K<K < \infty.
  • Local minorization: inf{x:V(x)rK}P(x,)αrν()\inf_{\{x : V(x) \le rK\}} P(x,\cdot) \ge \alpha_r \nu(\cdot) for some r>1r > 1.

The mechanism is elementary but effective: by Markov's inequality, P(x,{V>rK})1/rP(x,\{V > rK\}) \le 1/r, so every one-step transition places mass at least $1 - 1/r$ on the sublevel set where minorization holds. Consequently the two-step kernel satisfies a global minorization with constant VV0, which drives all subsequent bounds.

Hyper-V uniform ergodicity

The main theorem establishes, for any nondecreasing concave VV1,

VV2

and, for functions dominated by VV3,

VV4

The second bound is the defining feature of hyper-VV5 uniform ergodicity: geometric decay of bias uniformly over unbounded test functions controlled by VV6, independent of initialization. The proof combines the induced two-step global minorization with an oscillation contraction argument, using Jensen's inequality and concavity of VV7 to bound VV8 via VV9.

Compared with specializing Hairer and Mattingly's Harris-type theorem to this setting, the paper identifies two advantages. First, the multiplicative constant for test functions satisfying infxXP(x,)αν()\inf_{x \in X} P(x,\cdot) \ge \alpha \nu(\cdot)0 is infxXP(x,)αν()\inf_{x \in X} P(x,\cdot) \ge \alpha \nu(\cdot)1, versus constants of order at least infxXP(x,)αν()\inf_{x \in X} P(x,\cdot) \ge \alpha \nu(\cdot)2 under the classical approach, which can be very large when the minorization constant is small. Second, when the minorization constant scales as infxXP(x,)αν()\inf_{x \in X} P(x,\cdot) \ge \alpha \nu(\cdot)3 with infxXP(x,)αν()\inf_{x \in X} P(x,\cdot) \ge \alpha \nu(\cdot)4—a regime the authors argue is typical in statistical applications, where minorization scaling is often exponentially poor—the rate from the new theorem is strictly smaller than the optimized Hairer–Mattingly rate. This comparison depends on the assumption that infxXP(x,)αν()\inf_{x \in X} P(x,\cdot) \ge \alpha \nu(\cdot)5, which holds for power-law or faster decay of infxXP(x,)αν()\inf_{x \in X} P(x,\cdot) \ge \alpha \nu(\cdot)6 in infxXP(x,)αν()\inf_{x \in X} P(x,\cdot) \ge \alpha \nu(\cdot)7.

Minimax optimal rates

A refined analysis tracks the interaction between visits to the minorization region infxXP(x,)αν()\inf_{x \in X} P(x,\cdot) \ge \alpha \nu(\cdot)8 and its complement, lifting the dynamics to a two-dimensional system of Doeblin-type distances governed by a nonnegative matrix infxXP(x,)αν()\inf_{x \in X} P(x,\cdot) \ge \alpha \nu(\cdot)9. Solving the resulting linear recursion via Cayley–Hamilton yields an explicit bound PV(x)KPV(x) \le K0 whose asymptotic rate is the Perron root PV(x)KPV(x) \le K1, where PV(x)KPV(x) \le K2. Because the minorization set here (PV(x)KPV(x) \le K3) is smaller than the sets required in prior lifting-based analyses, the rate improves whenever PV(x)KPV(x) \le K4.

The paper then proves these bounds are minimax optimal over the class PV(x)KPV(x) \le K5 of all kernels satisfying the uniform drift and minorization conditions: both the finite-time upper bound and the asymptotic rate PV(x)KPV(x) \le K6 are attained. The lower-bound construction uses a five-state chain in which excursions away from the absorbing state mimic the matrix recursion exactly. This optimality statement is relative to the chosen framework—it does not claim optimality among all kernels admitting uniform ergodicity by other means—but it shows no sharper bound is possible using only the drift and minorization parameters PV(x)KPV(x) \le K7.

A corollary transfers the weighted bounds to Wasserstein distances: if PV(x)KPV(x) \le K8 for some metric PV(x)KPV(x) \le K9 making xXx \in X0 complete and separable, then xXx \in X1. Notably, this requires no contractivity or Lipschitz assumptions on the kernel, in contrast to standard Wasserstein ergodicity results based on coupling contractions. The authors concede that in settings where no local minorization exists, the argument must be extended to local Wasserstein contraction, which they leave open.

Applications to Gibbs samplers

Pólya–Gamma sampler. For Bayesian logistic regression with Gaussian prior, the marginal chain on xXx \in X2 is shown to be hyper-xXx \in X3 uniformly ergodic for xXx \in X4. The drift constant is explicit,

xXx \in X5

obtained from xXx \in X6. The local minorization follows from a density comparison of Pólya–Gamma densities at xXx \in X7 versus xXx \in X8 evaluated at the boundary xXx \in X9, giving K<K < \infty0. The result yields uniform total variation convergence, uniform control of expectations of functions bounded by K<K < \infty1, and—apparently for the first time for this model—uniform convergence in K<K < \infty2-Wasserstein distance. Simulation studies with K<K < \infty3, K<K < \infty4 show that K<K < \infty5 decays rapidly in K<K < \infty6, indicating substantial practical benefit from the sharper rates.

Kolmogorov–Gamma sampler. An analogous result holds for regression with continuous proportion data, with K<K < \infty7, exploiting monotonicity of K<K < \infty8. Again, hyper-K<K < \infty9 uniform ergodicity and Wasserstein convergence are established with substantially less effort than the global minorization arguments used previously.

Independence Metropolis–Hastings. This example delineates the framework's limits. For the IMH kernel with proposal inf{x:V(x)rK}P(x,)αrν()\inf_{\{x : V(x) \le rK\}} P(x,\cdot) \ge \alpha_r \nu(\cdot)0, taking inf{x:V(x)rK}P(x,)αrν()\inf_{\{x : V(x) \le rK\}} P(x,\cdot) \ge \alpha_r \nu(\cdot)1 gives a valid uniform drift condition inf{x:V(x)rK}P(x,)αrν()\inf_{\{x : V(x) \le rK\}} P(x,\cdot) \ge \alpha_r \nu(\cdot)2 (proved via inf{x:V(x)rK}P(x,)αrν()\inf_{\{x : V(x) \le rK\}} P(x,\cdot) \ge \alpha_r \nu(\cdot)3). However, any local minorization on sublevel sets forces inf{x:V(x)rK}P(x,)αrν()\inf_{\{x : V(x) \le rK\}} P(x,\cdot) \ge \alpha_r \nu(\cdot)4—the same condition required for a global minorization. In a toy example with exponential target and tilted exponential proposal, the drift condition additionally fails for inf{x:V(x)rK}P(x,)αrν()\inf_{\{x : V(x) \le rK\}} P(x,\cdot) \ge \alpha_r \nu(\cdot)5 even though global minorization holds for all inf{x:V(x)rK}P(x,)αrν()\inf_{\{x : V(x) \le rK\}} P(x,\cdot) \ge \alpha_r \nu(\cdot)6. Thus the framework does not always simplify verification; its value here lies in yielding the stronger weighted conclusion rather than easier proofs.

Qualitative convergence from the invariant measure's form

For two-variable data-augmentation Gibbs samplers whose invariant measure has the Gaussian conditional form

inf{x:V(x)rK}P(x,)αrν()\inf_{\{x : V(x) \le rK\}} P(x,\cdot) \ge \alpha_r \nu(\cdot)7

the paper shows that qualitative hyper-inf{x:V(x)rK}P(x,)αrν()\inf_{\{x : V(x) \le rK\}} P(x,\cdot) \ge \alpha_r \nu(\cdot)8 uniform ergodicity (with inf{x:V(x)rK}P(x,)αrν()\inf_{\{x : V(x) \le rK\}} P(x,\cdot) \ge \alpha_r \nu(\cdot)9) follows directly from finiteness and continuity of the normalizing function r>1r > 10 and boundedness of r>1r > 11—no explicit convergence analysis of the chain is required. Compactness of sublevel sets plus continuity of the conditional density produce the local minorization automatically. Concrete instances include the Pólya–Gamma and Kolmogorov–Gamma posteriors (where r>1r > 12 is uniformly bounded) and Bayesian multivariate linear regression with inverse-Wishart priors, where boundedness of r>1r > 13 reduces to a bound involving the maximum likelihood estimator r>1r > 14. These conclusions are qualitative only—they assert existence of a rate without computing it—and the sufficient condition r>1r > 15 may fail in models with unbounded sufficient statistics.

Limitations and open questions

Several caveats bear directly on the results. The optimality claims hold within the class of kernels characterized solely by r>1r > 16; kernels admitting better rates through additional structure are not ruled out. The improvement over Hairer–Mattingly rates hinges on the scaling assumption r>1r > 17, which is argued to be typical but is not verified in general. The IMH example shows the drift-plus-local-minorization route can be strictly harder than global minorization, so the framework is not universally applicable. The Wasserstein extension presupposes existence of a minorization condition; replacing it with local Wasserstein contraction remains open, as does a fuller characterization of the consequences of hyper-r>1r > 18 uniform ergodicity in transport metrics. Finally, whether a global minorization condition can ever be inferred from the form of the invariant measure alone—as the qualitative results do for hyper-r>1r > 19 uniform ergodicity—is left unresolved.

Conclusion

The paper establishes that a uniform drift condition paired with a local minorization implies hyper-P(x,{V>rK})1/rP(x,\{V > rK\}) \le 1/r0 uniform ergodicity with explicit, minimax optimal rates within the proposed framework, bypassing global minorization while strengthening the conclusion from total variation to weighted and Wasserstein control. Applications to the Pólya–Gamma and Kolmogorov–Gamma Gibbs samplers demonstrate both simplified verification and strictly stronger guarantees than previously available, while the independence Metropolis–Hastings analysis honestly delimits the scope of the method. The qualitative results linking convergence to the algebraic form of data-augmented posteriors offer a practical shortcut for applied convergence analysis, at the cost of forgoing explicit rates (2608.12738).

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.