- The paper proves that a uniform drift condition combined with local minorization produces hyper-V uniform ergodicity, with geometric convergence for total variation and functions controlled by a Lyapunov function.
- The paper derives explicit and minimax-optimal convergence rates within the drift–minorization framework, including Wasserstein bounds that require no kernel contractivity or Lipschitz assumptions.
- The paper applies the theory to Pólya–Gamma and Kolmogorov–Gamma Gibbs samplers, establishing stronger weighted and Wasserstein guarantees while identifying limitations for independence Metropolis–Hastings chains.
This paper develops a framework for establishing a strengthened form of uniform ergodicity for Markov chains, termed hyper-V uniform ergodicity, from two structurally simple assumptions: a uniform drift condition and a local minorization condition (2608.12738). The central contribution is that these assumptions, which require only local control of the transition kernel on sublevel sets of a Lyapunov function V, suffice to recover—and strengthen—the guarantees classically obtained from a global (Doeblin) minorization condition. The framework yields explicit convergence rates, minimax optimality results within the proposed class of kernels, Wasserstein extensions, and applications to the Pólya–Gamma and Kolmogorov–Gamma Gibbs samplers.
From global minorization to drift plus local minorization
Classical uniform ergodicity requires a global minorization condition infx∈XP(x,⋅)≥αν(⋅), which enforces uniform mixing across the entire state space. Verifying such a condition demands uniform control of the kernel and is often difficult on unbounded or high-dimensional spaces; moreover, the resulting bounds primarily control bounded test functions.
The paper replaces this with:
- Uniform drift: PV(x)≤K for all x∈X, for some K<∞.
- Local minorization: {x:V(x)≤rK}infP(x,⋅)≥αrν(⋅) for some r>1.
The mechanism is elementary but effective: by Markov's inequality, P(x,{V>rK})≤1/r, so every one-step transition places mass at least $1 - 1/r$ on the sublevel set where minorization holds. Consequently the two-step kernel satisfies a global minorization with constant V0, which drives all subsequent bounds.
The main theorem establishes, for any nondecreasing concave V1,
V2
and, for functions dominated by V3,
V4
The second bound is the defining feature of hyper-V5 uniform ergodicity: geometric decay of bias uniformly over unbounded test functions controlled by V6, independent of initialization. The proof combines the induced two-step global minorization with an oscillation contraction argument, using Jensen's inequality and concavity of V7 to bound V8 via V9.
Compared with specializing Hairer and Mattingly's Harris-type theorem to this setting, the paper identifies two advantages. First, the multiplicative constant for test functions satisfying infx∈XP(x,⋅)≥αν(⋅)0 is infx∈XP(x,⋅)≥αν(⋅)1, versus constants of order at least infx∈XP(x,⋅)≥αν(⋅)2 under the classical approach, which can be very large when the minorization constant is small. Second, when the minorization constant scales as infx∈XP(x,⋅)≥αν(⋅)3 with infx∈XP(x,⋅)≥αν(⋅)4—a regime the authors argue is typical in statistical applications, where minorization scaling is often exponentially poor—the rate from the new theorem is strictly smaller than the optimized Hairer–Mattingly rate. This comparison depends on the assumption that infx∈XP(x,⋅)≥αν(⋅)5, which holds for power-law or faster decay of infx∈XP(x,⋅)≥αν(⋅)6 in infx∈XP(x,⋅)≥αν(⋅)7.
Minimax optimal rates
A refined analysis tracks the interaction between visits to the minorization region infx∈XP(x,⋅)≥αν(⋅)8 and its complement, lifting the dynamics to a two-dimensional system of Doeblin-type distances governed by a nonnegative matrix infx∈XP(x,⋅)≥αν(⋅)9. Solving the resulting linear recursion via Cayley–Hamilton yields an explicit bound PV(x)≤K0 whose asymptotic rate is the Perron root PV(x)≤K1, where PV(x)≤K2. Because the minorization set here (PV(x)≤K3) is smaller than the sets required in prior lifting-based analyses, the rate improves whenever PV(x)≤K4.
The paper then proves these bounds are minimax optimal over the class PV(x)≤K5 of all kernels satisfying the uniform drift and minorization conditions: both the finite-time upper bound and the asymptotic rate PV(x)≤K6 are attained. The lower-bound construction uses a five-state chain in which excursions away from the absorbing state mimic the matrix recursion exactly. This optimality statement is relative to the chosen framework—it does not claim optimality among all kernels admitting uniform ergodicity by other means—but it shows no sharper bound is possible using only the drift and minorization parameters PV(x)≤K7.
A corollary transfers the weighted bounds to Wasserstein distances: if PV(x)≤K8 for some metric PV(x)≤K9 making x∈X0 complete and separable, then x∈X1. Notably, this requires no contractivity or Lipschitz assumptions on the kernel, in contrast to standard Wasserstein ergodicity results based on coupling contractions. The authors concede that in settings where no local minorization exists, the argument must be extended to local Wasserstein contraction, which they leave open.
Applications to Gibbs samplers
Pólya–Gamma sampler. For Bayesian logistic regression with Gaussian prior, the marginal chain on x∈X2 is shown to be hyper-x∈X3 uniformly ergodic for x∈X4. The drift constant is explicit,
x∈X5
obtained from x∈X6. The local minorization follows from a density comparison of Pólya–Gamma densities at x∈X7 versus x∈X8 evaluated at the boundary x∈X9, giving K<∞0. The result yields uniform total variation convergence, uniform control of expectations of functions bounded by K<∞1, and—apparently for the first time for this model—uniform convergence in K<∞2-Wasserstein distance. Simulation studies with K<∞3, K<∞4 show that K<∞5 decays rapidly in K<∞6, indicating substantial practical benefit from the sharper rates.
Kolmogorov–Gamma sampler. An analogous result holds for regression with continuous proportion data, with K<∞7, exploiting monotonicity of K<∞8. Again, hyper-K<∞9 uniform ergodicity and Wasserstein convergence are established with substantially less effort than the global minorization arguments used previously.
Independence Metropolis–Hastings. This example delineates the framework's limits. For the IMH kernel with proposal {x:V(x)≤rK}infP(x,⋅)≥αrν(⋅)0, taking {x:V(x)≤rK}infP(x,⋅)≥αrν(⋅)1 gives a valid uniform drift condition {x:V(x)≤rK}infP(x,⋅)≥αrν(⋅)2 (proved via {x:V(x)≤rK}infP(x,⋅)≥αrν(⋅)3). However, any local minorization on sublevel sets forces {x:V(x)≤rK}infP(x,⋅)≥αrν(⋅)4—the same condition required for a global minorization. In a toy example with exponential target and tilted exponential proposal, the drift condition additionally fails for {x:V(x)≤rK}infP(x,⋅)≥αrν(⋅)5 even though global minorization holds for all {x:V(x)≤rK}infP(x,⋅)≥αrν(⋅)6. Thus the framework does not always simplify verification; its value here lies in yielding the stronger weighted conclusion rather than easier proofs.
For two-variable data-augmentation Gibbs samplers whose invariant measure has the Gaussian conditional form
{x:V(x)≤rK}infP(x,⋅)≥αrν(⋅)7
the paper shows that qualitative hyper-{x:V(x)≤rK}infP(x,⋅)≥αrν(⋅)8 uniform ergodicity (with {x:V(x)≤rK}infP(x,⋅)≥αrν(⋅)9) follows directly from finiteness and continuity of the normalizing function r>10 and boundedness of r>11—no explicit convergence analysis of the chain is required. Compactness of sublevel sets plus continuity of the conditional density produce the local minorization automatically. Concrete instances include the Pólya–Gamma and Kolmogorov–Gamma posteriors (where r>12 is uniformly bounded) and Bayesian multivariate linear regression with inverse-Wishart priors, where boundedness of r>13 reduces to a bound involving the maximum likelihood estimator r>14. These conclusions are qualitative only—they assert existence of a rate without computing it—and the sufficient condition r>15 may fail in models with unbounded sufficient statistics.
Limitations and open questions
Several caveats bear directly on the results. The optimality claims hold within the class of kernels characterized solely by r>16; kernels admitting better rates through additional structure are not ruled out. The improvement over Hairer–Mattingly rates hinges on the scaling assumption r>17, which is argued to be typical but is not verified in general. The IMH example shows the drift-plus-local-minorization route can be strictly harder than global minorization, so the framework is not universally applicable. The Wasserstein extension presupposes existence of a minorization condition; replacing it with local Wasserstein contraction remains open, as does a fuller characterization of the consequences of hyper-r>18 uniform ergodicity in transport metrics. Finally, whether a global minorization condition can ever be inferred from the form of the invariant measure alone—as the qualitative results do for hyper-r>19 uniform ergodicity—is left unresolved.
Conclusion
The paper establishes that a uniform drift condition paired with a local minorization implies hyper-P(x,{V>rK})≤1/r0 uniform ergodicity with explicit, minimax optimal rates within the proposed framework, bypassing global minorization while strengthening the conclusion from total variation to weighted and Wasserstein control. Applications to the Pólya–Gamma and Kolmogorov–Gamma Gibbs samplers demonstrate both simplified verification and strictly stronger guarantees than previously available, while the independence Metropolis–Hastings analysis honestly delimits the scope of the method. The qualitative results linking convergence to the algebraic form of data-augmented posteriors offer a practical shortcut for applied convergence analysis, at the cost of forgoing explicit rates (2608.12738).