Matching lower bounds for global burn-in

Establish matching lower bounds for the global burn-in required by synchronous quantile temporal-difference learning in tabular distributional reinforcement learning.

Background

The paper proves a global high-probability last-iterate bound for synchronous quantile temporal-difference learning (QTD). Its analysis separates a global localization phase from a local stochastic-concentration phase, and shows that the deterministic transient and localization time can depend on the smallest Bellman-target density, which may deteriorate with the number of quantiles. The authors explicitly leave unresolved whether this burn-in dependence is unavoidable by requesting matching lower bounds.

References

Several extensions remain open. It would be useful to obtain matching lower bounds for the global burn-in, to study asynchronous and Markovian sampling, and to determine which parts of the positive-semigroup argument survive under function approximation.

A Finite Sample Analysis for Quantile Temporal Difference Learning in Distributional Reinforcement Learning  (2608.27313 - Cheng et al., 27 Aug 2026) in Section Conclusions