QTD under asynchronous and Markovian sampling

Analyze quantile temporal-difference learning under asynchronous and Markovian sampling and determine finite-sample behavior comparable to the synchronous generative-model setting.

Background

The established result concerns synchronous QTD with samples generated independently for each state at each iteration. The conclusion identifies asynchronous updates and Markovian sampling as unresolved extensions, which would require handling update delays, temporal dependence, or both while preserving the global localization and variance-sensitive local concentration properties used in the paper.

References

Several extensions remain open. It would be useful to obtain matching lower bounds for the global burn-in, to study asynchronous and Markovian sampling, and to determine which parts of the positive-semigroup argument survive under function approximation.

A Finite Sample Analysis for Quantile Temporal Difference Learning in Distributional Reinforcement Learning  (2608.27313 - Cheng et al., 27 Aug 2026) in Section Conclusions

Another direction is to compare raw last iterates with Polyak--Ruppert averages at finite time: averaging improves the rate in $T$ but can introduce an additional inverse-Jacobian factor and hence a different dependence on extreme-quantile densities.

A Finite Sample Analysis for Quantile Temporal Difference Learning in Distributional Reinforcement Learning  (2608.27313 - Cheng et al., 27 Aug 2026) in Section Conclusions