QTD under asynchronous and Markovian sampling
Analyze quantile temporal-difference learning under asynchronous and Markovian sampling and determine finite-sample behavior comparable to the synchronous generative-model setting.
References
Several extensions remain open. It would be useful to obtain matching lower bounds for the global burn-in, to study asynchronous and Markovian sampling, and to determine which parts of the positive-semigroup argument survive under function approximation.
— A Finite Sample Analysis for Quantile Temporal Difference Learning in Distributional Reinforcement Learning
(2608.27313 - Cheng et al., 27 Aug 2026) in Section Conclusions
Another direction is to compare raw last iterates with Polyak--Ruppert averages at finite time: averaging improves the rate in $T$ but can introduce an additional inverse-Jacobian factor and hence a different dependence on extreme-quantile densities.
— A Finite Sample Analysis for Quantile Temporal Difference Learning in Distributional Reinforcement Learning
(2608.27313 - Cheng et al., 27 Aug 2026) in Section Conclusions