A Finite Sample Analysis for Quantile Temporal Difference Learning in Distributional Reinforcement Learning
Abstract: We establish a global finite-sample guarantee for synchronous quantile temporal-difference learning (QTD) in tabular distributional reinforcement learning. The proof separates two stability mechanisms. A global comparison argument, based on the order monotonicity of reward cumulative distribution functions and the contraction of the distributional Bellman operator, brings an arbitrarily initialized iterate into a local neighborhood. Inside that neighborhood, we linearize the QTD mean field. Its Jacobian is a nonsingular -matrix, and the associated positive semigroup permits a variance-sensitive martingale analysis. For stepsizes with , the leading last-iterate fluctuation is of order and has no polynomial dependence on the number of quantiles. The deterministic transient and the required burn-in can still depend on the smallest Bellman-target density, which is of order in the worst case. The result therefore distinguishes sharply between the local stochastic fluctuation and the global sample complexity.
Paper Prompts
Sign up for free to create and run prompts on this paper.