Papers
Topics
Authors
Recent
Search
2000 character limit reached

A Finite Sample Analysis for Quantile Temporal Difference Learning in Distributional Reinforcement Learning

Published 27 Aug 2026 in stat.ML and cs.LG | (2608.27313v1)

Abstract: We establish a global finite-sample guarantee for synchronous quantile temporal-difference learning (QTD) in tabular distributional reinforcement learning. The proof separates two stability mechanisms. A global comparison argument, based on the order monotonicity of reward cumulative distribution functions and the WW_\infty contraction of the distributional Bellman operator, brings an arbitrarily initialized iterate into a local neighborhood. Inside that neighborhood, we linearize the QTD mean field. Its Jacobian is a nonsingular MM-matrix, and the associated positive semigroup permits a variance-sensitive martingale analysis. For stepsizes αt=c(t+1)<sup>aα_t=c(t+1)<sup>{-a} with a(1/2,1)a\in(1/2,1), the leading last-iterate fluctuation is of order O~(T<sup>a/2/1γ)\widetilde O\bigl(T<sup>{-a/2}/\sqrt{1-γ}\bigr) and has no polynomial dependence on the number of quantiles. The deterministic transient and the required burn-in can still depend on the smallest Bellman-target density, which is of order m<sup>1m<sup>{-1} in the worst case. The result therefore distinguishes sharply between the local stochastic fluctuation and the global sample complexity.

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Tweets

Sign up for free to view the 1 tweet with 4 likes about this paper.