Statistical Efficiency and Inference of Quantile Distributional Reinforcement Learning
Published 9 Jul 2026 in stat.ML and cs.LG | (2607.08444v1)
Abstract: In this paper, we study quantile-based distributional reinforcement learning from the perspective of statistical efficiency. We focus on distributional policy evaluation, whose goal is to characterize the return distribution, namely the distribution of discounted cumulative rewards under a given policy. To obtain a finite-dimensional representation of the return distribution, we consider the quantile fixed point η<em>m induced by the quantile-projected distributional Bellman equation. Assuming access to a generative model, we construct an estimator ηm<sup>(n) based on an empirical Markov decision process. For a fixed number of quantiles m, we establish a non-asymptotic error bound for ηm<sup>(n) and ηm under the supremum W</em>∞ metric, showing that the estimation error scales as O(m/n) with respect to m and n. This implies that the quantile-based distributional policy evaluation problem can be solved with sample efficiency, achieving the optimal parametric n convergence rate. We derive the asymptotic distribution of the quantile parameters n(θm<sup>(n)−θm) and characterize the semiparametric efficiency bound, which is attained by our estimator. Beyond the fixed-dimensional setting, we investigate the asymptotic regime in which the number of quantiles diverges. We characterize the limit covariance structure and show that it matches the semiparametric efficiency bound of the nonparametric model for distributional policy evaluation, showing that quantile-based estimators remain asymptotically efficient in the infinite-dimensional limit. Finally, we establish a Berry--Esseen theorem for smooth functionals n(ηm<sup>(n)(s)−ηm(s))f, thereby providing a foundation for statistically valid inference on functionals of the quantile-projected return distribution.
The paper establishes a sharp non-asymptotic error bound of O(√(m/n)) for quantile-based RL estimators, ensuring optimal statistical efficiency.
The paper demonstrates asymptotic normality and semiparametric efficiency through a closed-form covariance analysis of the quantile fixed point estimator.
The paper provides practical guidance on sample complexity and inference, enabling robust uncertainty quantification in distributional reinforcement learning.
Statistical Efficiency and Inference in Quantile Distributional RL: A Technical Summary
This paper investigates the statistical properties and efficiency of quantile-based distributional reinforcement learning (DRL), focusing primarily on the estimation and inference of return distributions under finite data settings. The analysis spans finite-sample (non-asymptotic) error bounds, asymptotic behavior, and the limiting efficiency of quantile-parametric representations when the number of quantile bins increases to infinity.
Quantile Projection and Distributional Bellman Operators
The central object of interest is the quantile fixed point: a finite-dimensional summary of the return distribution (the discounted sum of rewards under a fixed policy) obtained by quantile projection of the distributional Bellman operator. Unlike categorical methods, quantile DRL parameterizes the return distribution via quantile functions, introducing a fundamentally nonlinear, non-smooth fixed-point system with analytical challenges distinct from its categorical counterparts, owing to its inverse CDF structure and the non-smoothness of quantile operators.
Non-Asymptotic Statistical Guarantees
The paper defines a quantile fixed point estimator, constructed from empirical rewards and transitions in a generative model setting. A central theoretical contribution is the establishment of a sharp non-asymptotic bound on the estimation error between the empirical estimator ηm(n) and the population quantile fixed point ηm, assessed with the supremum ∞-Wasserstein metric.
Key result: The estimation error scales as O(m/n), where m is the number of quantile bins and n the number of generative queries per state-action pair. This implies quantile DRL achieves the optimal parametric rate for fixed m and is statistically efficient in terms of the sample complexity for distributional policy evaluation.
Figure 1: The statistical error $\Bar{W}_\infty(\bm\eta_m^{(n)},\bm\eta_m)$ as a function of n for various m, corroborating ηm0 scaling of error.
The theorem encompasses both situations where reward densities are bounded away from zero and cases with vanishing boundary density, with explicit dependence on problem parameters.
Asymptotic Normality, Efficiency, and Limiting Regimes
In the asymptotic regime, the estimator for the quantile parameters is shown to be ηm1-consistent and asymptotically normal. The authors derive the asymptotic covariance in closed form, involving the local sensitivity (Jacobian) of the fixed-point equation and the sampling variability of the projected Bellman operator.
Semiparametric Efficiency: A particularly strong claim is established: the quantile fixed point estimator exactly attains the semiparametric lower bound for estimating return distribution functionals in the quantile parametrization, for any fixed ηm2. Moreover, as ηm3 increases, the limiting covariance operator matches the efficiency bound of the infinite-dimensional (nonparametric) distributional model. Thus, quantile discretization introduces no asymptotic efficiency gap—even in the continuum limit.
This operator-theoretic convergence is technically non-trivial, requiring careful handling of the singularities arising at distributional boundaries and a formal correspondence between measure and quantile-function representations.
Distributional Inference and Berry–Esseen Rates
The paper further develops inference procedures for smooth functionals of the quantile-projected return distributions. Confidence intervals can be constructed for integral functionals of the form ηm4 (where ηm5 is smooth), with variance estimated plug-in.
Berry–Esseen Bound: For the Gaussian approximation of these functionals, a non-classical Berry–Esseen rate of ηm6 is obtained, reflecting the inherent nonlinearity and non-smoothness of the implicit quantile fixed-point solution. This rate is slower than classical quantile estimation (ηm7), due to the absence of a Bahadur representation and the presence of indicator functions in the Bellman equation.
Figure 2: The QQ plot of the standardized estimator ηm8 at large ηm9 for various ∞0 shows numerical agreement with Gaussianity.
Practical Implications and Theoretical Consequences
Efficiency Justification for Quantile DRL: The results rigorously justify the use of quantile-based distributional RL by showing no statistical efficiency loss, even with high quantile resolution, providing a theoretical explanation for empirical findings in deep quantile RL.
Confidence Intervals and Inference: Statistically valid inference is enabled for a broad class of functionals of the return distribution, with practical procedures operable via plug-in variance estimators (e.g., based on kernel density estimation under mild regularity assumptions).
Sample Complexity Guidance: The non-asymptotic results provide explicit guidance for the sample complexity requirements in model-based DRL setups aiming for distributional evaluation.
Directions for Future Research
The work highlights open problems and future directions:
Improved Gaussian Approximation Rates: Sharpening the ∞1-dependence in the Berry–Esseen bounds, aiming for rates closer to classical parametric regimes.
Extension Beyond Generative Models: Transferring these techniques to purely offline or online RL settings, such as off-policy quantile temporal difference learning, where data is not i.i.d. or fully generative, is non-trivial due to the compounded estimation and approximation errors.
Conclusion
The paper bridges a notable gap in the theoretical literature of quantile-based DRL by providing rigorous, comprehensive statistical guarantees—non-asymptotic error bounds, asymptotic normality, exact semiparametric efficiency, and quantitative central limit rates—for the quantile-parametric estimation of return distributions. The results clarify that quantile-based methods are both statistically sound and practically robust, with implications extending to uncertainty quantification, risk-sensitive learning, and downstream applications in safety-critical RL.