---
title: Polyak-Ruppert Averaged Q-learning
url: https://www.emergentmind.com/topics/polyak-ruppert-averaged-q-learning
type: topic
---

# Polyak-Ruppert Averaged Q-learning

Searching arXiv for the cited papers to ground the article with fresh references.
Polyak–Ruppert averaged Q-learning is the practice of replacing the last Q-learning iterate by an average of iterates, most commonly the Cesàro average
\[
\bar Q_T=\frac{1}{T}\sum_{t=1}^T Q_t,
\]
or, in some regimes, a tail average over the final segment of the trajectory. In the stochastic-approximation view of Q-learning, averaging suppresses transient bias and exposes a more regular fluctuation structure, which has led to central limit theorems, functional central limit theorems, non-asymptotic concentration bounds, and parameter-free optimal-rate results across synchronous, asynchronous, discounted, average-reward, and function-approximation settings [2112.14582, 2505.21796, 2508.05984, 2509.18964, 2604.07323, 2605.17678].

## 1. Basic formulation and averaging schemes

In discounted tabular Q-learning, the asynchronous update studied in several recent analyses updates only the visited coordinate:
\[
Q_{t+1}(s_t,a_t)=Q_t(s_t,a_t)+\alpha_t\Bigl(r_t+\gamma\max_{a\in A}Q_t(s_{t+1},a)-Q_t(s_t,a_t)\Bigr),
\]
with all other entries unchanged [2604.07323]. A closely related formulation writes
\[
Q_{k+1}(s,a)=Q_k(s,a)+\alpha_k\mathbf{1}_{\{(s,a)=(S_k,A_k)\}}
\Bigl(\mathcal R(S_k,A_k)+\gamma\max_{a'}Q_k(S_k',a')-Q_k(s,a)\Bigr),
\]
which is the asynchronous tabular recursion used in a high-probability concentration analysis [2505.21796]. In synchronous tabular settings with a generative model, all state-action coordinates are updated simultaneously through an empirical Bellman operator \(\widehat T_t\), and the averaged estimator remains
\[
\bar Q_T=\frac1T\sum_{t=1}^T Q_t
\]
[2112.14582].

The most common averaging rule is full-trajectory Polyak–Ruppert averaging. In that setting, inference and concentration are formulated in terms of \(\bar Q_T-Q^*\), or equivalently in terms of the normalized partial sum
\[
\frac{1}{\sqrt{T}}\sum_{t=1}^T (Q_t-Q^*).
\]
This equivalence is exploited in both asymptotic and non-asymptotic analyses [2112.14582, 2509.18964].

Recent work shows that the averaging window is itself a design choice rather than a fixed convention. For power-law decay to zero step sizes \(\eta_{t,n}=\eta(1-t/n)^\nu\), including the linear-decay-to-zero case \(\nu=1\), the usual full Polyak–Ruppert average is described as not appropriate, because early iterates behave too much like constant-step-size iterates and may not be centered well enough around \(Q^\star\). In that regime, the relevant estimator is a tail average over the last \(\lfloor c n^{\nu/(\nu+1)}\rfloor\) iterates [2604.04218].

A related variant, sample-averaged Q-learning, combines iterate averaging with within-update batch averaging. Its update uses a batch-averaged empirical Bellman operator, while inference is still based on the time-average
\[
\bar{\mathbf Q}_T=\frac{1}{T}\sum_{t=1}^T \mathbf Q_t.
\]
This places Polyak–Ruppert averaging inside a broader family of averaged Q-learning estimators [2410.10737].

## 2. Asymptotic statistical theory in synchronous tabular settings

A detailed asymptotic theory was established for synchronous, tabular, discounted Q-learning under a generative model [2112.14582]. The central object is the standardized partial-sum process
\[
\mathbb T_T(r)=\frac{1}{\sqrt T}\sum_{t=1}^{\lfloor Tr\rfloor}(Q_t-Q^*),\qquad r\in[0,1].
\]
Under bounded fourth moments, a Lipschitz or gap condition controlling the greedy-policy switch, and slowly decaying step sizes such as \(\eta_t=t^{-\alpha}\) with \(\alpha\in(1/2,1)\), this process converges weakly to a rescaled Brownian motion:
\[
\mathbb T_T(\cdot)\overset{w}{\longrightarrow}\operatorname{Var}^{1/2}\mathbb B_D(\cdot).
\]
At \(r=1\), this yields the ordinary central limit theorem
\[
\sqrt T(\bar Q_T-Q^*)\Rightarrow N(0,\operatorname{Var})
\]
[2112.14582].

The same analysis identifies averaged Q-learning as a regular asymptotically linear estimator. Specifically,
\[
\sqrt T(\bar Q_T-Q^*)=\frac{1}{\sqrt T}\sum_{t=1}^T (-\gamma^{\pi^*})^{-1}Z_t+o_p(1),
\]
where \(Z_t=(r_t-r)+\gamma(P_t-P)V^*\). The influence function is therefore
\[
\phi_t=(-\gamma^{\pi^*})^{-1}Z_t,
\]
and the estimator attains the semiparametric efficiency bound among regular asymptotically linear estimators in that model [2112.14582].

A practical consequence is fully online inference through a self-normalized statistic constructed from the full partial-sum path. The limiting distribution is pivotal and does not require explicit estimation of the asymptotic covariance, which differentiates the method from batch-mean or covariance-plug-in procedures [2112.14582].

The finite-sample side of the same theory gives a non-asymptotic \(\ell_\infty\) analysis for \(\mathbb E\|\bar Q_T-Q^*\|_\infty\). For polynomial step sizes, the dominant term matches the instance-dependent lower bound in the tabular setting, and in the worst case the leading term yields the optimal minimax sample complexity \(T=O(D/((1-\gamma)^3\varepsilon^2))\) for \(\varepsilon\)-accuracy in \(\ell_\infty\) [2112.14582]. Entropy-regularized Q-learning admits analogous asymptotic and finite-sample results without the hard-max Lipschitz assumption, because the soft Bellman operator is smooth and has a unique fixed point [2112.14582].

## 3. High-probability concentration and the stochastic-approximation viewpoint

A recent general theorem treats Polyak–Ruppert averaging as a modular stochastic-approximation construction rather than a Q-learning-specific device [2505.21796]. The base recursion is
\[
x_{k+1}=(1-\alpha_k)x_k+\alpha_k F(x_k,w_{k+1}),\qquad
\alpha_k=\frac{\alpha}{(k+h)^\xi},\;\; h>1,\;\alpha>0,\;\xi\in[0,1],
\]
and the averaged iterate is
\[
y_k=\frac{1}{k+1}\sum_{i=0}^k x_i.
\]
The theorem assumes a norm \(\|\cdot\|_c\) whose square is \(M\)-smooth, a fixed point \(x^*\) of the mean operator \(\bar F\), invertibility of \(J_{\bar F}(x^*)-I\), local pseudo-smoothness of \(F(x,w)-x\), and sub-Gaussian noise at the fixed point and in the Jacobian [2505.21796].

The theorem is deliberately modular. If, for every \(\delta'\in(0,1)\), the non-averaged iterates satisfy
\[
\|x_i-x^*\|_c^2\le \alpha_i f_\xi(\delta',k)\qquad \forall\,0\le i\le k
\]
with probability at least \(1-\delta'\), then the averaged iterate obeys
\[
\|y_k-x^*\|_c^2\le \left(\sqrt{\tilde\epsilon(k,\delta/2)}+\sqrt{\bar\epsilon(k,\delta/2)}\right)^2
\]
with probability at least \(1-\delta\). The term \(\tilde\epsilon\) is the main averaging term and contains the \(1/(k+1)\) rate with leading \(\log(1/\delta)\) concentration, whereas \(\bar\epsilon\) is a transient term associated with the part of the trajectory preceding entry into the local pseudo-smoothness region. If the operator is globally smooth, or if \(R=\infty\), then \(\bar\epsilon\) disappears entirely. The leading behavior is summarized as
\[
\|y_k-x^*\|_c^2=\mathcal O\!\left(\frac{\log(1/\delta)}{k}\right)+\text{lower-order terms}
\]
[2505.21796].

Specialized to asynchronous tabular Q-learning with i.i.d. samples, the same framework uses the average
\[
\bar Q_k=\frac{1}{k+1}\sum_{i=0}^k Q_i.
\]
The sampling distribution is induced by a behavior policy \(\pi_b\), and the exploration parameter is
\[
\rho_b:=\min_{s,a}\{\mu^{\pi_b}(s)\pi_b(a\mid s)\}.
\]
A key structural condition is greedy uniqueness:
for every \(s\), the greedy action \(a^*(s)=\arg\max_{a'}Q^*(s,a')\) is unique. This yields the local pseudo-smoothness required by the general theorem. With \(\|\cdot\|_p\), \(p\ge 2\), the analysis uses \(M=p-1\) and \(u_{c2}=1\), and the mean operator satisfies an explicit contraction bound whose factor is
\[
\gamma_c=(|\mathcal S||\mathcal A|)^{1/p}\left(1-(1-\gamma)\rho_b\right)
\]
provided
\[
p>p_{\min}:=\frac{\ln(|\mathcal S||\mathcal A|)}{\ln\!\left(\frac{1}{1-(1-\gamma)\rho_b}\right)}\ge 2.
\]
With step size \(\alpha_k=\alpha/(k+h)^{1/2}\) and \(\alpha/\sqrt h<1\), the averaged Q-learning iterate satisfies
\[
\|\bar{Q}_k-Q^*\|_\infty^2 \le \frac{12R_{\max}^2}{(1-\gamma)^4\rho_b^2}
\frac{8\log(2/\delta)+|\mathcal S||\mathcal A|}{k+1}
+\tilde{\mathcal O}(k^{-5/4})\mathcal O(\log(1/\delta))
\]
with probability at least \(1-\delta\) [2505.21796].

That result is positioned as the first high-probability bound on the full error distribution for averaged Q-learning, with a step size that does not depend on \(\delta\) [2505.21796]. A plausible implication is that Polyak–Ruppert averaging can be analyzed in Q-learning without choosing a confidence-specific schedule, provided a sufficiently sharp raw-iterate bound is available.

The broader stochastic-approximation background for such results includes linear analyses in which averaged iterates admit exact covariance formulas in the CLT and directional non-asymptotic concentration inequalities whose leading term matches that covariance up to universal constants [2004.04719].

## 4. Asynchronous updates, Markovian sampling, and Gaussian approximation

The asynchronous case is fundamentally different from the synchronous generative-model setting because the data are Markovian and only one coordinate is updated at each iteration. One line of work formulates the trajectory \(y_k=(s_k,a_k,s_{k+1})\) as an irreducible and aperiodic finite-state Markov chain with stationary distribution \(\tilde\mu\), defines the state-action visitation matrix
\[
D=\mathrm{diag}\{p(s,a)\},
\]
and measures exploration through
\[
\rho=\min_{(s,a)\in S\times A}p(s,a).
\]
Under the local quadratic regularity condition
\[
\|(P^\pi-P^{\pi^*})(Q-Q^*)\|_\infty\le L\|Q-Q^*\|_\infty^2,
\]
polynomial step sizes \(\alpha_k=\alpha(k+b)^{-\beta}\) with \(\beta\in(0.5,1)\), and the matrix
\[
A:=D-\gamma DP^{\pi^*},
\]
the normalized averaged error converges to a Gaussian law with covariance \(A^{-1}\Sigma A^{-\top}\), and the partial-sum process converges weakly in \(\mathcal D[0,1]\) to a scaled Brownian motion [2509.18964]. The same work gives a non-asymptotic central limit theorem in 1-Wasserstein distance; its simplified corollary with \(\beta=2/3\) has rate
\[
\tilde O\!\left(\frac{(|S||A|)^{1/2}}{K^{1/6}\rho^2(1-\gamma)^3}\right),
\]
up to the manuscript’s formatting imperfections [2509.18964].

A more refined Gaussian approximation result assumes that the chain of transition triples \(\bar z_t=(s_t,a_t,s_{t+1})\) is uniformly geometrically ergodic and studies the usual Polyak–Ruppert average
\[
\bar Q_n=\frac1n\sum_{t=1}^n Q_t.
\]
With polynomial stepsize \(\alpha_t=c_0(t+k_0)^{-\omega}\), \(\omega\in(1/2,1)\), the finite-sample distribution of \(\sqrt n(\bar Q_n-Q^*)\) is approximated by a Gaussian law over hyper-rectangles, with the best leading rate obtained at \(\omega=2/3\):
\[
n^{-1/6}\log^4(nSA)
\]
up to logarithmic factors [2604.07323]. The limiting covariance is
\[
\Sigma_\infty=G^{-1}\Sigma_{\mathrm{Bell}}G^{-\top},\qquad
G=I-D_\mu(I-\gamma P^{\pi^\star}),
\]
and the proof relies on a Poisson-equation decomposition of the Markovian noise into a martingale-difference component plus smaller remainders [2604.07323].

These Markovian analyses correct a common oversimplification. Averaged Q-learning is not restricted to i.i.d. or synchronous sampling. The available theory covers non-IID trajectories, exploration-dependent covariances, and process-level weak convergence, although at the cost of stronger mixing and local regularity assumptions [2509.18964, 2604.07323].

## 5. Step-size design, tail averaging, and inferential regimes

The choice of learning-rate schedule determines not only optimization dynamics but also the correct averaging window. For the power-law decay-to-zero family
\[
\eta_{t,n}=\eta(1-t/n)^\nu,\qquad \nu>0,
\]
with linear decay to zero as the special case \(\nu=1\), the iterates behave like a constant-step-size method for \(t\le cn\) with \(c<1\), because \(\eta_{t,n}\asymp 1\) there, and only stabilize near the end where \(\eta_{t,n}\to 0\) [2604.04218]. In this triangular-array setting, the correct Polyak–Ruppert estimator is not the full average over \(1,\dots,n\), but a tail average over the last \(\lfloor c n^{\nu/(\nu+1)}\rfloor\) iterates [2604.04218].

The non-asymptotic theory for that regime shows two phases. The first is fast initial forgetting, with \(\exp(-ct)\)-type decay. The second is a terminal statistical error of order
\[
n^{-\nu/(2(\nu+1))},
\]
which is analogous in spirit to the \(t^{-\alpha/2}\) error of polynomially decaying schedules [2604.04218]. Under Assumptions 1 and 2, compactness of \(Q_0,Q^\star\in K\), and
\[
0< \eta < \frac{2(1-\gamma)}{(1-\gamma)^2 + 2(p-1)\gamma^2},
\]
the tail average satisfies the central limit theorem
\[
n^{\frac{\nu}{2(\nu+1)}}(\bar Q_n-Q^\star)\overset{w}{\to}N(0,\Sigma)
\]
for any \(c>0\) and \(\nu\ge 1/p\) [2604.04218]. The covariance \(\Sigma\) is stated to be positive definite or semidefinite and independent of \(n\), but an exact closed-form expression is described as highly intractable [2604.04218].

That same analysis establishes a strong invariance principle for the tail partial sums and derives a time-uniform Gaussian approximation. Because direct estimation of \(\Sigma\) is difficult, the paper argues that the strong approximation hints at an easily implementable Gaussian bootstrap procedure by running multiple independent chains of the approximating Gaussian recursion in parallel [2604.04218].

The main interpretive distinction is therefore not merely between averaged and non-averaged Q-learning, but between full averaging and tail averaging under different step-size geometries. A plausible implication is that “Polyak–Ruppert averaged Q-learning” is not a single estimator class with a universal windowing rule; the averaging operator must be matched to the learning-rate schedule.

## 6. Extensions beyond standard discounted tabular Q-learning

Recent work extends Polyak–Ruppert averaging to settings where the Bellman operator is a nonlinear semi-norm contraction rather than a standard norm contraction. In that framework, the update has mean map \(H\) satisfying
\[
\#1{H(x)-H(y)}\le \beta\#1{x-y},\qquad \beta\in[0,1),
\]
which includes average-reward Q-learning in the span semi-norm
\[
\|x\|_{sp}=\max_i x(i)-\min_i x(i)
\]
and discounted Q-learning in \(\|\cdot\|_\infty\) [2508.05984]. The difficulty is that the span semi-norm is not monotone, which had obstructed earlier Polyak–Ruppert analyses. The key technical device is to couple the semi-norm contraction with a suitably induced monotone norm and to recast the averaged error as a linear recursion plus a nonlinear perturbation [2508.05984]. With the universal step size
\[
\alpha_t=\frac{1}{(t+1)^\alpha},\qquad \alpha\in(1/2,1),
\]
the averaged iterate achieves parameter-free optimal rates:
\[
\mathbb E\#1{\bar Q_T-Q^*}=\tilde O\!\left(\frac{1}{\sqrt{T}}\right),
\qquad
\mathbb E\#1{\bar Q_T-Q^*}=\tilde O\!\left(\frac{1}{\sqrt{NT}}\right)
\]
with \(N\) independent agents [2508.05984]. The stated framework covers synchronous and asynchronous updates, single-agent and distributed settings, and both simulator-generated and Markovian data [2508.05984].

Entropy-regularized asynchronous Q-learning with linear function approximation furnishes another extension. There, the parameter update is
\[
\theta_{t+1}=\theta_t+\alpha_t\,\phi(s_t,a_t)\Big(r(s_t,a_t)+\gamma V_{\theta_t}^\lambda(s_{t+1})-Q_{\theta_t}(s_t,a_t)\Big),
\]
with polynomially decaying step sizes \(\alpha_t=c_0(t+k_0)^{-\omega}\), \(\omega\in(1/2,1)\), and Polyak–Ruppert average
\[
\bar\theta_n=\frac1n\sum_{t=1}^n\theta_t.
\]
Under uniform geometric ergodicity, bounded features, and regularity of the projected soft Bellman equation, the normalized averaged error \(\sqrt n(\bar\theta_n-\theta^\star)\) admits a Gaussian approximation in convex distance. Optimizing over \(\omega\) gives \(\omega=3/4\) and rate
\[
\tilde O(n^{-1/4})
\]
up to polylogarithmic factors [2605.17678]. The asymptotic covariance is
\[
\Sigma_\infty=G^{-1}\Sigma_\varepsilon G^{-\top},\qquad G=-\nabla F(\theta^\star),
\]
and entropy regularization smooths the Bellman operator and guarantees uniqueness of the soft optimal policy, avoiding degeneracy that can obstruct central limit theorems in ordinary Q-learning [2605.17678].

Sample-averaged Q-learning provides a different extension in which averaging occurs both within updates and across time. With batch size \(B_t=B\lceil t^\beta\rceil\), step size \(\eta_t=\eta t^{-\rho}\), \(\rho\in(1/2,1)\), \(\beta\in[0,2\rho-1)\), and iterate average \(\bar{\mathbf Q}_T\), the asymptotic normalization becomes
\[
\frac{T}{\sqrt{\sum_{t=1}^T B_t^{-1}}},
\]
so the batch schedule directly changes the rate. The resulting functional central limit theorem supports random-scaling confidence intervals that avoid direct estimation of the covariance matrix \(\Omega=\mathbf G^{-1}\Sigma\mathbf G^{-1}\) [2410.10737].

| Setting | Averaging object | Representative result |
|---|---|---|
| Synchronous discounted tabular | Full average \(\bar Q_T\) | FCLT, RAL property, semiparametric efficiency [2112.14582] |
| Asynchronous discounted tabular | Full average \(\bar Q_n\) | Non-asymptotic CLT and Gaussian approximation under Markovian sampling [2509.18964, 2604.07323] |
| PD2Z/LD2Z schedules | Tail average | CLT and strong approximation for the tail window [2604.04218] |
| Semi-norm contractions | Full average \(\bar Q_T\) | Parameter-free \(\tilde O(1/\sqrt{T})\) optimal rates [2508.05984] |
| Entropy-regularized linear FA | Full average \(\bar\theta_n\) | High-dimensional Gaussian approximation in convex distance [2605.17678] |
| Batch/sample-averaged Q-learning | Time-average \(\bar{\mathbf Q}_T\) with within-update batching | FCLT and random-scaling inference [2410.10737] |

Across these variants, the common role of Polyak–Ruppert averaging is stable: it converts a recursively updated Q-estimator into an object with sharper asymptotic variance, tractable Gaussian limits, or improved finite-sample concentration. The main points of divergence are equally stable: the admissible step-size schedules, the required local smoothness or uniqueness assumptions, the sampling model, and whether the correct averaging operator is the full Cesàro average or a tail window.

Source: https://www.emergentmind.com/topics/polyak-ruppert-averaged-q-learning