Fastest Convergence for Q-learning (1707.03770v2)

Published 12 Jul 2017 in cs.SY, cs.LG, and math.OC

Abstract: The Zap Q-learning algorithm introduced in this paper is an improvement of Watkins' original algorithm and recent competitors in several respects. It is a matrix-gain algorithm designed so that its asymptotic variance is optimal. Moreover, an ODE analysis suggests that the transient behavior is a close match to a deterministic Newton-Raphson implementation. This is made possible by a two time-scale update equation for the matrix gain sequence. The analysis suggests that the approach will lead to stable and efficient computation even for non-ideal parameterized settings. Numerical experiments confirm the quick convergence, even in such non-ideal cases. A secondary goal of this paper is tutorial. The first half of the paper contains a survey on reinforcement learning algorithms, with a focus on minimum variance algorithms.

Citations (36)

View on Semantic Scholar

Summary

We haven't generated a summary for this paper yet.

Summarize Now

Related Papers

Stability of Q-Learning Through Design and Optimism (2023)
Zap Q-Learning With Nonlinear Function Approximation (2019)
Convex Q-Learning, Part 1: Deterministic Optimal Control (2020)
Generalized Speedy Q-learning (2019)
Zap Q-Learning for Optimal Stopping Time Problems (2019)