Minimax Q-learning Control for Linear Systems Using the Wasserstein Metric

Published 14 Oct 2020 in math.OC, cs.SY, and eess.SY | (2010.06794v4)

Abstract: Stochastic optimal control usually requires an explicit dynamical model with probability distributions, which are difficult to obtain in practice. In this work, we consider the linear quadratic regulator (LQR) problem of unknown linear systems and adopt a Wasserstein penalty to address the distribution uncertainty of additive stochastic disturbances. By constructing an equivalent deterministic game of the penalized LQR problem, we propose a Q-learning method with convergence guarantees to learn an optimal minimax controller.