Papers
Topics
Authors
Recent
Search
2000 character limit reached

SQT -- std QQ-target

Published 3 Feb 2024 in cs.LG and cs.AI | (2402.05950v3)

Abstract: Std QQ-target is a conservative, actor-critic, ensemble, QQ-learning-based algorithm, which is based on a single key QQ-formula: QQ-networks standard deviation, which is an "uncertainty penalty", and, serves as a minimalistic solution to the problem of overestimation bias. We implement SQT on top of TD3/TD7 code and test it against the state-of-the-art (SOTA) actor-critic algorithms, DDPG, TD3 and TD7 on seven popular MuJoCo and Bullet tasks. Our results demonstrate SQT's QQ-target formula superiority over TD3's QQ-target formula as a conservative solution to overestimation bias in RL, while SQT shows a clear performance advantage on a wide margin over DDPG, TD3, and TD7 on all tasks.

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.