Papers
Topics
Authors
Recent
Search
2000 character limit reached

Constant regret in general games via higher-order optimism

Published 3 Sep 2026 in cs.LG and cs.GT | (2609.04113v1)

Abstract: We introduce an uncoupled learning algorithm which, when employed by all players of an arbitrary NN-player normal form game with up to KK actions per player, guarantees O(N<sup>3log<sup>2</sup></sup>K)O(N<sup>3\log<sup>2</sup></sup> K) individual regret, uniformly over the horizon of play. The proposed algorithm - which we call higher-order optimism with discounting (HOOD) is a variant of optimistic follow-the-regularized-leader (OptFTRL) that combines a discounted (N+1)(N+1)-th order predictor with entropic regularization over a suitable "lifting" of the game's strategy space. This combination of ingredients is purposefully designed to dampen large oscillations of the induced sequence of play in a controlled manner, removing in this way a key stumbling block of previous attempts to achieve constant regret in general games. Our approach bears several striking similarities to the concurrent - and completely independent - work of Liu, Farina, and Ozdaglar (arXiv:2608.31166), who very recently derived an O(N<sup>21log<sup>4</sup></sup>K)O(N<sup>{21}\log<sup>{4}</sup></sup> K) regret bound through the use of higher-order optimism and an exponential moving average estimator.

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.