Improve the dependence on the number of players and actions

Determine whether the quadratic dependence on the number of players and the squared logarithmic dependence on the number of actions in the HOOD algorithm's bounded individual-regret guarantee for arbitrary finite normal-form games can be reduced or improved.

Background

The paper establishes that the HOOD (higher-order optimism with discounting) algorithm achieves uniformly bounded individual regret of order O(N3 log2 K) in arbitrary N-player finite normal-form games with at most K actions per player. Relative to the preceding O(N log2 K log T) bound, this removes the horizon dependence but introduces an additional quadratic factor in the number of players. The authors explicitly identify reducing this N2 price and improving the log2 K dependence as unresolved research directions.

References

Natural future research directions are whether this $\nPlayers2$ price can be reduced, whether the $\log2 \nPures$ dependence can be improved, and whether bounded regret is possible in broader game classes, and for stronger notions of regret, such as swap regret.

Constant regret in general games via higher-order optimism  (2609.04113 - Abbadi et al., 3 Sep 2026) in Section 5, Concluding remarks

Another important question is whether we can obtain bounded regret from simpler dynamics---such as the classical \acs{OFTRL} update, for a suitable choice of regularizer and step size---and whether these dynamics retain good regret bounds in the game's parameters.

Constant regret in general games via higher-order optimism  (2609.04113 - Abbadi et al., 3 Sep 2026) in Section 5, Concluding remarks