Papers
Topics
Authors
Recent
Search
2000 character limit reached

Online Learning for Uninformed Markov Games: Empirical Nash-Value Regret and Non-Stationarity Adaptation

Published 6 Feb 2026 in cs.LG, cs.GT, and stat.ML | (2602.07205v1)

Abstract: We study online learning in two-player uninformed Markov games, where the opponent's actions and policies are unobserved. In this setting, Tian et al. (2021) show that achieving no-external-regret is impossible without incurring an exponential dependence on the episode length HH. They then turn to the weaker notion of Nash-value regret and propose a V-learning algorithm with regret O(K<sup>2/3)O(K<sup>{2/3}) after KK episodes. However, their algorithm and guarantee do not adapt to the difficulty of the problem: even in the case where the opponent follows a fixed policy and thus O(K)O(\sqrt{K}) external regret is well-known to be achievable, their result is still the worse rate O(K<sup>2/3)O(K<sup>{2/3}) on a weaker metric. In this work, we fully address both limitations. First, we introduce empirical Nash-value regret, a new regret notion that is strictly stronger than Nash-value regret and naturally reduces to external regret when the opponent follows a fixed policy. Moreover, under this new metric, we propose a parameter-free algorithm that achieves an O(minK+(CK)<sup>1/3,LK)O(\min {\sqrt{K} + (CK)<sup>{1/3},\sqrt{LK}}) regret bound, where CC quantifies the variance of the opponent's policies and LL denotes the number of policy switches (both at most O(K)O(K)). Therefore, our results not only recover the two extremes -- O(K)O(\sqrt{K}) external regret when the opponent is fixed and O(K<sup>2/3)O(K<sup>{2/3}) Nash-value regret in the worst case -- but also smoothly interpolate between these extremes by automatically adapting to the opponent's non-stationarity. We achieve so by first providing a new analysis of the epoch-based V-learning algorithm by Mao et al. (2022), establishing an O(ηC+K/η)O(ηC + \sqrt{K/η}) regret bound, where ηη is the epoch incremental factor. Next, we show how to adaptively restart this algorithm with an appropriate ηη in response to the potential non-stationarity of the opponent, eventually achieving our final results.

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.