Sublinear pathwise Markov-policy regret in continuing partial-feedback stochastic games
Establish whether decentralized learners can achieve sublinear regret against the pathwise Markov-policy benchmark in general stochastic games under continuing interaction and partial feedback.
References
To the best of our knowledge, whether decentralized learners can achieve sublinear regret against this pathwise Markov-policy benchmark in general stochastic games in the continuing, partial-feedback setting studied here remains open.
— Equilibrium in Multi-Agent Reinforcement Learning
(2608.22840 - D'Andrea et al., 24 Aug 2026) in Section 1, Introduction