Sublinear pathwise Markov-policy regret in continuing partial-feedback stochastic games

Establish whether decentralized learners can achieve sublinear regret against the pathwise Markov-policy benchmark in general stochastic games under continuing interaction and partial feedback.

Background

The paper distinguishes the unresolved benchmark of competing with a fixed Markov-policy deviation in hindsight from the adaptive Markov coarse regret (AMCR) criterion introduced in the paper. In a continuing stochastic game, changing a player’s Markov policy changes the subsequent state trajectory and therefore the data used by all learners, so the payoff of a counterfactual deviation cannot generally be inferred from the realized history without specifying how the learners would behave on that counterfactual path.

The authors instead establish equilibrium guarantees for AMCR and show that vanishing AMCR leads to Markov Bayes coarse correlated equilibrium. They explicitly leave open whether the stronger pathwise Markov-policy regret benchmark can be achieved by decentralized learners in the continuing, partial-feedback setting.

References

To the best of our knowledge, whether decentralized learners can achieve sublinear regret against this pathwise Markov-policy benchmark in general stochastic games in the continuing, partial-feedback setting studied here remains open.

Equilibrium in Multi-Agent Reinforcement Learning  (2608.22840 - D'Andrea et al., 24 Aug 2026) in Section 1, Introduction