Optimal regret rate for full information asymmetry

Determine whether the optimal regret rate for Problem C, in which players cannot observe one another’s actions and receive independent rewards, is \(\sqrt{T}\) or \(T^{2/3}\).

Background

Problem C combines unobserved actions with independent rewards, preventing the implicit coordination mechanisms used in Problems A and B. The paper therefore analyzes two explore-then-commit algorithms, mEXC and mEXC-Bellman, and obtains a T2/3T^{2/3}-type regret bound. This is worse than the T\sqrt{T} rate established for the other two information-asymmetry settings, leaving unresolved whether the deterioration is fundamental or merely an artifact of the explore-then-commit design.

References

Whether the optimal rate for Problem C is $\sqrt{T}$ or $T{2/3}$ is left open.

Decentralized Multi-Player Q-Learning in Episodic Markov Decision Processes with Information Asymmetry  (2608.12753 - Xu et al., 13 Aug 2026) in Remark following Theorem 1, Section 3 (Problem C: Full Information Asymmetry)

Open directions include sharper bounds for Problem C, extensions to function approximation (linear MDPs), matching lower bounds under each asymmetry model, and regret--communication tradeoffs.

Decentralized Multi-Player Q-Learning in Episodic Markov Decision Processes with Information Asymmetry  (2608.12753 - Xu et al., 13 Aug 2026) in Section Conclusion