Self-Play PSRO: Toward Optimal Populations in Two-Player Zero-Sum Games (2207.06541v1)

Published 13 Jul 2022 in cs.GT, cs.LG, and cs.MA

Abstract: In competitive two-agent environments, deep reinforcement learning (RL) methods based on the \emph{Double Oracle (DO)} algorithm, such as \emph{Policy Space Response Oracles (PSRO)} and \emph{Anytime PSRO (APSRO)}, iteratively add RL best response policies to a population. Eventually, an optimal mixture of these population policies will approximate a Nash equilibrium. However, these methods might need to add all deterministic policies before converging. In this work, we introduce \emph{Self-Play PSRO (SP-PSRO)}, a method that adds an approximately optimal stochastic policy to the population in each iteration. Instead of adding only deterministic best responses to the opponent's least exploitable population mixture, SP-PSRO also learns an approximately optimal stochastic policy and adds it to the population as well. As a result, SP-PSRO empirically tends to converge much faster than APSRO and in many games converges in just a few iterations.

Authors (6)

Stephen McAleer (41 papers)
JB Lanier (6 papers)
Kevin Wang (41 papers)
Pierre Baldi (89 papers)
Roy Fox (39 papers)
Tuomas Sandholm (119 papers)

Citations (17)

View on Semantic Scholar

Summary

We haven't generated a summary for this paper yet.

Summarize Now

Self-Play PSRO: Toward Optimal Populations in Two-Player Zero-Sum Games (2207.06541v1)

Summary

Related Papers