Policy Iteration for Pareto-Optimal Policies in Stochastic Stackelberg Games (2405.06689v1)

Published 7 May 2024 in cs.GT, cs.LG, cs.MA, and math.OC

Abstract: In general-sum stochastic games, a stationary Stackelberg equilibrium (SSE) does not always exist, in which the leader maximizes leader's return for all the initial states when the follower takes the best response against the leader's policy. Existing methods of determining the SSEs require strong assumptions to guarantee the convergence and the coincidence of the limit with the SSE. Moreover, our analysis suggests that the performance at the fixed points of these methods is not reasonable when they are not SSEs. Herein, we introduced the concept of Pareto-optimality as a reasonable alternative to SSEs. We derive the policy improvement theorem for stochastic games with the best-response follower and propose an iterative algorithm to determine the Pareto-optimal policies based on it. Monotone improvement and convergence of the proposed approach are proved, and its convergence to SSEs is proved in a special case.

References (18)

Summary

We haven't generated a summary for this paper yet.

Summarize Now

Tweets

https://twitter.com/gastronomy/status/1790232737569296553

https://twitter.com/realmofresearch/status/1790591374217314376

https://twitter.com/econ_cs/status/1790231129712156990

Policy Iteration for Pareto-Optimal Policies in Stochastic Stackelberg Games (2405.06689v1)

Summary

Related Papers

Tweets