MDP Geometry, Normalization and Reward Balancing Solvers (2407.06712v4)

Published 9 Jul 2024 in cs.LG and math.OC

Abstract: We present a new geometric interpretation of Markov Decision Processes (MDPs) with a natural normalization procedure that allows us to adjust the value function at each state without altering the advantage of any action with respect to any policy. This advantage-preserving transformation of the MDP motivates a class of algorithms which we call Reward Balancing, which solve MDPs by iterating through these transformations, until an approximately optimal policy can be trivially found. We provide a convergence analysis of several algorithms in this class, in particular showing that for MDPs for unknown transition probabilities we can improve upon state-of-the-art sample complexity results.

Summary

We haven't generated a summary for this paper yet.

Summarize Now

Tweets

https://twitter.com/alexolshevsky1/status/1857427481814266244

https://twitter.com/YPaschalidis/status/1857890457688105129

MDP Geometry, Normalization and Reward Balancing Solvers (2407.06712v4)

Summary

Related Papers

Tweets