Polynomial iteration bound for Howard’s policy iteration with logarithmic-bit rewards
Determine whether Howard’s policy iteration admits a polynomial iteration bound on deterministic discounted Markov decision processes with at most two actions per state and nonnegative integer rewards encoded with O(log N) bits, independently of the discount factor.
References
For deterministic discounted MDPs with at most two actions per state and nonnegative integer rewards encoded with O(\log N) bits, \citet{MukherjeeKalyanakrishnan2025} give an iteration upper bound of \exp(O(\sqrt N(\log N){3/2}))=\nobreak\exp(o(N)), independently of the discount. This rules out the exponential behavior in Theorem~\ref{thm:exponential} but leaves open whether a polynomial iteration bound holds.
— Policy Iteration Is Not Strongly Polynomial for Deterministic Markov Decision Processes: The Price of Algorithmic Anarchy
(2609.40147 - Zhong et al., 30 Sep 2026) in Section 1, Introduction, paragraph following Theorem 1.2