Polynomial iteration bound for Howard’s policy iteration with logarithmic-bit rewards

Determine whether Howard’s policy iteration admits a polynomial iteration bound on deterministic discounted Markov decision processes with at most two actions per state and nonnegative integer rewards encoded with O(log N) bits, independently of the discount factor.

Background

The paper discusses deterministic discounted Markov decision processes with at most two actions per state and nonnegative integer rewards whose encodings use O(log N) bits. Prior work cited in the paper establishes an upper bound of exp(O(√N(log N){3/2})) iterations for Howard’s policy iteration, independently of the discount factor. This subexponential upper bound excludes the fully exponential behavior proved elsewhere in the paper, but it does not determine whether the iteration complexity is polynomial. The paper’s stretched-exponential lower bound demonstrates superpolynomial behavior under the same reward-size restriction, so the stated question remains unresolved within the cited context.

References

For deterministic discounted MDPs with at most two actions per state and nonnegative integer rewards encoded with O(\log N) bits, \citet{MukherjeeKalyanakrishnan2025} give an iteration upper bound of \exp(O(\sqrt N(\log N){3/2}))=\nobreak\exp(o(N)), independently of the discount. This rules out the exponential behavior in Theorem~\ref{thm:exponential} but leaves open whether a polynomial iteration bound holds.

— Policy Iteration Is Not Strongly Polynomial for Deterministic Markov Decision Processes: The Price of Algorithmic Anarchy  (2609.40147 - Zhong et al., 30 Sep 2026) in Section 1, Introduction, paragraph following Theorem 1.2