Unrestricted supremum for target discounted-sum objectives in MDPs

Determine the unrestricted optimal probability of satisfying a target discounted-sum objective in a finite Markov decision process, without restricting strategies to finite memory.

Background

For a finite Markov decision process, the target discounted-sum objective requires the generated infinite path to have discounted sum exactly equal to a specified target. The paper establishes computability and finite-memory attainability for the supremum restricted to finite-memory strategies, as well as for the infimum over all strategies.

The unrestricted supremum is stronger: strategies may use unbounded or infinite memory. The paper explicitly leaves this optimization problem unresolved because solving it would subsume the general target discounted-sum existence problem. The paper further shows that, when an unrestricted optimal strategy exists, either pseudo-polynomially bounded finite memory suffices or every optimal strategy requires infinite memory.

References

For the supremum, computing the unrestricted optimal probability remains open.

— Target Discounted Sum Problem on Markov Chains with Applications to Markov Decision Processes  (2609.03670 - Bertrand et al., 3 Sep 2026) in Section 1, Applications