Extension to Markovian Trajectories

Extend the convergence theory for the single-loop, entropy-regularized, uncentered Natural Actor-Critic algorithm from conditionally independent sampling to Markovian trajectories.

Background

The paper analyzes the single-loop, entropy-regularized Natural Actor-Critic algorithm primarily under conditionally independent sampling from the discounted visitation measure, with an extension to exploratory restart distributions. Although Markovian sampling is important for practical reinforcement-learning trajectories and has been studied in related actor–critic analyses, the presented theory does not establish convergence guarantees under this sampling mechanism.

The authors explicitly identify extending the theory to Markovian trajectories as future work, making this an unresolved methodological and theoretical problem rather than a general aspiration.

References

While our current analysis utilizes conditionally independent sampling (with an exploratory $\nu$-sampling extension), we leave the extension of our theory to Markovian trajectories for future work.

Unregularized Convergence of Single-Loop, Entropy-Regularized Natural Actor-Critic  (2608.19587 - Tan, 20 Aug 2026) in Section 1, Related Works

It remains an interesting open question whether the last-iterate rate can also be improved to $\tilde{\mathcal{O}(T_{total}{-2/3})$ via alternative analytical techniques or algorithmic designs.

Unregularized Convergence of Single-Loop, Entropy-Regularized Natural Actor-Critic  (2608.19587 - Tan, 20 Aug 2026) in Remark 6.3, Section 6.3, Application to Tabular Setting