Extension to Markovian Trajectories
Extend the convergence theory for the single-loop, entropy-regularized, uncentered Natural Actor-Critic algorithm from conditionally independent sampling to Markovian trajectories.
References
While our current analysis utilizes conditionally independent sampling (with an exploratory $\nu$-sampling extension), we leave the extension of our theory to Markovian trajectories for future work.
— Unregularized Convergence of Single-Loop, Entropy-Regularized Natural Actor-Critic
(2608.19587 - Tan, 20 Aug 2026) in Section 1, Related Works
It remains an interesting open question whether the last-iterate rate can also be improved to $\tilde{\mathcal{O}(T_{total}{-2/3})$ via alternative analytical techniques or algorithmic designs.
— Unregularized Convergence of Single-Loop, Entropy-Regularized Natural Actor-Critic
(2608.19587 - Tan, 20 Aug 2026) in Remark 6.3, Section 6.3, Application to Tabular Setting