Robustness of policy geometry beyond the attained Dog-Stand regime
Determine whether substantially stronger training, a different optimizer schedule, or an algorithm other than PPO preserves the observed entropy-dependent policy geometry in bounded continuous-control policies.
References
The Dog conclusions are also tied to the attained performance regime. Canonical stochastic returns span roughly 464--594 on 1000-step episodes. Whether substantially stronger training, a different optimizer schedule, or another algorithm preserves the same geometry remains open.
— Where Entropy Is Measured Matters: Policy Geometry in Bounded Continuous-Control PPO
(2608.24488 - He et al., 25 Aug 2026) in Section 6, subsection “Scope”