Robustness of policy geometry beyond the attained Dog-Stand regime

Determine whether substantially stronger training, a different optimizer schedule, or an algorithm other than PPO preserves the observed entropy-dependent policy geometry in bounded continuous-control policies.

Background

The paper studies diagonal Gaussian policies with state-independent learned log standard deviations, bounded actions, and PPO training in two MuJoCo-family simulation tasks: MyoLeg and Dog-Stand. It finds that measuring entropy in latent Gaussian space versus executed-action space produces distinct mean-placement and variance regimes, with executed-action entropy generally yielding more interior policy means.

The authors explicitly qualify this conclusion as tied to the attained Dog-Stand performance regime, where canonical stochastic returns range approximately from 464 to 594 over 1000-step episodes. They leave unresolved whether the same policy-geometry ordering and mechanism would persist under substantially stronger training, a different optimizer schedule, or a different reinforcement-learning algorithm.

References

The Dog conclusions are also tied to the attained performance regime. Canonical stochastic returns span roughly 464--594 on 1000-step episodes. Whether substantially stronger training, a different optimizer schedule, or another algorithm preserves the same geometry remains open.

— Where Entropy Is Measured Matters: Policy Geometry in Bounded Continuous-Control PPO  (2608.24488 - He et al., 25 Aug 2026) in Section 6, subsection “Scope”