Identifiability of continuous actions from policy gradients
Establish an identifiability result for Gaussian continuous-control policies that characterizes when advantage-weighted residual information in policy gradients permits recovery of the underlying continuous action, accounting for covariance parameterization and the availability of advantage-sign information.
References
For Gaussian continuous-control policies, gradients may reveal advantage-weighted residual structure, but identifiability depends on both the covariance parameterization and the advantage-sign information. Consequently, the categorical argmin/argmax rule does not apply directly, and the corresponding identifiability statement is left to future work.
— Temporal Gradient Inversion for Private Trajectory Reconstruction in Embodied Reinforcement Learning
(2609.30258 - Bhujel et al., 24 Sep 2026) in Appendix, Section Limitations