Identifiability of continuous actions from policy gradients

Establish an identifiability result for Gaussian continuous-control policies that characterizes when advantage-weighted residual information in policy gradients permits recovery of the underlying continuous action, accounting for covariance parameterization and the availability of advantage-sign information.

Background

The paper’s theoretical action-recovery result applies to categorical policies over finite discrete action spaces and yields an argmin/argmax recovery rule from the policy-head gradient structure. The authors explicitly note that this rule does not directly extend to Gaussian continuous-control policies.

For continuous actions, the paper identifies two unresolved factors: the covariance parameterization of the Gaussian policy and whether information about the sign of the advantage is available. A corresponding identifiability theorem or recovery characterization is not provided and is left for future work.

References

For Gaussian continuous-control policies, gradients may reveal advantage-weighted residual structure, but identifiability depends on both the covariance parameterization and the advantage-sign information. Consequently, the categorical argmin/argmax rule does not apply directly, and the corresponding identifiability statement is left to future work.

— Temporal Gradient Inversion for Private Trajectory Reconstruction in Embodied Reinforcement Learning  (2609.30258 - Bhujel et al., 24 Sep 2026) in Appendix, Section Limitations