Formal Differential-Privacy Guarantees for HAO Training
Establish a formal $(\epsilon, \delta)$-differential privacy guarantee for the Homomorphic Advantage Operator framework by applying per-example gradient clipping and a privacy accountant to its Gaussian-noise training mechanism.
References
In these experiments, the noise is added to the clipped mini-batch gradient; a formal $(\epsilon, \delta)$ guarantee additionally requires per-example clipping and a privacy accountant, which we leave to future work.
— Homomorphic Advantage Operator: Stabilizing Reinforcement Learning Under Fully Homomorphic Encryption Constraints
(2610.02074 - Nadhir et al., 1 Oct 2026) in Section 3.3, “Client-Server Architecture and Privacy”