Formal Differential-Privacy Guarantees for HAO Training

Establish a formal $(\epsilon, \delta)$-differential privacy guarantee for the Homomorphic Advantage Operator framework by applying per-example gradient clipping and a privacy accountant to its Gaussian-noise training mechanism.

Background

The Homomorphic Advantage Operator framework combines homomorphic-encryption-based reinforcement learning with clipped weight updates and optional Gaussian noise modeled on DP-SGD. The paper reports empirical stability under such noise, including a 0% boundary-breach rate for noise levels up to σ=1.0\sigma=1.0. However, the implementation adds noise to clipped mini-batch gradients rather than performing the per-example clipping required for a formal differential-privacy analysis.

The unresolved task is therefore to provide a rigorous (ϵ,δ)(\epsilon, \delta) guarantee, which would require per-example clipping and a privacy accountant. Such a result would quantify protection against gradient-inversion attacks rather than merely demonstrate empirical robustness to DP-style noise.

References

In these experiments, the noise is added to the clipped mini-batch gradient; a formal $(\epsilon, \delta)$ guarantee additionally requires per-example clipping and a privacy accountant, which we leave to future work.

— Homomorphic Advantage Operator: Stabilizing Reinforcement Learning Under Fully Homomorphic Encryption Constraints  (2610.02074 - Nadhir et al., 1 Oct 2026) in Section 3.3, “Client-Server Architecture and Privacy”