Formal privacy guarantees for attention freezing
Establish whether freezing the attention parameters in GPT-2-class transformer models satisfies formal $(\varepsilon,\delta)$-differential privacy guarantees and determine how attention freezing composes with DP-SGD.
References
Open questions include whether attention freezing satisfies formal $(\varepsilon,\delta)$-DP guarantees, how it composes with DP-SGD, and how the single-step calibration extends to multi-step FedAvg.
— AEGIS: Attention-Embedding Gradient Isolation Shield - Triple-Channel Gradient Masking for Privacy-Preserving Federated LLM Fine-Tuning
(2608.19534 - Tao et al., 20 Aug 2026) in Section 6, Conclusion