Robustness to unseen generators and heterogeneous audio types
Establish audio deepfake detectors that achieve robust performance against unseen generation methods and consistent detection accuracy across diverse audio types, including speech, environmental sound, singing, and music.
References
The results show that large-scale self-supervised representations, condition-aware augmentation, multi-crop inference, and structured fusion or routing are central to generalization, while generator-specific robustness and consistent performance across diverse audio types remain unresolved.
— AT-ADD: A Benchmark and Challenge for Robust and All-Type Audio Deepfake Detection
(2608.23437 - Xie et al., 24 Aug 2026) in Abstract