Robustness to unseen generators and heterogeneous audio types

Establish audio deepfake detectors that achieve robust performance against unseen generation methods and consistent detection accuracy across diverse audio types, including speech, environmental sound, singing, and music.

Background

The AT-ADD benchmark evaluates two complementary settings: robust speech deepfake detection under unseen generators, recording-condition variation, signal perturbations, and replay effects, and type-agnostic detection across speech, sound, singing, and music. The reported results show that high aggregate performance does not guarantee uniform generalization: generator-specific failures remain pronounced, particularly for difficult speech generators such as BigVGAN, while performance varies substantially across audio types. The paper therefore identifies both generator-specific robustness and balanced cross-type generalization as unresolved technical challenges for future audio forensic systems.

References

The results show that large-scale self-supervised representations, condition-aware augmentation, multi-crop inference, and structured fusion or routing are central to generalization, while generator-specific robustness and consistent performance across diverse audio types remain unresolved.

AT-ADD: A Benchmark and Challenge for Robust and All-Type Audio Deepfake Detection  (2608.23437 - Xie et al., 24 Aug 2026) in Abstract