Effectiveness of Wav2Vec2 in LLM-based deepfake voice detection

Determine whether self-supervised Wav2Vec2 audio representations are as effective for deepfake voice detection in LLM-based architectures as they are in standalone deepfake detection systems.

Background

The paper notes that self-supervised models such as Wav2Vec2 are highly effective audio encoders in standalone deepfake voice detection systems. However, LLM-based detectors require audio representations to be mapped into the semantic space of a LLM, and the suitability of different audio encoders for this cross-modal configuration is not established. The unresolved issue is therefore whether Wav2Vec2 retains its effectiveness when integrated into an LLM-based deepfake voice detector.

References

However, it remains an open question whether they are equally effective in LLM-based architectures.

Textual Acoustic Grounding for Generalizable LLM-Based Deepfake Voice Detection  (2608.30622 - Kheir et al., 31 Aug 2026) in Section 1, Introduction