MEGA Hub

AffectDF: The Most Comprehensive Benchmark for Speech Deepfake Detection against Emotionally Expressive Attacks

Authors

Do you know Aurosweta Mahapatra?You can claim authorship or link another user.Do you know Xiutian Zhao?You can claim authorship or link another user.Do you know Shreeram Suresh Chandra?You can claim authorship or link another user.Do you know Zihan Zhang?You can claim authorship or link another user.Do you know Zongyang Du?You can claim authorship or link another user.Do you know Ismail Rasim Ulgen?You can claim authorship or link another user.Do you know Kong Aik Lee?You can claim authorship or link another user.Do you know Nicholas Andrews?You can claim authorship or link another user.Do you know Carlos Busso?You can claim authorship or link another user.Do you know Berrak Sisman?You can claim authorship or link another user.

Abstract

Speech deepfake detection (SDD) systems achieve strong performance on conventional benchmarks; however, existing datasets provide limited coverage of emotionally expressive and recent large audio-language model (LALM)-based attacks. Existing emotional spoofing datasets are also limited in scale and attack diversity, typically covering only voice conversion (VC) or text-to-speech (TTS) attacks. We introduce AffectDF, the most comprehensive benchmark for emotionally expressive speech deepfakes, spanning TTS, VC, emotional VC, and LALM-based spoofing attacks across both acted and spontaneous emotional speech. AffectDF contains approximately 260 hours of speech generated using 21 spoofing attacks across five emotional states. We benchmark state-of-the-art SDD systems under conventional and emotional spoofing conditions, including LALM-based detectors evaluated with both inference-only prompting and supervised fine-tuning. Our experiments reveal severe robustness degradation when models trained on conventional benchmarks are evaluated on AffectDF, with several systems approaching near-random performance. Surprisingly, even large-scale emotional training does not consistently improve cross-domain robustness, indicating that current SDD systems fail to learn generalized spoof representations under emotional and prosodic variability. Robustness further varies substantially across emotional states, attack families, and acted vs spontaneous emotional speech conditions. These findings expose fundamental limitations of current SDD systems and establish AffectDF as a benchmark for developing more robust spoof detection models.

Community

00