MEGA Hub

Tracking the Trend in How Speech Synthesizers Deceive People

Authors

Do you know Milan Šalko?You can claim authorship or link another user.Do you know Anton Firc?You can claim authorship or link another user.Do you know Kamil Malinka?You can claim authorship or link another user.Do you know Vojtěch Staněk?You can claim authorship or link another user.Do you know Martin Perešini?You can claim authorship or link another user.Do you know Filip Pleško?You can claim authorship or link another user.Do you know Jakub Reš?You can claim authorship or link another user.

Abstract

Advances in speech synthesis have made deepfake audio highly realistic. Earlier studies reported 70-80% human detection accuracy, but relied primarily on older synthesizers. We compare human detection for three selected voice synthesis tools released in 2019, 2022, and 2024 with 82 IT professionals, and benchmark humans against six pretrained detectors on the same material. For fully synthetic speech (full spoofs), the F1 score drops from about 90% for RTVC and YourTTS to 48% for ElevenLabs, although listeners were explicitly warned that deepfakes were present. For partial spoofing, where only one sentence of an utterance is altered, strict accuracy falls to 9%, and listeners classify the synthetic sentence as bona fide 77% of the time. Humans and detectors fail in complementary ways, and neither reliably localizes short manipulations. Additionally, listeners increasingly mislabel bona fide speech as fake, eroding trust in unmanipulated audio. These findings show that human perception alone is unreliable for the selected modern and partial-spoof conditions and motivate procedural verification, provenance, watermarking, and segment-level detection.

Community

00

Publication notes

Author note
Accepted at the 6th Symposium on Security and Privacy in Speech Communication (SPSC 2026)