Tracking the Trend in How Speech Synthesizers Deceive People
Updated study shows human ability to detect synthetic speech has dropped significantly as synthesisers improve, reinforcing the urgency for automated deepfake-audio detection in enterprise voice channels.
Summary written by editorial AI · Source link below
arXiv:2608.19959v1 Announce Type: new Abstract: Advances in speech synthesis have made deepfake audio highly realistic. Earlier studies reported 70-80% human detection accuracy, but relied primarily on older synthesizers. We compare human detection for three selected voice synthesis tools released in 2019, 2022, and 2024 with 82 IT professionals, and benchmark humans against six pretrained detectors on the same material. For fully synthetic speech (full spoofs), the F1 score drops from about 90
Editorial Analysis
With vishing and CEO-fraud attacks leveraging ever-more-realistic voice clones, enterprises relying on human judgement for voice authentication face rapidly growing exposure.
Reassess voice-based authentication and authorisation workflows for susceptibility to modern speech synthesis and accelerate adoption of automated liveness-detection controls.
Human listeners are increasingly unable to distinguish synthetic from real speech, elevating the risk of voice-clone fraud in executive communication channels.
Forward-looking interpretation drafted by editorial AI under human review — not a reproduction of the source. See methodology.
External link — opens at arXiv Crypto & Security in a new tab.
More from the Research Desk
- ShadowPath: Lookup-Private Credential Status Verification over Authenticated State21 Aug
- QUASAR: A Quantum-Classical Neural Network for SAR Satellite Physical-Layer Authentication21 Aug
- Survival of~the~Stealthiest: Evolving Low-Entropy Ransomware via~Genetic Algorithms21 Aug
- Better Call Graphs: A New Dataset of Function Call Graphs for Malware Classification21 Aug
- WaveVerif: Acoustic Side-Channel based Verification of Robotic Workflows21 Aug