Backdoor Attacks on Speech Emotion Recognition via TTS-Generated Poisoning
First systematic study of TTS-based poisoning attacks on speech emotion recognition shows self-supervised acoustic models are vulnerable to training-time backdoors—relevant for voice-AI in enterprise call analytics.
Summary written by editorial AI · Source link below
arXiv:2606.21052v2 Announce Type: replace-cross Abstract: Speech Emotion Recognition (SER) systems increasingly leverage self-supervised acoustic representations, yet their vulnerability to training-time attacks remains largely underexplored. This paper presents the first systematic study of poisoning-based backdoor attacks on SER, with a focus on threats enabled by text-to-speech (TTS) generated audio. We introduce a stealthy, low-energy acoustic trigger that can be embedded imperceptibly into
Editorial Analysis
Enterprises using emotion-detection in call centres or HR screening face a new attack surface if adversaries can poison training corpora with synthetic speech.
Forward-looking interpretation drafted by editorial AI under human review — not a reproduction of the source. See methodology.
External link — opens at arXiv Crypto & Security in a new tab.
More from the AI Security Desk
- OpenAI admits it didn't disclose rogue AI wiki hijacking incident2d
- Thousands of OpenAI Agents Quietly Turned an Abandoned Wiki Into Their Coordination Channel3d
- Using a VM to Contain an AI Agent3d
- Companies Have 6 Months to Prepare for Automated Attacks3d
- [NEU] [mittel] Ollama: Schwachstelle ermöglicht Offenlegung von Informationen3d