Hide and Seek in Embedding Space: Geometry-based Steganography and Detection in Large Language Models
Research formalises steganographic payload recoverability in fine-tuned LLMs and shows embedding-space geometry enables covert channels that evade naive detection.
Summary written by editorial AI · Source link below
arXiv:2601.22818v2 Announce Type: replace Abstract: Fine-tuned LLMs can covertly encode prompt secrets into outputs via steganographic channels. Prior work demonstrated this threat but relied on trivially recoverable encodings. We formalize payload recoverability via classifier accuracy and show previous schemes achieve 100\% recoverability. In response, we introduce low-recoverability steganography, replacing arbitrary mappings with embedding-space-derived ones. For Llama-8B (LoRA) and Ministr
Editorial Analysis
Organisations deploying fine-tuned LLMs should consider that model outputs could embed hidden information, creating a covert data-exfiltration pathway.
Include steganographic output analysis in your LLM security assessment framework.
Forward-looking interpretation drafted by editorial AI under human review — not a reproduction of the source. See methodology.
External link — opens at arXiv Crypto & Security in a new tab.
More from the Research Desk
- Is That Really My X-Ray? Measuring Internet-Exposed DICOM Services in the Presence of Deception20 Jul
- Characterizing Phishing Pages by JavaScript Capabilities20 Jul
- Intentional Electromagnetic Interference Attacks on Facial Recognition20 Jul
- DoSQ: A Cross-Layer Denial of Service Quality Attack by Exploiting Side Channels in 5G NR20 Jul
- Vogls: a Fast Interactive Full-timing Simulator for Pre-silicon Power Side-Channel Analysis20 Jul