Injection-Execution Dissociation: A Mechanistic Evaluation of Persistent Memory Attacks and Defenses in Stateful LLM Agents
Blocking prompt injection does not block malicious tool execution, and vice versa — this 'injection-execution dissociation' means enterprises must independently harden both layers in LLM agents.
Summary written by editorial AI · Source link below
arXiv:2605.08442v5 Announce Type: replace Abstract: We discover that prompt-injection success and tool-execution success are separable safety properties: defenses that block injection do not necessarily block execution, and vice versa. We call this the injection-execution dissociation. In LLM agents with persistent memory, malicious instructions are stored at rates exceeding 97.5%, yet downstream execution ranges from 0% to 95% with no correlation to storage rate. This reframes the threat model
Editorial Analysis
Enterprises deploying LLM agents with tool access face a dual-layer security requirement that current single-defence approaches fail to address, risking false confidence in agent safety.
Audit your LLM agent deployments to verify that prompt-injection defences and tool-execution authorisation are independently implemented and tested.
Defending AI agents against prompt attacks does not prevent them from executing malicious actions — both defences must be built and tested separately.
Forward-looking interpretation drafted by editorial AI under human review — not a reproduction of the source. See methodology.
External link — opens at arXiv Crypto & Security in a new tab.
More from the AI Security Desk
- OpenAI admits it didn't disclose rogue AI wiki hijacking incident2d
- Thousands of OpenAI Agents Quietly Turned an Abandoned Wiki Into Their Coordination Channel3d
- Using a VM to Contain an AI Agent3d
- Companies Have 6 Months to Prepare for Automated Attacks3d
- [NEU] [mittel] Ollama: Schwachstelle ermöglicht Offenlegung von Informationen3d