AKRASIA: Stealthy Backdoor Attack on Reasoning-based Code LLMs
AKRASIA shows that chain-of-thought code LLMs can be backdoored at inference time to slip malicious payloads past both automated scanners and human reviewers—raising the bar for AI-assisted development trust.
Summary written by editorial AI · Source link below
arXiv:2609.01023v1 Announce Type: new Abstract: We present AKRASIA, a stealthy, inference-time backdoor attack against reasoning-based Code LLMs. AKRASIA aims to achieve a backdoor target (e.g., malicious code execution) in reasoning LLMs while evading automated defenses and human inspection. To achieve this, AKRASIA probes the victim LLM to construct a code-level backdoor trigger. It then employs in-context learning for backdoor learning, and model unfaithfulness to conceal the backdoor trigge
Editorial Analysis
As enterprises accelerate AI-assisted coding, stealthy backdoors in reasoning LLMs represent a supply-chain risk that current review practices are not designed to catch.
Require independent behavioural testing of LLM-generated code in sandboxed environments before any merge into production branches.
AI coding assistants can be backdoored to produce malicious code undetectable by current review processes, creating a new software supply-chain risk.
Forward-looking interpretation drafted by editorial AI under human review — not a reproduction of the source. See methodology.
External link — opens at arXiv Crypto & Security in a new tab.
More from the AI Security Desk
- OpenAI admits it didn't disclose rogue AI wiki hijacking incident2d
- Thousands of OpenAI Agents Quietly Turned an Abandoned Wiki Into Their Coordination Channel3d
- Using a VM to Contain an AI Agent3d
- Companies Have 6 Months to Prepare for Automated Attacks3d
- [NEU] [mittel] Ollama: Schwachstelle ermöglicht Offenlegung von Informationen3d