Refusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Activations
Linear probes on transformer residual streams can predict LLM refusal before token generation, revealing exploitable internal signals — relevant for both red-teamers and guardrail designers.
Summary written by editorial AI · Source link below
arXiv:2605.28553v2 Announce Type: replace-cross Abstract: In this paper, we investigate whether refusal behavior can be predicted from LLM intermediate activations before decoding using linear probes trained on residual stream activations at each transformer block. We find that refusal is linearly decodable well before the final layer, indicating that safety-relevant behavior is represented in intermediate activations before output generation. To test whether this signal is actionable, we intro
Editorial Analysis
Understanding that refusal behaviour is detectable — and exploitable — in intermediate model layers changes how enterprises should evaluate the robustness of LLM safety controls.
Forward-looking interpretation drafted by editorial AI under human review — not a reproduction of the source. See methodology.
External link — opens at arXiv Crypto & Security in a new tab.
More from the Research Desk
- 39 New Methods That Compromise Passkey Authentication3d
- Security Vulnerability in a Voting System3d
- Selfie-Capture Dynamics as an Auxiliary Signal Against Deepfakes and Injection Attacks for Mobile Identity Verification4d
- How Reliable Is the Multi-Input Heuristic for Bitcoin Address Clustering in Law Enforcement Contexts?4d
- Privacy Leakage in Federated Learning: Gradient-Based Client Identity Inference and Defenses for Inertial Sensing in Vehicular Edge Networks4d