CachePrune: Privacy-Aware and Fine-Grained KV Cache Sharing for Efficient LLM Inference
Multi-tenant LLM services sharing cached computations across users create data leakage risks that could expose confidential enterprise prompts to competitors.
Summary written by editorial AI · Source link below
arXiv:2605.23640v1 Announce Type: new Abstract: Large Language Models (LLMs) rely on Key-Value (KV) caching to accelerate inference, and many serving systems further share the KV cache across users' requests to reduce redundant computation. While widely adopted, unrestricted cross-user sharing introduces side-channel vulnerabilities, allowing an adversary to infer user inputs by probing for cache reuse. Existing defenses disable sharing entirely to prevent leakage; yet such a coarse-grained str
External link — opens at arXiv Crypto & Security in a new tab.
More from the Cloud Desk
- New NadMesh Botnet Hunts Exposed AI Services for Cloud Keys and Kubernetes Tokens17 Jul
- Google Bets 'Agentic Defense' Strategy Can Outpace Attackers17 Jul
- {\epsilon}-Indistinguishability In Moving Target Defense: Framework, Algorithms, And Cloud Case Studies16 Jul
- The Risk of Exposed Cloud Functions and How to Harden15 Jul
- The Red Agent POV: The One Boolean That Broke a B2B Platform’s Credit System15 Jul