Memorization Is Not Extraction: Tight Differential-Privacy Bounds and Audit Blind Spots
Researchers formalise the distinction between LLM memorization and data extraction, showing that differential privacy guards against some definitions but leaves blind spots — important for enterprises relying on DP as a privacy silver bullet.
Summary written by editorial AI · Source link below
arXiv:2608.27782v1 Announce Type: new Abstract: Memorization in large language models is measured through a zoo of definitions whose formal relations are unknown, and differential privacy (DP) is treated as a proxy against all of them at once. We pin down the exact DP constant for the two that carry the practical weight, counterfactual memorization and adaptive extraction, and show that they do not control each other. Under $f$-DP, every adaptive extraction protocol with list budget $m$ succeed
Editorial Analysis
Enterprises using differential privacy to protect training data may overestimate their protection; this work shows DP's coverage is definition-dependent, creating residual extraction risk.
Forward-looking interpretation drafted by editorial AI under human review — not a reproduction of the source. See methodology.
External link — opens at arXiv Crypto & Security in a new tab.
More from the Research Desk
- 39 New Methods That Compromise Passkey Authentication3d
- Security Vulnerability in a Voting System3d
- Selfie-Capture Dynamics as an Auxiliary Signal Against Deepfakes and Injection Attacks for Mobile Identity Verification4d
- How Reliable Is the Multi-Input Heuristic for Bitcoin Address Clustering in Law Enforcement Contexts?4d
- Privacy Leakage in Federated Learning: Gradient-Based Client Identity Inference and Defenses for Inertial Sensing in Vehicular Edge Networks4d