COGNITION: From Evaluation to Defense against Multimodal LLM CAPTCHA Solvers
Large language models can now bypass visual CAPTCHAs at scale, fundamentally undermining a core web security assumption relied upon by enterprise authentication flows.
Summary written by editorial AI · Source link below
arXiv:2512.02318v4 Announce Type: replace Abstract: This paper studies how multimodal large language models (MLLMs) undermine the security guarantees of visual CAPTCHA. We identify the attack surface where an adversary can cheaply automate CAPTCHA solving using off-the-shelf models. We evaluate 7 representative MLLMs on 18 real-world CAPTCHA task types, measuring single-shot accuracy, success under limited retries, end-to-end latency, and per-solve cost. We further validate our findings through
Editorial Analysis
Organizations relying on CAPTCHA as a bot protection layer may face automated attacks that bypass this defense entirely, requiring immediate authentication strategy review.
Audit web applications using CAPTCHA-only bot protection and implement multi-layered behavioral detection mechanisms.
Traditional bot protection methods are becoming obsolete as AI systems can now solve visual puzzles designed to distinguish humans from machines.
Forward-looking interpretation drafted by editorial AI under human review — not a reproduction of the source. See methodology.
External link — opens at arXiv Crypto & Security in a new tab.
More from the AI Security Desk
- Hugging Face warns an autonomous AI agent hacked its network20 Jul
- Jailbreak Foundry: From Papers to Runnable Attacks for Reproducible Benchmarking20 Jul
- Hidden in Thought: Transferable Chain-of-Thought Artifacts Induce Harmful Behavior20 Jul
- Poison to Detect: Detection of Targeted Overfitting in Federated Learning20 Jul
- Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation20 Jul