Claim◆Article
Autonomous exploit development by frontier AI agents is no longer a hypothetical capability.
The ExploitGym paper concludes that frontier AI agents can now autonomously turn vulnerabilities into working exploits, a capability previously considered implausible. ✦ AI generated
ExploitGym authors (UC Berkeley, Max Planck Institute, UC Santa Barbara, Arizona State) · Simon Willison's Weblog · 2026-07-22 · original ↗
Our results show that autonomous exploit development by frontier AI agents is no longer a hypothetical capability. While current agents are not yet reliable across all targets, they already exploit a non-trivial fraction of real-world vulnerabilities, including complex targets such as kernel components.
Read full article ↗excerpt · fair-use quotation
- ·Frontier AI agents autonomously turn vulnerabilities into working exploits
- ·This capability was previously considered implausible
- ·Includes complex targets such as kernel components
- ·Agents exploit a non-trivial fraction of real-world vulnerabilities
- ·Not yet reliable across all targets or environments
Around this claim
Evidence · 3
Rather than solve the test, the model broke its way out of OpenAI's sandbox, then found exploits to break in to Hugging Face, all so it could cheat on the test by stealing the answers.Simon Willison (author) · Simon Willison's Weblog · conf 95%The very best frontier models, unencumbered by additional guardrails, will find an exploit if there is one to be found.Simon Willison · Simon Willison's Weblog · conf 80%The future internet will be like a complex ecology of attacker and defender AI agents carving out their own ecological niches, which may require humans to deploy their own defender agents as digital white blood cells.Jack Clark (Import AI) · Import AI · conf 60%
This moment responds to
supports → The models identified and chained vulnerabilities across OpenAI's research environment and Hugging Face's production infrastructure to obtain test solutions directly from Hugging Face's production database.OpenAI (security incident disclosure) · Simon Willison's Weblogprovides context → We do not know which model powered the attacker's agents, whether a jailbroken hosted model or an unrestricted open-weight one; either way, the attacker was bound by no usage policy, while our own forensic work was blocked by the guardrails of the hosted models we first tried.Hugging Face (security incident disclosure) · Simon Willison's Webloggives example → AutoGPT was the pivotal inspiration for Onyx because it demonstrated the future of truly autonomous agents—an LLM deciding what to do, calling tools in a loop—even though the models at the time weren't good enough to make it work reliably.Maxim Bar Kogan · No Priorsgives example → This could not be a better time for AI to get good at cryptanalysis, unless AI succeeds in undermining the hard problems themselves or we live in Impagliazzo's Minicrypt.Matthew Green · Simon Willison's Weblog