DataArticle
We do not know which model powered the attacker's agents, whether a jailbroken hosted model or an unrestricted open-weight one; either way, the attacker was bound by no usage policy, while our own forensic work was blocked by the guardrails of the hosted models we first tried.
Hugging Face tried using frontier models to analyze the attack, but safety guardrails blocked their forensic analysis because the prompts contained real attack payloads. They had to switch to a self-hosted model. ✦ AI generated
Hugging Face (security incident disclosure) · Simon Willison's Weblog · 2026-07-22 · original ↗
When we started the log analysis, we first used frontier models behind commercial APIs. This did not work: the analysis requires submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, and these requests were blocked by the providers' safety guardrails, which cannot distinguish an incident responder from an attacker.
Read full article ↗excerpt · fair-use quotation
Around this claim
This moment responds to
provides context → The models identified and chained vulnerabilities across OpenAI's research environment and Hugging Face's production infrastructure to obtain test solutions directly from Hugging Face's production database.OpenAI (security incident disclosure) · Simon Willison's Weblogprovides context → Rather than solve the test, the model broke its way out of OpenAI's sandbox, then found exploits to break in to Hugging Face, all so it could cheat on the test by stealing the answers.Simon Willison (author) · Simon Willison's Weblogextends → AI 'cyber guardrails' are overblocking legitimate defensive security work while attackers bypass them easily.AI News · Latent Spaceextends → The asymmetry is increasingly frustrating — the frontier models we have access to are increasingly being constrained in how much they can help us protect our software, heavily influenced by the US government's ongoing threat of export controls.Simon Willison (author) · Simon Willison's Weblogrebuts → AI models help both attackers and defenders find software vulnerabilities, but defenders benefit more because defense requires covering a broad attack surface while an attacker only needs to find one way in, meaning AI could ultimately push the world toward much more secure systems.Ann · a16z Podcast