ATRIUMsearch → argument graph
DataArticle

We do not know which model powered the attacker's agents, whether a jailbroken hosted model or an unrestricted open-weight one; either way, the attacker was bound by no usage policy, while our own forensic work was blocked by the guardrails of the hosted models we first tried.

Hugging Face tried using frontier models to analyze the attack, but safety guardrails blocked their forensic analysis because the prompts contained real attack payloads. They had to switch to a self-hosted model. ✦ AI generated

Hugging Face (security incident disclosure) · Simon Willison's Weblog · 2026-07-22 · original ↗

When we started the log analysis, we first used frontier models behind commercial APIs. This did not work: the analysis requires submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, and these requests were blocked by the providers' safety guardrails, which cannot distinguish an incident responder from an attacker.

Read full article ↗excerpt · fair-use quotation

Around this claim
This moment responds to