ClaimArticle
We've pretty much mitigated every attack in the main categories of risk we care about, such as prompt injection and data exfiltration, to the point that these risks are far lower than the average human reviewer.
Cat Wu asserts Anthropic has largely mitigated prompt injection and data exfiltration attacks in Claude Code auto mode, with risks now far lower than an average human reviewer, with evals forthcoming. ✦ AI generated
Cat Wu · Simon Willison's Weblog · 2026-08-08 · original ↗
Elicited by
“how they run Claude Code safely within Anthropic (given the threat of prompt injection)”
We’re going to publish some evals in the coming weeks, but we’ve pretty much mitigated every attack. [...] for the main categories of risks that we’re concerned about, like prompt injection and data exfiltration, the risks are far lower than the average human reviewer.
Read full article ↗excerpt · fair-use quotation
Around this claim
This moment responds to
supports → In a third-party Trajectory Labs evaluation of 72 held-out indirect prompt injection scenarios across the latest Claude Code and Codex versions, none of the 720 attack attempts succeeded against Claude Fable 5, Opus 5, or Sonnet 5 running auto mode.The article author · Simon Willison's Weblogsupports → In a test across 1,053 paid testers where a permission prompt was swapped for a clearly dangerous command, only 13.6% of the humans refused the harmful action, whereas auto mode would have blocked 89% of those actions.The article author · Simon Willison's Weblogsupports → Auto mode is a better solution than asking humans to constantly approve actions, because confirmation fatigue makes human approval clearly unsafe.The article author · Simon Willison's Weblogsupports → Attacks aimed at a model's interior — model theft, training-data extraction, and poisoning — are bounded and largely mitigated for most developers, ranking low for initial effort compared to risks around external actions.Article author (GLM pipeline) · ByteByteGo Newsletter