ATRIUMsearch → argument graph
ClaimArticle

We've pretty much mitigated every attack in the main categories of risk we care about, such as prompt injection and data exfiltration, to the point that these risks are far lower than the average human reviewer.

Cat Wu asserts Anthropic has largely mitigated prompt injection and data exfiltration attacks in Claude Code auto mode, with risks now far lower than an average human reviewer, with evals forthcoming. ✦ AI generated

Cat Wu · Simon Willison's Weblog · 2026-08-08 · original ↗

Elicited by

how they run Claude Code safely within Anthropic (given the threat of prompt injection)

We’re going to publish some evals in the coming weeks, but we’ve pretty much mitigated every attack. [...] for the main categories of risks that we’re concerned about, like prompt injection and data exfiltration, the risks are far lower than the average human reviewer.

Read full article ↗excerpt · fair-use quotation

Around this claim