Data◆Article
In a third-party Trajectory Labs evaluation of 72 held-out indirect prompt injection scenarios across the latest Claude Code and Codex versions, none of the 720 attack attempts succeeded against Claude Fable 5, Opus 5, or Sonnet 5 running auto mode.
The article reports Anthropic's commissioned third-party evaluation by Trajectory Labs: across 72 held-out indirect prompt injection scenarios, none of 720 attack attempts succeeded against auto mode on Claude Fable 5, Opus 5, or Sonnet 5. ✦ AI generated
The article author · Simon Willison's Weblog · 2026-08-08 · original ↗
We commissioned an evaluation from a third party, Trajectory Labs, who tested different models within the latest publicly available versions of Claude Code and Codex as of July 17th 2026. They tested 72 indirect prompt injection scenarios held out from Anthropic. [...] In this evaluation, none of the 720 attack attempts succeeded against Claude Fable 5, Opus 5, or Sonnet 5 running auto mode.
Read full article ↗excerpt · fair-use quotation
- ·Third-party eval, held-out scenarios from Anthropic
- ·Tested latest Claude Code and Codex (as of July 17, 2026)
- ·72 indirect prompt injection scenarios, 720 attacks total
- ·None of 720 attempts succeeded
- ·Claude Fable 5, Opus 5, Sonnet 5 all fully resistant
- ·Evaluation ran under auto mode across all three
Around this claim
Evidence · 2
We've pretty much mitigated every attack in the main categories of risk we care about, such as prompt injection and data exfiltration, to the point that these risks are far lower than the average human reviewer.Cat Wu · Simon Willison's Weblog · conf 80%Auto mode is a better solution than asking humans to constantly approve actions, because confirmation fatigue makes human approval clearly unsafe.The article author · Simon Willison's Weblog · conf 70%