ATRIUMsearch → argument graph
DataArticle

In a third-party Trajectory Labs evaluation of 72 held-out indirect prompt injection scenarios across the latest Claude Code and Codex versions, none of the 720 attack attempts succeeded against Claude Fable 5, Opus 5, or Sonnet 5 running auto mode.

The article reports Anthropic's commissioned third-party evaluation by Trajectory Labs: across 72 held-out indirect prompt injection scenarios, none of 720 attack attempts succeeded against auto mode on Claude Fable 5, Opus 5, or Sonnet 5. ✦ AI generated

The article author · Simon Willison's Weblog · 2026-08-08 · original ↗

We commissioned an evaluation from a third party, Trajectory Labs, who tested different models within the latest publicly available versions of Claude Code and Codex as of July 17th 2026. They tested 72 indirect prompt injection scenarios held out from Anthropic. [...] In this evaluation, none of the 720 attack attempts succeeded against Claude Fable 5, Opus 5, or Sonnet 5 running auto mode.

Read full article ↗excerpt · fair-use quotation

Around this claim