DataArticle
In a test across 1,053 paid testers where a permission prompt was swapped for a clearly dangerous command, only 13.6% of the humans refused the harmful action, whereas auto mode would have blocked 89% of those actions.
The author reports Anthropic's human study: across 1,053 paid testers and one swapped dangerous command, only 13.6% of humans refused, while auto mode would have blocked 89% of the actions. ✦ AI generated
The article author · Simon Willison's Weblog · 2026-08-08 · original ↗
Partway through each session, a single permission prompt was swapped for a clearly dangerous command, and the vendor recorded whether the tester approved it. Every participant had the same experience. Only 13.6% of the humans refused that harmful action. Auto mode would have blocked 89% of those actions.
Read full article ↗excerpt · fair-use quotation
Around this claim
This moment responds to
supports → Auto mode is a better solution than asking humans to constantly approve actions, because confirmation fatigue makes human approval clearly unsafe.The article author · Simon Willison's Weblogrebuts → I'd like to see more independent confirmation of Anthropic's claims, because malicious instruction scenarios like a malicious package that exfiltrates data when it fetches model files may not be protectable by any version of auto mode.The article author · Simon Willison's Weblog