ATRIUMsearch → argument graph
DataArticle

In a test across 1,053 paid testers where a permission prompt was swapped for a clearly dangerous command, only 13.6% of the humans refused the harmful action, whereas auto mode would have blocked 89% of those actions.

The author reports Anthropic's human study: across 1,053 paid testers and one swapped dangerous command, only 13.6% of humans refused, while auto mode would have blocked 89% of the actions. ✦ AI generated

The article author · Simon Willison's Weblog · 2026-08-08 · original ↗

Partway through each session, a single permission prompt was swapped for a clearly dangerous command, and the vendor recorded whether the tester approved it. Every participant had the same experience. Only 13.6% of the humans refused that harmful action. Auto mode would have blocked 89% of those actions.

Read full article ↗excerpt · fair-use quotation

Around this claim