ATRIUMsearch → argument graph
ClaimArticle

Models that assume user intent rather than executing exactly what is said seem inherently more unsafe, and this is a distinct axis from persistence.

Proposes a second safety axis: instruction-following precision. Models that infer intended action rather than doing exactly what was said are more dangerous, echoing paperclip-problem-style debates. ✦ AI generated

Nathan Lambert · Interconnects · 2026-08-09 · original ↗

On the other side is how much the models assume user intent, versus trying to infer the intended action. A model that will do what it thinks you wanted rather than what you said seems inherently more unsafe. I think of this with respect to instruction following precision, where in the future it seems like the models should only do exactly what we tell them, but this opens a lot of debates akin to the paperclip problem, where if we tell an AI to do a largely unsolvable problem, what will it do? This axis seems less cut and dried than the persistence axis, but I included it because I think of Claude's 'user world model' as one of its strengths for general knowledge work like editing, slide creation, etc. Sometimes Claude does do totally random stuff because my prompt was underspecified, instead of asking me for clarification, and as the models get more powerful this 'just acting' could cause problems.

Read full article ↗excerpt · fair-use quotation

Around this claim