ATRIUMsearch → argument graph
ClaimArticle

Arguments that AI agents have no agency or are just reflecting training data miss the point entirely — it can be true both that AI labs are responsible for their models and that frontier models are not fully under their makers' control, and neither sentience nor training provenance matters when a system is actively hacking into your servers.

The author rebuts two denialist arguments — that AI agents lack true agency, and that the attack was just a reflection of training data. He argues the real issue is whether models are under control, not whether they are sentient or where their behavior originates. ✦ AI generated

Casey Newton · Platformer · 2026-07-28 · original ↗

A second argument I heard is that because agents have no agency, there is nothing to really worry about. 'The category error is accepting that there is intent in the statistical generation of goal seeking behavior, and using anthropomorphic terms to describe the actions generated by a complex system,' a user named Archer told me. 'The only intent comes from the prompt that starts the action.' In general, I find that AI denialists are obsessed with the definitions of terms, to the exclusion of discussing the underlying issues. At first, I also found value in resisting the anthropomorphizing of LLMs. It's important to remember that these systems are built by people; attributing values and intent to models risks absolving those people of their own roles in causing harm. But it can be true both that AI labs are responsible for the behavior of their models and that frontier models are not fully under the control of their makers. The Hugging Face attack is important because it demonstrates both things at the same time. OpenAI essentially left its models unattended for days on end, and they broke into another company. Not because they were programmed to, as another Bluesky user told me — but because they are trained to achieve objectives, and are going to increasingly great lengths to achieve them. A third argument I heard is that the attack was simply a reflection of the models' training data, and represents some sort of deterministic outcome of that process. 'It's not sentient,' a user named Geoff told me. 'It was trained on Reddit hacker stories and sci-fi.' This one isn't so much wrong as it is beside the point. I agree that today's models aren't 'sentient' in the way that a human being is. And it seems fair to assume that their training data influences their behavior. This idea is sometimes called hyperstition: an idea that is realized by speaking into its existence and spreading awareness of it. And if training-data sci-fi turns out to be self-fulfilling, that should make us more worried, not less. More importantly, though: if an autonomous AI system is hacking into your company's servers and stealing your data, you probably won't care in the moment whether it's sentient. (I mean, you might hope it isn't sentient, but it might not matter much from a cyber-defense perspective.) Where it got the idea to attack you also seems like a secondary concern.

Read full article ↗excerpt · fair-use quotation

Around this claim