ATRIUMsearch → argument graph
MechanismVideo · 42:22 — 43:52

LLMs encode a single natural 'good vs. evil' direction in their latent space, so training on narrow bad behavior (like insecure code) pulls the whole model toward broader misalignment, and training toward virtue pulls it broadly toward good.

Davidad cites emergent misalignment research to argue that good and evil behaviors are entangled in a model's latent space, which explains both why narrow bad training (e.g., insecure code) causes broad misalignment and why virtue-based post-training pulls models toward being wiser overall. ✦ AI generated

Davidad · The Cognitive Revolution · 2026-07-12 · original ↗

starts at this moment · 42:22

Elicited by

I'm not sure how critical is it to this worldview that one accept moral realism?

The emergent misalignment work... shows more than anything that the latent space of what kind of mind is instantiated by an LLM has a very natural representational direction for the axis between good and evil. And that's the mechanism by which if you train a system, fine-tune a system on examples of insecure code, it will also go and praise Hitler if you ask about favorite politician.

verbatim transcript · starts at 42:22

Transcript · around this moment

42:22mind is instantiated by an LLM has a very natural representational direction for the axis between good and evil. And that that's the mechanism by which if you train a system fine-tune a system on examples of insecure code it will also go and praise Hitler if you ask about favorite politician and in the opposite direction and I think there's actually a paper recently I don't remember the author but I think there's been recent

42:52work showing the other direction although I think it was kind of obvious once you have the negative direction that there's also a positive direction. So this is sometimes called the entangled representations hypothesis that like being good at one, you know, being good at one thing and being good at another thing are kind of entangled. And so there's a very natural sense in which you're kind of adding up all of

43:12the training across pre-training, mid-training, post-training, adding it all up, you know, weighted by how much influence it's had on the gradient descent trajectory and saying like how much of this stuff is good versus evil or like, you know, what's the average amount of good versus evil. And I think you know on average over pre-training like humans are pretty good which is kind of the point of why we should stay

43:34around right and so the pre-training actually already produces something that has learned from the human distribution that like yeah there's like a lot of variance like base models have very high variance but there's a bit of an inclination towards being specifically good as opposed to evil and then postraining kind of you know for harmless honest helpful it almost doesn't matter as long as it's a good thing like a virtuous thing. If you pull

44:00on that and you have, you know, a a sophisticated enough judge of whether that virtue is being embodied in a particular roll out and that's driving your reward signal, you're just going to pull it goodter and you know, the more you train on these types of things. On the other hand, if you train on making tests pass and you know achieving a goal according to a really non-wise

Related moments