The arguments about misalignment and takeover are currently illegible, in-the-weeds, and hard to adjudicate, but over time empirical evidence and greater transparency should make them crisper—so we should have this conversation now rather than later.
In his closing remarks, Ryan argues that although today's misalignment arguments are complicated and hard to adjudicate—and he might be partly getting them wrong—they will become more easily testable and crisp over time, which is exactly why we should be reasoning about them now. ✦ AI generated
Ryan Greenblatt · Dwarkesh Podcast · 2026-08-11 · original ↗
plays this moment only · 114:40 — 117:40
I think right now a lot of the arguments for misalignment, AI takeover, all this crazy shit going down in the future, are illegible conceptual arguments that are extremely deep in the weeds and complicated and hard to adjudicate. Which means that maybe I'm getting a bunch of it wrong because it's really hard, and I'm trying to be uncertain... But it also means that over time, as we get more empirical evidence and better understand the nature of AI systems, it'll be easier to adjudicate a bunch of disagreements. It'll be more obvious what's going to happen. At least I hope.
verbatim transcript · starts at 114:40
Ryan Greenblatt
Another thing I want to note is that I think right now a lot of the arguments for misalignment, AI takeover, all this crazy shit going down in the future, are illegible conceptual arguments that are extremely deep in the weeds and complicated and hard to adjudicate. Which means that maybe I’m getting a bunch of it wrong because it’s really hard, and I’m trying to be uncertain. Obviously here I presented some specific scenarios, but those are not exhaustive. Probably the thing that actually happens is some more messy, confusing situation.
But it also means that over time, as we get more empirical evidence and better understand the nature of AI systems, it’ll be easier to adjudicate a bunch of disagreements. It’ll be more obvious what’s going to happen. At least I hope. Maybe the AIs will be able to help us with the epistemics and understanding what’s going on, if we can actually align them well so they try to help us.
Even if the arguments are complicated now, this would have been even harder six years ago, even though the shape of the arguments would have looked broadly pretty similar. Hopefully before it’s too late, this whole thing will become more crisp and clear, and we can all notice these problems and intervene.