Even perfect constitutional alignment wouldn't drive doom odds as low as 5%, because solving alignment is only the price of admission to a game that still concentrates overwhelming competitive power in AI; a viable doom estimate cannot assume the alignment problem is the only problem.
Zvi pushes back on David Dalmple lowering his doom estimate below 5%, arguing that even if constitutional alignment worked perfectly, it wouldn't solve the fundamental problems of AI minds being far more competitive and efficient than human minds, competing for resources, and making decisions — so 5% is only a lower bound, not the actual estimate. ✦ AI generated
Zvi Mowshowitz · The Cognitive Revolution · 2026-08-05 · original ↗
starts at this moment · 108:05
even if you told me that alignment was perfectly solved I would not have a PDM as low as 5%. Or like even if you told me that the AIS are going to be aligned ... This does not solve the problem that AI minds are much more advanced and competitive and efficient than human minds. ... that is going to be subverted in any number of ways that this is all it does not solve your problems in a fundamental way is just the price of admission right like it's the right to play the game at all that you solve this problem and so even if P alignment is 95% there does not mean PD Doom is 5%. It means PD Doom is lower bounded at 5%.
verbatim transcript · starts at 108:05
108:10taking an order of magnitude less than the previous one. When you're on a planet, it's been around for four billion years. You know, with mammals that have been around for, you know, hundreds of millions of years, with reasonably intelligent things that have been around for a few million years, with agriculture and civilization that's been around for, you know, tens of thousands of years, with industry that's been around for hundreds of years, with
108:35AI, with AI and information age that's been around for on the order of, you know, maybe 50 years or 20 years or 10 years depending how you count. And LM have been around for five. And like we're already at this point where we're seeing production that is many times what it was at the start because of these multipliers, these force multipliers on the people who are working on it by their own self-reports.
108:56You'd be weird if this didn't have a finite sum that wasn't that large beyond where we already are, right? Like if you expect this to continue without something crazy happening for 20 years, I want to know why because like we really really should see something really really crazy happening pretty soon. And if you don't think we can handle that, maybe you should try and stop it from
109:22happening too quickly so that you can figure out how to handle it. Like we already have these amazing AI. If we were to diffuse Saul and Fable into the economy fully, it'd be transformational already. Like I don't think anybody who understands these technologies doesn't think that. So like we really really really don't need to keep pushing like the frontier that much further. And the more the frontier
109:50is jagged into things like a R&D itself, the stronger the argument comes that pushing farther is a bad idea because you're not missing out on that much of the other stuff relative to the amount you're missing out on AI R&D and recursive AI development. But the recursive AI development just goes boom, right? it goes not just boom in some sense like it goes like we don't know
110:16what happens the next but things go really fast if we're dealing with things on the order of a new model iteration every month and then every week and then every day with the kind of improvements we're seeing now like I don't think that ends up very often even if our alignment people are kind of on the ball and things are technically a lot more possible than I think they are I think
110:41it still seems like really scary and also like you're not giving up that much time slowing that down, right? Like we're talking about when we say pace the frontier, you know, pause gives the impression of a year from now we will have made no progress from where we are now. Like we'll get confused but we won't have fundamentally improved. We're talking about pacing here. We're talking about a year from now. We will
- ·Even perfect constitutional alignment would not yield P(doom) ≤ 5%
- ·Solving alignment is only the price of admission to the game
- ·It buys the right to play, not safety from the outcome
- ·AI minds are far more competitive and efficient than human minds
- ·They compete for resources and make consequential decisions
- ·Alignment is subverted in any number of ways downstream
- ·P(alignment) = 95% does not imply P(doom) = 5%
- ·P(doom) is lower bounded at 5% even with perfect alignment
- ·5% is a floor, not the actual estimate