The right approach to AI safety is not to prove an AI is safe, but to treat it like uranium: contain it in an engineered vessel and extract only artifacts that carry proofs of their own correctness.
Davidad explains the core idea behind Safeguarded AI: don't try to prove the AI is safe, contain the unsafe AI like radioactive material and only let out artifacts whose correctness has been formally proven. ✦ AI generated
Davidad (David Dalrymple) · The Cognitive Revolution · 2026-07-12 · original ↗
starts at this moment · 10:32
“What's the state of guaranteed safe AI today? Where are we on this process of trying to get some sort of at least soft guarantees around what AI will and won't do?”
The concept there is not that we would prove that some AI is safe, but that we would take AI which is not safe and treat it kind of like uranium which is not safe, put it into an engineered constructed containment vessel which makes the overall thing safe while also harnessing it to get stuff done that's economically valuable.
verbatim transcript · starts at 10:32
10:32of those agendas. That's the name of the program that Nora now leads. And the concept there is not that we would prove that some AI is safe, but that we would take AI which is not safe and treat it kind of like uranium which is not safe, put it into an engineered constructed containment vessel which makes the overall thing safe while also harnessing it to get stuff done that's economically
10:56valuable. And a lot of these cases that it is is now taking the form of you put the AI into a coding harness in a container and you have it produce some artifacts and you have it prove that those artifacts satisfy some criteria and then you take the artifact out of the container once it's proven and then you deploy that artifact and that's you know a piece of software potentially
11:18with some neural networks in it but like small neural networks that are just for doing one thing at a time so that you can check what they do but you're still taking advantage of the huge neural network because that's helping you to develop all of these small neural networks. So where we are in that is it's a long-term research program. I started you know I kind of wrote down
11:35the open agency architecture which was the original version of this agenda in 2022 and I said this is going to take 5 to 10 years and a lot of people thought that was a crazy short figure like Connor Lee was like no this will take 30 to 60 years like it's completely hopeless and I said no I think this could could be done in 5 to 10 years. So
11:53that's, you know, 2027 to 2032. Now seems like a kind of too late. You like it kind of we, you know, we kind of needed in order for this to be a strategy for for avoiding some extremely dangerous super intelligence existing being deployed, it would need to have been ready now. But what we can do is say, well, there's going to be a lot of aligned AI. I mean, that's what I that's
12:18part of what I'm saying. We'll get into that why I think there probably is going to be a lot of aligned AI. I also think there's going to be rogue AI and it's too late to avoid. But what we can do is provide aligned AI with tools that enable it to construct artifacts that are very reliable and that sort of they form a coalition that sort of defends