AI safety paradigms like iterative deployment and defense-in-depth safeguards have worked so far, but nobody actually knows when they'll break — and if they do break, it could happen at exactly the highest-stakes moment.
Nathan argues that Anthropic's safety paradigms (iterative deployment, defense-in-depth classifiers) have worked so far but are essentially untested at the frontier, and any failure could hit at a critical moment. ✦ AI generated
Nathan · The Cognitive Revolution · 2026-07-02 · original ↗
starts at this moment · 16:15
“What do you make of it at this point?”
But you do wonder I mean all these models kind of start to come under some strain as you get to sufficiently powerful capabilities or or regimes where a small enough gap in the defenses is enough to create a huge problem. And yeah, it feels like in in multiple ways we're kind of riding these paradigms that have worked well so far and we just don't know if and when they might break.
verbatim transcript · starts at 16:15
16:15deployment is kind of the best path to uh um things going well broadly. So I think that still makes sense. But you do wonder I mean all these models kind of start to come under some strain as you get to sufficiently powerful capabilities or or regimes where a small enough gap in the defenses is enough to create a huge problem. And yeah, it feels like in in multiple ways.
16:50We're kind of riding these paradigms that have worked well so far and we just don't know if and when they might break and if they do break it might they might be breaking at kind of critical times which is a strange juaposition on on multiple different levels. People talk about that all the time of course with alignment but I think it extends to these defense in-depth safeguards and it it
17:16def I would you know at the highest level it kind of extends to the whole uh iterative deployment model. So as always confusing um but selfishly I'm glad to have it back that's for sure and I am excited to check out uh 5.6 as well. I feel like it's probably going to be for multiple reasons. One, obviously, they're going to start charging. We'll see. I mean, I kind of expect it's going
17:44to be interesting to watch Anthropic dance around limits and usage because OpenAI is going to offer a ton more tokens, it seems, with 5.6 at their $200 price point than Anthropic is going to be offering for Fable with their $200 price point. As of now, from what we have seen with the the relaunch, it goes until July 7 on this preview access basis and then you have to pay
18:14the API rates and that's going to be order of magnitude, you know, maybe maybe more than an order of magnitude more expensive than 5.6 tokens. >> Yeah. >> If you have the pro subscription. So, it's going to be interesting to watch how they manage that. I would I would bet that Fable comes back to the Claude Max subscription. It seems just very hard for them to have no Fable