Because AI model release cycles are now faster than the time it takes to properly test a model on very long-horizon tasks, labs genuinely don't know whether their current frontier models have topped out before the next one ships.
Nathan relays Noam Brown's concern that iteration speed has outpaced the time needed to test long-horizon performance, meaning labs may not know a model's true ceiling, and floats the idea of a model 'recall' program if long-running tests later reveal problems. ✦ AI generated
Nathan · The Cognitive Revolution · 2026-07-08 · original ↗
starts at this moment · 21:42
“so this is one of his complaints um you know I I don't see how how you resolve that Nathan view.”
he's like you might need to give these things a month or a couple months or you know what happens if you spend a million dollars with one of these models. He's he basically says we we don't know really if they top out and just calendar-wise from the time we're kind of done training it to our release time is enough to run a lot of standard tests but it's not enough to really do these long standing tests.
verbatim transcript · starts at 21:42
21:42know long longunning task. So he's like you might need to give these things a month or a couple months or you know what happens if you spend a million dollars with one of these models. He's he basically says we we don't know really if they top out and just calendar-wise from the time we're kind of done training it to our release time is enough to run a lot of standard tests
22:05but it's not enough to really do these long standing tests. So I I had even heard him kind of propose something along the lines of like a you know like a clawback or sort of a recall model recall program almost where and obviously this doesn't work in open source but it can work in a API uh paradigm where a model might get released you know day n after it's kind of been
22:30deemed to be ready that gives you n days head start to be running models on really long time horizon tests and potentially you need to be ready as a frontier company to see at, you know, day n plus 30 or whatever that like actually we're now starting to see some problems when we get these things into the super long running regime. And so we either need to
22:58recall and fix or we need to which would mean taking things offline which obviously is going to be kind of painful but you know it's either that or you know have like a a longer delay to launch or just fly blind. Um, and I think that's quite interesting the idea that like the, you know, it's literally we've, you know, quite a tipping point where the the iteration
23:23cycle is just plain shorter than the testing time horizon is a very weird world to find ourselves in. Um, but yeah, I'm excited to to try it. I've seen mostly pretty affusive praise for 5.6. Um, you said that Fable is still better and Matt Schumer also said something similar. You know, he was like, "It's great, but by and large, Fable is still better on most tasks that