Because model iteration cycles have become shorter than the time needed to run truly long-horizon tests, AI labs genuinely don't know whether their models have topped out in capability, which could justify a post-release 'recall' testing program.
Nathan relays Noam Brown's argument that the pace of new model releases now outstrips the time it takes to discover a model's true long-horizon performance ceiling, raising the idea of a model 'recall' program to catch problems that only emerge over very long runs. ✦ AI generated
Nathan · The Cognitive Revolution · 2026-07-08 · original ↗
starts at this moment · 21:42
“I I don't see how how you resolve that Nathan view.”
he basically says we we don't know really if they top out and just calendar-wise from the time we're kind of done training it to our release time is enough to run a lot of standard tests but it's not enough to really do these long standing tests.
verbatim transcript · starts at 21:42
21:42know long longunning task. So he's like you might need to give these things a month or a couple months or you know what happens if you spend a million dollars with one of these models. He's he basically says we we don't know really if they top out and just calendar-wise from the time we're kind of done training it to our release time is enough to run a lot of standard tests
22:05but it's not enough to really do these long standing tests. So I I had even heard him kind of propose something along the lines of like a you know like a clawback or sort of a recall model recall program almost where and obviously this doesn't work in open source but it can work in a API uh paradigm where a model might get released you know day n after it's kind of been
22:30deemed to be ready that gives you n days head start to be running models on really long time horizon tests and potentially you need to be ready as a frontier company to see at, you know, day n plus 30 or whatever that like actually we're now starting to see some problems when we get these things into the super long running regime. And so we either need to
22:58recall and fix or we need to which would mean taking things offline which obviously is going to be kind of painful but you know it's either that or you know have like a a longer delay to launch or just fly blind. Um, and I think that's quite interesting the idea that like the, you know, it's literally we've, you know, quite a tipping point where the the iteration
23:23cycle is just plain shorter than the testing time horizon is a very weird world to find ourselves in. Um, but yeah, I'm excited to to try it. I've seen mostly pretty affusive praise for 5.6. Um, you said that Fable is still better and Matt Schumer also said something similar. You know, he was like, "It's great, but by and large, Fable is still better on most tasks that