ATRIUMsearch → argument graph
ClaimVideo · 21:42 — 23:12

Because model iteration cycles are now shorter than the time it takes to test a model's true ceiling on long-running tasks, frontier labs genuinely don't know if today's models 'top out,' which is why OpenAI's Noam Brown has floated something like a model recall program.

Nathan relays Noam Brown's point that the pace of model releases has outrun the time needed to test models on very long-running tasks, so labs don't actually know if a model has topped out — raising the idea of a post-release model recall program. ✦ AI generated

Nathan · The Cognitive Revolution · 2026-07-08 · original ↗

starts at this moment · 21:42

Elicited by

I don't see how how you resolve that Nathan view.

he basically says we we don't know really if they top out and just calendar-wise from the time we're kind of done training it to our release time is enough to run a lot of standard tests but it's not enough to really do these long standing tests

verbatim transcript · starts at 21:42

Transcript · around this moment

21:42know long longunning task. So he's like you might need to give these things a month or a couple months or you know what happens if you spend a million dollars with one of these models. He's he basically says we we don't know really if they top out and just calendar-wise from the time we're kind of done training it to our release time is enough to run a lot of standard tests

22:05but it's not enough to really do these long standing tests. So I I had even heard him kind of propose something along the lines of like a you know like a clawback or sort of a recall model recall program almost where and obviously this doesn't work in open source but it can work in a API uh paradigm where a model might get released you know day n after it's kind of been

22:30deemed to be ready that gives you n days head start to be running models on really long time horizon tests and potentially you need to be ready as a frontier company to see at, you know, day n plus 30 or whatever that like actually we're now starting to see some problems when we get these things into the super long running regime. And so we either need to

22:58recall and fix or we need to which would mean taking things offline which obviously is going to be kind of painful but you know it's either that or you know have like a a longer delay to launch or just fly blind. Um, and I think that's quite interesting the idea that like the, you know, it's literally we've, you know, quite a tipping point where the the iteration

23:23cycle is just plain shorter than the testing time horizon is a very weird world to find ourselves in. Um, but yeah, I'm excited to to try it. I've seen mostly pretty affusive praise for 5.6. Um, you said that Fable is still better and Matt Schumer also said something similar. You know, he was like, "It's great, but by and large, Fable is still better on most tasks that

Around this claim