We genuinely don't know how intelligent the newest AI models are because no one has run them continuously long enough to properly evaluate their capabilities before the next model ships.
Gavin Baker highlights Noam Brown's point that intelligence benchmarks are becoming obsolete because no one lets a frontier model run long enough to discover its true capability ceiling before it's superseded. ✦ AI generated
Gavin Baker · BG2 Pod · 2026-06-11 · original ↗
starts at this moment · 44:03
“Say more about that. Why don't we know how smart they are?”
Because nobody has run Mythos for a year continuously. And we may never know how smart each generation of models actually is or was, but because we don't have time to appropriately evaluate their intelligence before the next model comes out. I mean, this is a profound statement.
verbatim transcript · starts at 44:03
44:03after the Fable 5 release, and Mythos is evidently even better. But I just think that Gnome Brown post from yesterday, polynomial, is so profound. And just the idea that we do not know how smart these models are. And we made >> Say more about that. Why don't we know how smart they are? >> Because nobody has run Mythos for a year continuously. And we may never know how
44:29smart each generation of models actually is or was, but because we don't have time to appropriately evaluate their intelligence before the next model comes out. I mean, this is a profound statement. And just just imagine, okay? So, I always say like when you think about FSD, just imagine a human being who never gets distracted, never gets tired, never talks on the phone in the car, never drinks and drives, never yells at their
44:55kids, never has to go to the backseat to give their baby a bottle. And like of course you would think that over time that is superior to humans who are distracted. I don't know how long How long can you think deeply about one topic, Brad? >> What do you Give me an hour. Give me an [laughter] hour. Give me an hour. >> A BIT. THAT MAKES me feel terrible cuz I think
45:15I can think deeply about one topic continuously before having a stray thought enter my mind for like maybe 5 minutes. Then I can come back to that. Imagine if Albert Einstein had been able instead of, you know, and maybe that maybe I maybe he could think for 3 hours at a time. Clearly an exceptional intellect. But imagine Albert Einstein had just thought about fundamental physics 24 hours a day.
45:40He doesn't have to eat, he doesn't have to sleep, he doesn't have to relax, he doesn't drink, >> never gets old, >> never gets old, >> never has diminished intelligence, >> and he thought for 1 year. I mean, we might already, you know, >> have solved a lot of these intractable problems. >> So, I just think that's an extraordinary thought. And just my takeaway was however bullish I was on compute before then,
- ·No one runs a frontier model a year straight
- ·Next model ships before evaluation finishes
- ·True intelligence ceiling stays undiscovered
- ·Benchmarks can't capture capability without time
- ·Nobody has run Mythos continuously for a year
- ·We may never know each model's actual intelligence
- ·Not enough time to evaluate before next release
- ·Baker calls this 'a profound statement'