We may never truly know how intelligent each generation of AI models is or was, because no one has run a model continuously for long enough (e.g., a year) to properly evaluate it before the next model replaces it.
Referencing a Noam Brown post, Gavin Baker argues that because frontier models are replaced so quickly, nobody has time to actually evaluate their full capability, meaning true intelligence levels remain unknown — an insight he calls profound. ✦ AI generated
Gavin Baker · BG2 Pod · 2026-06-11 · original ↗
starts at this moment · 44:03
“So, Gavin, what is this new class of model, right, Fable 5, ChatGPT 5.5? What does it mean for the race in superintelligence? Who's up? Who's down? Who's still on the frontier? Um give us your thoughts.”
Because nobody has run Mythos for a year continuously. And we may never know how smart each generation of models actually is or was, but because we don't have time to appropriately evaluate their intelligence before the next model comes out. I mean, this is a profound statement.
verbatim transcript · starts at 44:03
44:03after the Fable 5 release, and Mythos is evidently even better. But I just think that Gnome Brown post from yesterday, polynomial, is so profound. And just the idea that we do not know how smart these models are. And we made >> Say more about that. Why don't we know how smart they are? >> Because nobody has run Mythos for a year continuously. And we may never know how
44:29smart each generation of models actually is or was, but because we don't have time to appropriately evaluate their intelligence before the next model comes out. I mean, this is a profound statement. And just just imagine, okay? So, I always say like when you think about FSD, just imagine a human being who never gets distracted, never gets tired, never talks on the phone in the car, never drinks and drives, never yells at their
44:55kids, never has to go to the backseat to give their baby a bottle. And like of course you would think that over time that is superior to humans who are distracted. I don't know how long How long can you think deeply about one topic, Brad? >> What do you Give me an hour. Give me an [laughter] hour. Give me an hour. >> A BIT. THAT MAKES me feel terrible cuz I think
45:15I can think deeply about one topic continuously before having a stray thought enter my mind for like maybe 5 minutes. Then I can come back to that. Imagine if Albert Einstein had been able instead of, you know, and maybe that maybe I maybe he could think for 3 hours at a time. Clearly an exceptional intellect. But imagine Albert Einstein had just thought about fundamental physics 24 hours a day.
45:40He doesn't have to eat, he doesn't have to sleep, he doesn't have to relax, he doesn't drink, >> never gets old, >> never gets old, >> never has diminished intelligence, >> and he thought for 1 year. I mean, we might already, you know, >> have solved a lot of these intractable problems. >> So, I just think that's an extraordinary thought. And just my takeaway was however bullish I was on compute before then,
- ·Frontier models get replaced before full evaluation completes
- ·No model has run continuously for a full year
- ·True intelligence levels remain effectively unknown
- ·Cites Noam Brown: no model evaluated a year continuously
- ·Next generation arrives before prior one is understood
- ·Baker calls this observation 'profound'