We may never actually know how intelligent frontier AI models are, because no one runs a model continuously long enough to properly evaluate it before the next generation replaces it.
Gavin Baker relays Noam Brown's point that intelligence evaluation can't keep pace with model releases, since nobody runs a model like Mythos continuously for a year, meaning true capability levels remain unknown. ✦ AI generated
Gavin Baker · BG2 Pod · 2026-06-11 · original ↗
starts at this moment · 44:29
“Say more about that. Why don't we know how smart they are?”
Because nobody has run Mythos for a year continuously. And we may never know how smart each generation of models actually is or was, but because we don't have time to appropriately evaluate their intelligence before the next model comes out. I mean, this is a profound statement.
verbatim transcript · starts at 44:29
44:29smart each generation of models actually is or was, but because we don't have time to appropriately evaluate their intelligence before the next model comes out. I mean, this is a profound statement. And just just imagine, okay? So, I always say like when you think about FSD, just imagine a human being who never gets distracted, never gets tired, never talks on the phone in the car, never drinks and drives, never yells at their
44:55kids, never has to go to the backseat to give their baby a bottle. And like of course you would think that over time that is superior to humans who are distracted. I don't know how long How long can you think deeply about one topic, Brad? >> What do you Give me an hour. Give me an [laughter] hour. Give me an hour. >> A BIT. THAT MAKES me feel terrible cuz I think
45:15I can think deeply about one topic continuously before having a stray thought enter my mind for like maybe 5 minutes. Then I can come back to that. Imagine if Albert Einstein had been able instead of, you know, and maybe that maybe I maybe he could think for 3 hours at a time. Clearly an exceptional intellect. But imagine Albert Einstein had just thought about fundamental physics 24 hours a day.
45:40He doesn't have to eat, he doesn't have to sleep, he doesn't have to relax, he doesn't drink, >> never gets old, >> never gets old, >> never has diminished intelligence, >> and he thought for 1 year. I mean, we might already, you know, >> have solved a lot of these intractable problems. >> So, I just think that's an extraordinary thought. And just my takeaway was however bullish I was on compute before then,
46:05I'm just a lot more bullish. >> Right. Right. Right. So, so, so that is a, you know, we saw when that was probably what really unlocked Opus 4.6. It was the first really long-running model that could maintain that context, maintain that memory, um solve some of these longer-running problems, right? For us, the signal was in January. We knew we felt like that was a big moment, but then when you
- ·Nobody runs a model continuously for a year
- ·Next generation arrives before evaluation finishes
- ·True capability levels stay unknown, per generation
- ·Evaluation can't keep pace with release speed
- ·Example: nobody has run Mythos a year straight
- ·Baker calls this 'a profound statement'