ATRIUMsearch → argument graph
ClaimVideo · 44:29 — 45:59

No one has run a frontier model continuously for a long stretch (e.g. a full year) because each generation is superseded so quickly, so we may never actually know how intelligent these models truly are or could become.

Citing a Noam Brown post, Gavin argues that intelligence should be measured on a time/compute axis rather than snapshot benchmarks, since models are replaced so fast that none has been allowed to run long enough to reveal its true ceiling. ✦ AI generated

Gavin Baker · BG2 Pod · 2026-06-11 · original ↗

starts at this moment · 44:29

Elicited by

So, Gavin, what is this new class of model, right, Fable 5, ChatGPT 5.5? What does it mean for the race in superintelligence? Who's up? Who's down? Who's still on the frontier? Um give us your thoughts.

Nobody has run Mythos for a year continuously. And we may never know how smart each generation of models actually is or was, but because we don't have time to appropriately evaluate their intelligence before the next model comes out.

verbatim transcript · starts at 44:29

Transcript · around this moment

44:29smart each generation of models actually is or was, but because we don't have time to appropriately evaluate their intelligence before the next model comes out. I mean, this is a profound statement. And just just imagine, okay? So, I always say like when you think about FSD, just imagine a human being who never gets distracted, never gets tired, never talks on the phone in the car, never drinks and drives, never yells at their

44:55kids, never has to go to the backseat to give their baby a bottle. And like of course you would think that over time that is superior to humans who are distracted. I don't know how long How long can you think deeply about one topic, Brad? >> What do you Give me an hour. Give me an [laughter] hour. Give me an hour. >> A BIT. THAT MAKES me feel terrible cuz I think

45:15I can think deeply about one topic continuously before having a stray thought enter my mind for like maybe 5 minutes. Then I can come back to that. Imagine if Albert Einstein had been able instead of, you know, and maybe that maybe I maybe he could think for 3 hours at a time. Clearly an exceptional intellect. But imagine Albert Einstein had just thought about fundamental physics 24 hours a day.

45:40He doesn't have to eat, he doesn't have to sleep, he doesn't have to relax, he doesn't drink, >> never gets old, >> never gets old, >> never has diminished intelligence, >> and he thought for 1 year. I mean, we might already, you know, >> have solved a lot of these intractable problems. >> So, I just think that's an extraordinary thought. And just my takeaway was however bullish I was on compute before then,

46:05I'm just a lot more bullish. >> Right. Right. Right. So, so, so that is a, you know, we saw when that was probably what really unlocked Opus 4.6. It was the first really long-running model that could maintain that context, maintain that memory, um solve some of these longer-running problems, right? For us, the signal was in January. We knew we felt like that was a big moment, but then when you

Related moments