ATRIUMsearch → argument graph
PredictionVideo · 54:28 — 55:58

Pre-training scaling laws are fundamentally unlikely to stop working since they have held across 13 orders of magnitude of compute already; the real constraint will be the practical difficulty of testing ever-larger scales.

Nathan argues that pre-training scaling laws show no sign of breaking down after 13 orders of magnitude of compute, and expects them to keep holding — the bottleneck will be the logistics of testing bigger scales, not the underlying law. ✦ AI generated

Nathan Lambert · Lex Fridman · 2026-01-31 · original ↗

starts at this moment · 54:28

Elicited by

Your intuition about pre-training, if you scale the size of compute, will the models get better? Not whether it's financially viable but just from the law aspect of it, do you think the models will get smarter?

It's held for 13 orders of magnitude of compute, why would it ever end? So I think fundamentally it is pretty unlikely to stop, it's just eventually we're not even gonna be able to test the bigger scales because of all the problems that come with more compute.

verbatim transcript · starts at 54:28

Transcript · around this moment

54:28- Yeah. And I think that there's... And this sometimes comes off as almost like disillusionment from people, leadership at AI companies saying this, but they're like, "It's held for 13 orders of magnitude of compute, why would it ever end?" So I think fundamentally it is pretty unlikely to stop, it's just eventually we're not even gonna be able to test the bigger scales because of all the problems that come with more compute.

54:50I think that there's a lot of talk on how 2026 is a year when very large Blackwell compute clusters, like gigawatt-scale facilities at hyperscalers, are coming online. These were all contracts for power and data centers that were signed and sought out in 2022 and 2023. So before or right after ChatGPT. It took this two-to-three-year lead time to build these bigger clusters to train the models. While there's obviously

55:19immense interest in building even more data centers than that. So that is the crux that people are saying: these new clusters are coming. The labs are gonna have more compute for training. They're going to utilize this, but it's not a given. I've seen so much progress that I expect it, and I expect a little bit bigger models, and I expect... I would say it's more like we'll see a $2,000 subscription this year. We've seen

55:43$200 subscriptions. That could 10X again, and these are the kind of things that could come, and they're all downstream of this bigger model that offers just a little bit more cutting edge. - So, you know, it's reported that xAI is gonna hit that one-gigawatt scale early '26, and a full two gigawatts by year end. How do you think they'll utilize that in the context of scaling laws?

56:09Is a lot of that inference? Is a lot of that training? - It ends up being all of the above. So I think that all of your decisions when you're training a model come back to pre-training. So if you're going to scale RL on a model, you still need to decide on your architecture that enables this. We were talking about other architectures and using different types of attention, or a mixture of experts models.

Related moments