ATRIUMsearch → argument graph
ClaimVideo · 33:15 — 34:45

The larger a neural network gets, the less structured and less biased it should be — transformers dominate at trillion-parameter scale, while gated, biased architectures win at smaller, specialized scales.

Hasani argues there's a 'scale-to-bias' law: unstructured architectures like pure transformers win at massive scale, while smaller, resource-constrained models benefit from biased, gated, structured operators like those in liquid networks. ✦ AI generated

Ramin Hasani · The Cognitive Revolution · 2026-07-04 · original ↗

starts at this moment · 33:15

The larger neural networks that you make into infinite size, the more you want them to become less and less structured, and we've seen the success of transformers at the trillions of parameters... liquid neural networks, these alternative architectures, they have a sweet spot — there is a regime of parameters, up to let's say 100 billion, up to a trillion parameters, where you can just do better.

verbatim transcript · starts at 33:15

Transcript · around this moment

33:15talking about transformers being this revolutionary thing we're talking about maximum scale the reason why transformer architecture and attention mechanism is such a brilliant architecture is the fact that it is unstructured. There's no structure. When I when I'm talking about nested nonlinearities and all this crap that we have like in liquid normal networks, you don't have that in transformers, right? You have basically just matrix multiplication as the core

33:43functional kind of things that represent and it is unstructured. You can literally multiply any mat matrix into any matrix of any size. So the whole idea here is this. The larger neural networks that you make into infinite size. The larger neural networks you make the more want you want them to become less and less structured and we've seen like the success of transformers at the trillions of

34:08parameters. Now we are talking about tens of trillions of parameters like I mean the next generation of models that are going to come we're talking about trillions of and you can do that with a transformer architecture. As soon as you start adding a little bit of bias in that architecture at scale things become completely messed up. So we are talking about right now liquid neural networks like these alternative architectures

34:30that we we are talking about they have a scale to heat you know like there is a regime of parameters that you can just do better you know like let's say up to let's say 100 billion parameters up to a trillion parameter you know like that's kind of the range we are we are operating right now in this range the smaller is the model architecture the more you want to specialize them for a

34:52certain application to solve and they are actually like mathematically they're biased to actually solve a certain type of tasks better than any better than other types of architectures. So I would say biases on algorithms you know the more the more kind of biases you put on and what by biases what I mean is this adding a lot more nonlinearity adding like multiple gating levels on top of a

35:14neural network you know adding like recurrence you know like multiple different types of recurrence adding let's say convolutions like as a I mean convolution itself if you just keep the convolution or kind of neural networks they're also pretty unstructured because you can apply convolutions on any size of neural networks you know so that that that that question becomes like how much bias and what kind of problems you want

Around this claim