State space models rely on linear dynamics because current methods cannot properly scale nonlinear systems into parallelizable tensor computation.
Hasani explains that the core bottleneck in scaling architectures like liquid neural networks is that nonlinear relationships resist being cleanly tensorized for parallel computation, which is why state space models default to linear dynamics. ✦ AI generated
Ramin Hasani · The Cognitive Revolution · 2026-07-04 · original ↗
starts at this moment · 29:50
“So the obvious bitterpilled question would be how what if we take that exact paradigm and go to millions billions but what you must be hitting some bottlenecks along the way what are they?”
That's why the state space models that actually came out, these are all about around linear dynamics. You know, we're talking about like linear dynamical systems, right? And why those linear dynamical systems are important is just pure fact that we do not have proper ways to scale nonlinear systems.
verbatim transcript · starts at 29:50
29:50these are all about around linear dynamics. You know, we're talking about like linear dynamical systems, right? And why those linear dynamical systems are important is just pure fact that we do not have proper ways to scale nonlinear systems. You can apply nonlinearity to a to a linear dynamical systems. For example, in a let's say like a gated or let's say you like you can you can you can add a sigmoidal
30:15function after you perform the dynamical system like in a linear way you know like you compute the you compute all the matrices and everything like in parallel but then you can do a pointwise kind of application of a nonlinear operation on top of the entire tensor. That's what you can do. But what if the relationship between the parameters themselves are governed by some nonlinearity? So you cannot really use typical linear
30:38algebra. you you got to you got to always approximate that nonlinear system into a linear system then you would be able to actually parallelize the systems right does that does that make sense so that's the the fundamental bottleneck >> so how much does the closed form solution address that and how far has the with available computing resources how far has the original liquid network paradigm been able to scale so far to present?
31:13>> Fantastic question. So, one of the properties of liquid neural networks was the fact that they have multiple feedback mechanisms. You know, they're not just one format of feedback. You know, like they have like three layers of feedback like between two cells between in in the synaptic dynamic itself because we wanted to mimic like how how it is done in the brains, right? So, it has a multiple degrees of
31:33feedback. It has then a non basically nested nonlinearities on top of each other. That nested nonlinearity itself would also add a degrees of complexity. Even the closed form solution version of a liquid neural networks from a dynamics points of view, you can now instead of hundred neurons, you can have hundred thousands of neurons, maybe 1 million to 10 million neurons. But the nonlinearity it's you still have to compute the the