ATRIUMsearch → argument graph
Video · 2026-07-04 · 1h 48m · 42 moments

Intelligence on the Edge: Liquid AI's Ramin Hasani on the Search for Device-Native Foundation Models

✦ AI generated

timeline · colored by role

01
Context

The C. elegans worm's 300-neuron nervous system achieves control and dexterity that outperforms the best robotic systems, which is what inspired Liquid AI's approach of modeling neural networks directly on biological neuron dynamics.

Ramin Hasani explains that Liquid AI's origins trace to studying the C. elegans worm, whose 300-neuron nervous system outperforms robotic control systems, motivating a biologically-grounded approach to neural network design.

transcript

Ramin Hasani: this animal exhibits massive amount of kind of sensory reactive kind of behavior like amazing levels of control with 300 cells in its nervous system, you know, and that's was fascinating for us because this is much smaller than any neural network that performs control at that time like on on let's say like autonomous systems, you know, and it's wor can do better better, you know, dextrous movements and stuff

02
Example

Extremely small liquid neural networks—as few as 12, 19, and 30 neurons—were sufficient to autonomously parallel park a car, drive a car, and fly a drone, respectively.

Ramin Hassani recounts early MIT results showing tiny liquid neural networks with only a dozen or so neurons could handle real-world control tasks like parking and flying, far outperforming expectations for such small models.

transcript

Ramin Hasani: I mean early on the results were fascinating like we saw that with 12 neurons with with actually 12 neurons you could you could parallel park autonomously like a car you know like a small car with 19 neurons you could drive a car with with 30 neurons you can fly autonomously like navigating kind of a drone

explains mechanism · 1

03
Anecdote

With as few as 12 to 30 liquid neurons, we demonstrated a car autonomously parallel parking, driving, and a drone flying — control tasks normally requiring vastly larger neural networks.

Ramin Hasani recounts early MIT results showing that tiny liquid neural networks — as few as 12 neurons — could autonomously park, drive, or fly a drone, far outperforming what network size would suggest.

transcript

Ramin Hasani: with 12 neurons with with actually 12 neurons you could you could parallel park autonomously like a car you know like a small car with 19 neurons you could drive a car with with 30 neurons you can fly autonomously like navigating kind of a drone

04
Anecdote

With just 12 liquid neurons a neural network can autonomously parallel park a car, with 19 neurons it can drive a car, and with 30 neurons it can fly a drone.

Hasani recounts early MIT results showing that tiny liquid neural networks with only a handful of neurons could perform control tasks like autonomous parking, driving, and drone flight.

transcript

Ramin Hasani: So we I mean early on the results were fascinating like we saw that with 12 neurons with with actually 12 neurons you could you could parallel park autonomously like a car you know like a small car with 19 neurons you could drive a car with with 30 neurons you can fly autonomously like navigating kind of a drone

05
Example

With just 12 liquid neurons a system could autonomously parallel park a car; 19 neurons could drive a car; 30 neurons could autonomously fly a drone.

Hassani recounts early MIT results showing liquid neural networks, inspired by the 300-neuron nervous system of the C. elegans worm, could control cars and drones with only a handful of neurons.

transcript

Ramin Hasani: Early on the results were fascinating. We saw that with 12 neurons you could parallel park autonomously like a car, like a small car. With 19 neurons you could drive a car. With 30 neurons you can fly autonomously, navigating kind of a drone, and process this information with this a little bit more complex and elaborate version of these neural dynamics, which we called liquid neural networks.

06
Mechanism

In 2022, Liquid AI's team found the first-ever closed-form solution to the differential equations governing liquid neural network dynamics, a problem unsolved since Louis Lapicque's 1907 membrane potential equations, enabling the models to scale from hundreds to billions of neurons without numerical solvers.

Liquid AI solved a century-old open problem in neuronal dynamics equations in closed form, a breakthrough published in Nature Machine Intelligence that let them scale liquid networks from hundreds to billions of parameters.

transcript

Ramin Hasani: What if we just solve the whole system in closed form? Turns out the closed form solution for this type of equation hasn't existed since 1907. Around 2022, we solved the liquid neural network kind of interaction of neurons with each other in closed form for the first time. This became a Nature Machine Intelligence paper published in November of 2022.

07
Mechanism

We solved the closed-form solution for liquid neural network neuron dynamics in 2022, a problem that had gone unsolved since 1907, and this removed the need for numerical solvers so the networks can scale to billions of neurons.

Hassani describes how Liquid AI derived, for the first time in 2022, a closed-form solution to the differential equations governing liquid neural network dynamics — a problem open since 1907 — enabling scaling from hundreds of neurons to billions.

transcript

Ramin Hasani: we for the first time we actually solved that and in this was like 2022 Around 2022, we solved the liquid neural network kind of interaction of neurons with each other in closed form for the first time. This became a nature machine intelligence paper uh published in November of 2022. This is called the closed form continuous time uh systems you know like this is this liquid neural networks in closed form and uh the closed form I mean it has massive implications.

08
Fact

In 2022, Liquid AI solved the liquid neural network differential equations in closed form, a problem that had no known solution since 1907, which removed the need for numerical solvers and allowed scaling from hundreds to billions of neurons.

Hasani describes a 2022 breakthrough (published in Nature Machine Intelligence) that solved liquid neural network dynamics in closed form for the first time, eliminating numerical solvers and unlocking massive scalability.

transcript

Ramin Hasani: This became a nature machine intelligence paper uh published in November of 2022. This is called the closed form continuous time uh systems you know like this is this liquid neural networks in closed form and uh the closed form I mean it has massive implications. Why? Because now I don't need to use any numerical solvers to actually run a liquid neural networks.

09
Fact

In 2022, Liquid AI's team found the first-ever closed-form solution to the neuron-interaction differential equations underlying liquid neural networks, a problem rooted in a 1907 membrane-potential equation that had remained open for over a century.

Hassani describes solving, for the first time, the closed-form solution to liquid neural network dynamics—tracing the underlying math back to a 1907 equation for membrane potential—which was published in Nature Machine Intelligence in late 2022.

transcript

Ramin Hasani: Around 2022, we solved the liquid neural network kind of interaction of neurons with each other in closed form for the first time. This became a nature machine intelligence paper uh published in November of 2022. This is called the closed form continuous time uh systems you know like this is this liquid neural networks in closed form

10
Fact

In 2022 we found the first-ever closed-form solution to the liquid neural network differential equations — a problem open since 1907 — which let us scale from hundreds of neurons to billions without numerical solvers.

Hasani describes solving, for the first time, the closed-form dynamics of liquid neural networks in 2022 — a mathematical problem dating to 1907 — which removed the need for numerical solvers and unlocked massive scaling.

transcript

Ramin Hasani: This became a nature machine intelligence paper published in November of 2022. This is called the closed form continuous time systems... the closed form I mean it has massive implications. Why? Because now I don't need to use any numerical solvers to actually run a liquid neural networks. I can now have not only hundreds of neurons but now I can have billions of neurons next to each other.

11
Fact

Liquid AI found the closed-form solution to liquid neural network neuron dynamics in 2022, resolving a mathematical problem for this class of differential equations that had remained open since 1907.

Hasani describes solving in closed form a class of differential equations modeling neuron dynamics that had lacked a known closed-form solution since Louis Lapicque's 1907 work, published as a Nature Machine Intelligence paper in 2022.

transcript

Ramin Hasani: in this was like 2022 Around 2022, we solved the liquid neural network kind of interaction of neurons with each other in closed form for the first time. This became a nature machine intelligence paper uh published in November of 2022. This is called the closed form continuous time uh systems you know like this is this liquid neural networks in closed form

12
Fact

Liquid AI found the first closed-form solution for the differential equations governing liquid neural network neuron dynamics in 2022, a problem that had remained mathematically open since 1907.

Ramin Hassani describes how his team found, in 2022, the first-ever closed-form solution to the neuron-dynamics equations underlying liquid neural networks—unlocking the ability to scale from hundreds to billions of neurons.

transcript

Ramin Hassani: Around 2022, we solved the liquid neural network kind of interaction of neurons with each other in closed form for the first time. This became a nature machine intelligence paper uh published in November of 2022. This is called the closed form continuous time uh systems you know like this is this liquid neural networks in closed form and uh the closed form I mean it has massive implications.

13
Mechanism

Nonlinear neural network dynamics cannot be trivially converted into parallelizable tensor computations, which is the fundamental scaling bottleneck for liquid neural networks and the reason state space models were built on linear dynamics instead.

Hassani explains that the core technical bottleneck in scaling liquid neural networks is that nonlinear relationships resist being reshaped into parallel tensor/matrix computations, which is why alternative architectures like state space models chose linear dynamics.

transcript

Ramin Hassani: The main challenge is turning sequential computation into parallel computation. So when you have degree like when you start like going from a single neuron dynamics to multiple neural dynamics weight parameters of your system instead of being a scalers or vectors they become matrices and tensors. So now you're talking about matrix multiplication.

14
Mechanism

State space models rely on linear dynamics because current methods cannot properly scale nonlinear systems into parallelizable tensor computation.

Hasani explains that the core bottleneck in scaling architectures like liquid neural networks is that nonlinear relationships resist being cleanly tensorized for parallel computation, which is why state space models default to linear dynamics.

transcript

Ramin Hasani: That's why the state space models that actually came out, these are all about around linear dynamics. You know, we're talking about like linear dynamical systems, right? And why those linear dynamical systems are important is just pure fact that we do not have proper ways to scale nonlinear systems.

15
Claim

The larger a neural network gets, the less structured and less biased it should be — transformers dominate at trillion-parameter scale, while gated, biased architectures win at smaller, specialized scales.

Hasani argues there's a 'scale-to-bias' law: unstructured architectures like pure transformers win at massive scale, while smaller, resource-constrained models benefit from biased, gated, structured operators like those in liquid networks.

transcript

Ramin Hasani: The larger neural networks that you make into infinite size, the more you want them to become less and less structured, and we've seen the success of transformers at the trillions of parameters... liquid neural networks, these alternative architectures, they have a sweet spot — there is a regime of parameters, up to let's say 100 billion, up to a trillion parameters, where you can just do better.

gives example · 1

16
Claim

Transformers dominate at massive scale precisely because attention is mathematically unstructured, whereas the nested nonlinearities and gating structures liquid networks and other alternatives rely on only help at smaller, more specialized scales.

Hassani argues transformers win at the largest scales because attention is fundamentally 'unstructured' pure matrix multiplication, while structured, biased architectures like liquid networks only provide an edge at smaller scales.

transcript

Ramin Hasani: the reason why transformer architecture and attention mechanism is such a brilliant architecture is the fact that it is unstructured. There's no structure. When I when I'm talking about nested nonlinearities and all this crap that we have like in liquid normal networks, you don't have that in transformers, right? You have basically just matrix multiplication as the core functional kind of things that represent and it is unstructured.

17
Claim

As neural networks scale up in size, the optimal design becomes progressively less structured, which is why unbiased architectures like transformers dominate at the largest scales.

Hasani lays out a scale-dependent theory of architecture: small, specialized models benefit from added structural biases like gating, while at massive scale, unstructured operators like pure attention or matrix multiplication win out.

transcript

Ramin Hasani: The whole idea here is this. The larger neural networks that you make into infinite size. The larger neural networks you make the more want you want them to become less and less structured and we've seen like the success of transformers at the trillions of parameters.

18
Claim

As neural networks scale up, the optimal design becomes progressively less structured and less biased, which is why unstructured attention-based transformers dominate at massive scale while adding architectural biases badly hurts performance at that scale.

Hasani argues that transformers dominate at the largest scales precisely because they are unstructured, and that the ideal degree of architectural bias (gating, recurrence, nonlinearity) is inversely related to model scale.

transcript

Ramin Hasani: The larger neural networks that you make into infinite size. The larger neural networks you make the more want you want them to become less and less structured and we've seen like the success of transformers at the trillions of parameters.

19
Claim

The optimal degree of architectural bias (gating, structure, specialized operators) is a function of model scale: larger models perform best with unstructured, unbiased operators like pure attention or matmul, while smaller models benefit from added structural biases.

Hassani argues that as neural networks scale up, unstructured operators like transformers' attention win out, whereas adding hand-designed biases like gating actually hurts performance at very large scale—an insight drawn from Liquid's massive architecture search.

transcript

Ramin Hasani: The larger neural networks you make the more want you want them to become less and less structured and we've seen like the success of transformers at the trillions of parameters. Now we are talking about tens of trillions of parameters like I mean the next generation of models that are going to come we're talking about trillions of and you can do that with a transformer architecture.

gives example · 1explains mechanism · 1

20
Claim

The optimal amount of architectural bias (gating, nonlinearity, structure) in a neural network is inversely related to model scale: smaller, specialized models benefit from more biased/structured operators, while at very large scale (up to trillions of parameters) unstructured, unbiased architectures like pure attention and matrix multiplication win.

Hassani explains Liquid's key architectural finding: bias-heavy, structured architectures (gating, nonlinearity) win at small scale, but as models grow toward trillions of parameters, unstructured operators like pure attention and matrix multiplication dominate.

transcript

Ramin Hasani: The larger neural networks you make, the more you want them to become less and less structured, and we've seen the success of transformers at the trillions of parameters... we are talking about liquid neural networks, these alternative architectures — they have a scale to it. There is a regime of parameters that you can just do better, let's say up to 100 billion parameters, up to a trillion parameters. The smaller the model architecture, the more you want to specialize them for a certain application.

gives example · 1

21
Mechanism

Liquid AI built an automated foundation model design system (AFMD) that runs hardware-in-the-loop evolutionary search evaluated on real downstream applications across roughly 100 benchmarks, rather than proxy metrics like perplexity, specifically to eliminate human bias from architecture design.

Hassani describes AFMD, Liquid AI's in-house automated foundation model design system, which uses an evolution strategy with real hardware and downstream-task evaluation—rather than perplexity—to search for the best architecture for a given memory, latency, and speed constraint.

transcript

Ramin Hassani: There's a system that we developed in house. We call it automated automated foundation model design. You know, AFMD, you know, like that's that's a meta metal learning system that explo that puts a hardware in the loop and then tries out like many different operators with an evolution strategy. The criteria is an evolution strategy optimizing for a couple of things. optimizing for memory consumption on that device, optimizing for latency, optimizing for speed.

22
Mechanism

Liquid AI built an in-house automated architecture search system (AFMD) that puts target hardware in the loop and evaluates candidate architectures on real downstream application performance across ~100 benchmarks rather than proxy metrics like perplexity, specifically to remove human bias from architecture design.

Hassani describes AFMD, Liquid's automated foundation model design system, which uses an evolution strategy with real hardware and downstream-task evaluation to eliminate human bias from architecture choices — a deliberate reaction against how architecture decisions are typically made ad hoc by a small group of experts at frontier labs.

transcript

Ramin Hasani: There's a system that we developed in house. We call it automated foundation model design, AFMD — that's a meta-learning system that puts a hardware in the loop and then tries out many different operators with an evolution strategy... optimizing for memory consumption on that device, optimizing for latency, optimizing for speed, while no sacrifice on quality. When we talk about quality, perplexity is not the measure. It's actually the downstream applications that we care about... we're talking about 100 different benchmarks.

23
Mechanism

Liquid AI built an in-house automated system called AFMD that uses hardware-in-the-loop evolutionary search over dozens of operators, optimizing for memory, latency, and speed on real downstream tasks, specifically to remove human bias from architecture design.

Hassani explains Liquid's automated foundation model design (AFMD) system, a meta-learning process with hardware in the loop that evolves architectures against real downstream tasks and hardware constraints rather than proxy metrics, explicitly to eliminate human bias.

transcript

Ramin Hasani: There's a system that we developed in house. We call it automated automated foundation model design. You know, AFMD, you know, like that's that's a meta metal learning system that explo that puts a hardware in the loop and then tries out like many different operators with an evolution strategy. The criteria is an evolution strategy optimizing for a couple of things. optimizing for memory consumption on that device, optimizing for latency, optimizing for speed.

24
Mechanism

Liquid AI's automated architecture search system (AFMD) removes human bias by running a hardware-in-the-loop evolutionary search that optimizes for memory, latency, and speed while evaluating quality on real downstream tasks rather than proxy metrics like perplexity.

Hasani describes AFMD, Liquid AI's in-house automated foundation model design system, which uses an evolution strategy with hardware in the loop to search architectures, judged on downstream task performance across ~100 benchmarks rather than perplexity.

transcript

Ramin Hasani: There's a system that we developed in house. We call it automated automated foundation model design. You know, AFMD, you know, like that's that's a meta metal learning system that explo that puts a hardware in the loop and then tries out like many different operators with an evolution strategy. The criteria is an evolution strategy optimizing for a couple of things. optimizing for memory consumption on that device, optimizing for latency, optimizing for speed.

25
Mechanism

Liquid AI built an automated, hardware-in-the-loop architecture search system (AFMD) that uses an evolution strategy across dozens of operators to remove human bias from architecture design, optimizing directly for memory, latency, and speed on target hardware.

Hasani describes Liquid AI's in-house AFMD (automated foundation model design) system, which searches architecture space with hardware in the loop to eliminate the human biases that typically shape design decisions at foundation model labs.

transcript

Ramin Hasani: There's a system that we developed in house. We call it automated automated foundation model design. You know, AFMD, you know, like that's that's a meta metal learning system that explo that puts a hardware in the loop and then tries out like many different operators with an evolution strategy. The criteria is an evolution strategy optimizing for a couple of things. optimizing for memory consumption on that device, optimizing for latency, optimizing for speed.

26
Mechanism

We built an automated, hardware-in-the-loop evolutionary architecture search system (AFMD) specifically to remove human bias from foundation model design, since even top labs rely on small groups of researchers making subjective architecture calls.

Hasani explains Liquid AI's in-house AFMD system — an evolutionary, hardware-in-the-loop meta-learning search over dozens of architecture operators — designed to eliminate the human bias he says pervades architecture decisions even at top labs.

transcript

Ramin Hasani: There's a system that we developed in house. We call it automated foundation model design, AFMD — that's a meta learning system that puts a hardware in the loop and then tries out many different operators with an evolution strategy... optimizing for memory consumption on that device, optimizing for latency, optimizing for speed, while no sacrifice on quality.

27
Mechanism

Liquid AI built an automated, hardware-in-the-loop architecture search system (AFMD) specifically to eliminate human bias in architecture design, because even at top labs a small group of people is really just calling the shots based on personal intuition.

Hassani explains that Liquid built its AFMD automated architecture-search system with hardware-in-the-loop evolutionary optimization specifically to remove the human biases he says still drive architecture decisions even at top labs like Anthropic and OpenAI.

transcript

Ramin Hasani: we wanted to remove all the human biases early on as we are actually like building architectures. One of the things that we realize like culturally at companies like this is what I can tell you like even at largest foundation model labs uh in in the US right now entropic and open AI there are a bunch of people like people that are coming from the science like I call them the avengers of the architectures.

28
Claim

A year and a half before Mamba, our Liquid-S4 paper first introduced input-dependent SSMs — the gating mechanism that later became central to Mamba and other sequence architectures originated in our liquid neural network research.

Hasani traces the input-dependent gating mechanism now central to architectures like Mamba back to Liquid AI's own Liquid-S4 paper, published a year and a half earlier, framing it as a foundational discovery rooted in liquid neural network theory.

transcript

Ramin Hasani: That's a liquid structure — that's when I say this input dependent kind of thing that is actually coming into Mamba. So, not right before, but a year and a half before Mamba came out, we released a paper called Liquid S4... we were actually for the first time introducing this idea of input dependent SSMs.

29
Claim

Liquid AI published the Liquid S4 paper introducing input-dependent state space models a year and a half before Mamba was released, making the gating/input-dependence mechanism now central to many architectures an original Liquid AI contribution.

Hasani traces the input-dependent gating mechanism now common across architectures like Mamba back to Liquid AI's own Liquid S4 paper, published well before Mamba.

transcript

Ramin Hasani: a year and a half before Mamba come come out we released a paper called liquid S4 you know this paper Because if you just read the abstract of that paper, you know, like we are actually for the first time we're introducing this idea of input dependent SSMs. Basically input dependent SSM.

30
Fact

Liquid AI's Liquid S4 paper introduced input-dependent state space models a year and a half before Mamba, making input-dependent gating one of the company's foundational discoveries.

Hassani notes that Liquid AI's 'Liquid S4' paper introduced input-dependent state space models roughly eighteen months before Mamba, framing gating and input-dependence as a discovery the company made early and carried into later architectures.

transcript

Ramin Hasani: right before Mamba actually got out you know like actually not right before but a year and a half before Mamba come come out we released a paper called liquid S4 you know this paper because if you just read the abstract of that paper, you know, like we are actually for the first time we're introducing this idea of input dependent SSMs.

31
Claim

At smaller, specialized model scales, biased/structured components like gating mechanisms improve performance, but as models scale up they benefit from becoming maximally unstructured (pure attention, pure matrix multiplication)—architectural bias is a function of scale, not a universal good.

Confirming the host's reframe of 'attention is all you need,' Hassani explains that attention remains essential at scale, but gating on the remaining layers matters more than any specific exotic architecture—and that the right amount of architectural bias is a function of model size.

transcript

Ramin Hassani: No, I mean you're touching on the right things and this is basically it's like the regime you're operating in the goal of your system. Are you trying to build super intel? Are you trying to build like the most powerful version of the AI system? You need the most unbiased version of an algorithm. Now attention is a extremely rich unbiased format of algorithms.

32
Data

The mobile device market, worth about $500 billion a year and roughly as large as the AI data center buildout, is a massive, largely untapped opportunity for efficient on-device intelligence.

Hassani highlights that the mobile device market (~$500B/year) rivals the scale of the data center AI buildout, framing it as an enormous, underexploited substrate for on-device intelligence that Liquid is targeting.

transcript

Ramin Hasani: that market itself by the way just the mobile kind of business is $500 billion market is absolutely insane and um it is as as big as kind of the data center kind of market so you can imagine like There is there's a parallel here to be made like you know the efficient markets you know and also constrained intelligence market it's something that liquidi is going after.

33
Data

The on-device compute market—smartphones and laptops combined—represents roughly a trillion dollars of underutilized 'dark compute,' with the mobile business alone being as large as the AI data center market.

Hassani highlights that the mobile device market alone (~$500B) rivals the AI data center market in size, framing the vast installed base of underused edge compute as a massive opportunity for efficient, on-device intelligence.

transcript

Ramin Hasani: you've seen like that that market itself by the way just the mobile kind of business is $500 billion market is absolutely insane and um it is as as big as kind of the data center kind of market so you can imagine like There is there's a parallel here to be made like you know the efficient markets you know and also constrained intelligence market it's something that liquidi is going after

34
Data

The global mobile device business alone is a $500 billion market, roughly as large as the entire AI data center market, representing a massive underexploited substrate of 'dark compute' for on-device intelligence.

Hassani highlights that the mobile device market ($500B) rivals the data center AI buildout in size, underscoring the massive commercial opportunity in efficient, on-device 'dark compute' that Liquid AI is targeting.

transcript

Ramin Hasani: That market itself, by the way, just the mobile kind of business, is $500 billion market — is absolutely insane. And it is as big as kind of the data center kind of market. So you can imagine there is a parallel here to be made, like the efficient markets, and also constrained intelligence market — it's something that liquidi is going after.

35
Data

The global smartphone market alone is roughly $500 billion annually, comparable in scale to the AI data center buildout, representing a massive underused compute substrate that on-device intelligence can tap into.

Hassani points out that the mobile device market is roughly as large as the data center market in dollar terms, framing edge devices as a huge, underexploited substrate for AI that Liquid AI is targeting.

transcript

Ramin Hassani: you've seen like that that market itself by the way just the mobile kind of business is $500 billion market is absolutely insane and um it is as as big as kind of the data center kind of market so you can imagine like there is there's a parallel here to be made like you know the efficient markets you know and also constrained intelligence market it's something that liquidi is going after

36
Data

The mobile device market (~$500 billion annually) is comparable in scale to the AI data center market, representing a huge, largely untapped substrate of 'dark compute' for on-device intelligence.

Hasani points out that the mobile device market is roughly as large as the AI data center market, framing edge devices as a massive underutilized opportunity for efficient, on-device AI.

transcript

Ramin Hasani: you've seen like that that market itself by the way just the mobile kind of business is $500 billion market is absolutely insane and um it is as as big as kind of the data center kind of market so you can imagine like There is there's a parallel here to be made like you know the efficient markets you know and also constrained intelligence market it's something that liquidi is going after

37
Data

The mobile phone and laptop market, worth about $500 billion a year, is as large as the AI data center market, representing a massive, largely untapped opportunity for on-device constrained intelligence.

Hasani frames the roughly $500 billion annual mobile device market as an enormous, underexploited parallel opportunity to the data center AI buildout — the 'constrained intelligence market' that Liquid AI is targeting.

transcript

Ramin Hasani: you've seen that market itself, by the way — just the mobile kind of business is a $500 billion market, is absolutely insane, and it is as big as the data center kind of market. So you can imagine there's a parallel here to be made — the efficient markets, and also constrained intelligence market — it's something that Liquid is going after.

38
Prediction

Current AI algorithms will not achieve the human brain's intelligence-per-watt efficiency because human intelligence emerged from an immensely long evolutionary process combining diverse learning mechanisms beyond next-token prediction.

Asked about the ultimate limits of miniaturized on-device intelligence, Hassani argues today's algorithms can't match the brain's efficiency because human intelligence is the product of eons of evolution layering multiple learning mechanisms, not just next-token-style prediction.

transcript

Ramin Hasani: I don't believe with the current set of algorithms like we would be able to get close to the let's say intelligence per watt that human brain is actually like providing right you're not going to get there why because I believe human brain over the we also have to like consider the amount of energy that went into design of humans you know as a whole biological evolution is actually a very very long kind of process

39
Prediction

Current AI algorithms will not reach the human brain's intelligence-per-watt efficiency, because the brain's efficiency is the product of an immensely long evolutionary process, not just something replicable through next-token prediction at smaller scale.

Hassani predicts that with today's algorithmic paradigm, on-device AI won't match the brain's intelligence-per-watt efficiency, since human intelligence reflects eons of evolutionary 'pre-training' and multiple learning mechanisms beyond emergent in-context learning from next-token prediction.

transcript

Ramin Hasani: I don't believe with the current set of algorithms like we would be able to get close to the let's say intelligence per watt that human brain is actually like providing right you're not going to get there why because I believe human brain over the we also have to like consider the amount of energy that went into design of humans you know as a whole biological evolution is actually a very very long kind of process.

rebuts · 1

40
Prediction

Current AI algorithms will not reach the intelligence-per-watt efficiency of the human brain, because human intelligence emerged from a diverse set of mechanisms shaped by biological evolution, not merely from next-token prediction.

Hasani argues that today's algorithmic paradigm, centered on next-token prediction, cannot match the brain's energy efficiency because human intelligence draws on multiple learning mechanisms (reinforcement learning, simulation, Bayesian reasoning) forged over evolutionary time.

transcript

Ramin Hasani: I don't believe with the current set of algorithms like we would be able to get close to the let's say intelligence per watt that human brain is actually like providing right you're not going to get there why because I believe human brain over the we also have to like consider the amount of energy that went into design of humans

41
Prediction

Even with continued scaling, current AI algorithms will not reach the human brain's intelligence-per-watt efficiency, because human intelligence resulted from a long evolutionary process that encoded multiple learning algorithms, not just the emergent, gradient-descent-like in-context learning that arises from next-token prediction.

Asked about the ultimate limits of miniaturized on-device intelligence, Hassani argues today's algorithms won't match the brain's efficiency because human intelligence is the product of a vastly long evolutionary process that embedded many distinct learning mechanisms, not just next-token prediction.

transcript

Ramin Hassani: I don't believe with the current set of algorithms like we would be able to get close to the let's say intelligence per watt that human brain is actually like providing right you're not going to get there why because I believe human brain over the we also have to like consider the amount of energy that went into design of humans you know as a whole biological evolution is actually a very very long kind of process

42
Prediction

With current AI algorithms, it will not be possible to reach the intelligence-per-watt efficiency of the human brain, because that efficiency is the product of a vastly long biological evolutionary process encoding multiple diverse learning mechanisms (in-context learning, reinforcement learning, mental simulation, Bayesian inference) beyond what next-token prediction alone can produce.

Hassani argues that today's AI paradigm, dominated by next-token prediction, cannot approach the brain's intelligence-per-watt because human intelligence emerged from eons of evolution encoding multiple distinct learning algorithms, not just one emergent gradient-descent-like mechanism.

transcript

Ramin Hasani: I don't believe with the current set of algorithms we would be able to get close to the intelligence per watt that human brain is actually providing. You're not going to get there. Why? Because I believe human brain — we also have to consider the amount of energy that went into design of humans... biological evolution is actually a very, very long kind of process.

Highlight slides
Liquid Neural Networks: Extreme Efficiency✦ from: With just 12 liquid neurons a neural network can autonomously parallel park a car, with 19 neurons it can drive a car, and with 30 neurons it can fly a drone.Why It Mattered✦ from: With just 12 liquid neurons a neural network can autonomously parallel park a car, with 19 neurons it can drive a car, and with 30 neurons it can fly a drone.Liquid Neural Networks: Solved in Closed Form✦ from: We solved the closed-form solution for liquid neural network neuron dynamics in 2022, a problem that had gone unsolved since 1907, and this removed the need for numerical solvers so the networks can scale to billions of neurons.From Theory to Publication✦ from: We solved the closed-form solution for liquid neural network neuron dynamics in 2022, a problem that had gone unsolved since 1907, and this removed the need for numerical solvers so the networks can scale to billions of neurons.Closed-Form Breakthrough (2022)✦ from: In 2022, Liquid AI solved the liquid neural network differential equations in closed form, a problem that had no known solution since 1907, which removed the need for numerical solvers and allowed scaling from hundreds to billions of neurons.First Closed-Form Solution for Liquid Neural Networks✦ from: In 2022, Liquid AI's team found the first-ever closed-form solution to the neuron-interaction differential equations underlying liquid neural networks, a problem rooted in a 1907 membrane-potential equation that had remained open for over a century.A 115-Year-Old Math Problem, Solved✦ from: Liquid AI found the first closed-form solution for the differential equations governing liquid neural network neuron dynamics in 2022, a problem that had remained mathematically open since 1907.First Closed-Form Solution to Liquid Neural Network Dynamics✦ from: In 2022 we found the first-ever closed-form solution to the liquid neural network differential equations — a problem open since 1907 — which let us scale from hundreds of neurons to billions without numerical solvers.Closed-Form Liquid Neural Networks (2022)✦ from: Liquid AI found the closed-form solution to liquid neural network neuron dynamics in 2022, resolving a mathematical problem for this class of differential equations that had remained open since 1907.A Century-Old Problem, Solved✦ from: In 2022, Liquid AI's team found the first-ever closed-form solution to the neuron-interaction differential equations underlying liquid neural networks, a problem rooted in a 1907 membrane-potential equation that had remained open for over a century.Publication✦ from: Liquid AI found the closed-form solution to liquid neural network neuron dynamics in 2022, resolving a mathematical problem for this class of differential equations that had remained open since 1907.Why It Mattered✦ from: In 2022, Liquid AI solved the liquid neural network differential equations in closed form, a problem that had no known solution since 1907, which removed the need for numerical solvers and allowed scaling from hundreds to billions of neurons.Why the Closed-Form Solution Matters✦ from: Liquid AI found the first closed-form solution for the differential equations governing liquid neural network neuron dynamics in 2022, a problem that had remained mathematically open since 1907.Why It Matters: No More Numerical Solvers✦ from: In 2022 we found the first-ever closed-form solution to the liquid neural network differential equations — a problem open since 1907 — which let us scale from hundreds of neurons to billions without numerical solvers.The Scale-to-Bias Law✦ from: The larger a neural network gets, the less structured and less biased it should be — transformers dominate at trillion-parameter scale, while gated, biased architectures win at smaller, specialized scales.Why Transformers Win at Scale✦ from: Transformers dominate at massive scale precisely because attention is mathematically unstructured, whereas the nested nonlinearities and gating structures liquid networks and other alternatives rely on only help at smaller, more specialized scales.Structured vs. Unstructured Architectures✦ from: Transformers dominate at massive scale precisely because attention is mathematically unstructured, whereas the nested nonlinearities and gating structures liquid networks and other alternatives rely on only help at smaller, more specialized scales.Where Liquid Networks Win✦ from: The larger a neural network gets, the less structured and less biased it should be — transformers dominate at trillion-parameter scale, while gated, biased architectures win at smaller, specialized scales.Scale Determines Optimal Architectural Bias✦ from: The optimal degree of architectural bias (gating, structure, specialized operators) is a function of model scale: larger models perform best with unstructured, unbiased operators like pure attention or matmul, while smaller models benefit from added structural biases.Why Transformers Keep Winning at Scale✦ from: The optimal degree of architectural bias (gating, structure, specialized operators) is a function of model scale: larger models perform best with unstructured, unbiased operators like pure attention or matmul, while smaller models benefit from added structural biases.Architectural Bias Scales Inversely with Model Size✦ from: The optimal amount of architectural bias (gating, nonlinearity, structure) in a neural network is inversely related to model scale: smaller, specialized models benefit from more biased/structured operators, while at very large scale (up to trillions of parameters) unstructured, unbiased architectures like pure attention and matrix multiplication win.Where Each Regime Wins✦ from: The optimal amount of architectural bias (gating, nonlinearity, structure) in a neural network is inversely related to model scale: smaller, specialized models benefit from more biased/structured operators, while at very large scale (up to trillions of parameters) unstructured, unbiased architectures like pure attention and matrix multiplication win.AFMD: Automated Foundation Model Design✦ from: Liquid AI's automated architecture search system (AFMD) removes human bias by running a hardware-in-the-loop evolutionary search that optimizes for memory, latency, and speed while evaluating quality on real downstream tasks rather than proxy metrics like perplexity.AFMD: Automated Foundation Model Design✦ from: Liquid AI built an in-house automated architecture search system (AFMD) that puts target hardware in the loop and evaluates candidate architectures on real downstream application performance across ~100 benchmarks rather than proxy metrics like perplexity, specifically to remove human bias from architecture design.Judged on Real Tasks, Not Proxies✦ from: Liquid AI's automated architecture search system (AFMD) removes human bias by running a hardware-in-the-loop evolutionary search that optimizes for memory, latency, and speed while evaluating quality on real downstream tasks rather than proxy metrics like perplexity.Real Performance, Not Proxy Metrics✦ from: Liquid AI built an in-house automated architecture search system (AFMD) that puts target hardware in the loop and evaluates candidate architectures on real downstream application performance across ~100 benchmarks rather than proxy metrics like perplexity, specifically to remove human bias from architecture design.Judged on Real Tasks, Not Proxies✦ from: Liquid AI's automated architecture search system (AFMD) removes human bias by running a hardware-in-the-loop evolutionary search that optimizes for memory, latency, and speed while evaluating quality on real downstream tasks rather than proxy metrics like perplexity.AFMD: Automated Foundation Model Design✦ from: Liquid AI built an automated, hardware-in-the-loop architecture search system (AFMD) that uses an evolution strategy across dozens of operators to remove human bias from architecture design, optimizing directly for memory, latency, and speed on target hardware.Optimization Criteria✦ from: Liquid AI built an automated, hardware-in-the-loop architecture search system (AFMD) that uses an evolution strategy across dozens of operators to remove human bias from architecture design, optimizing directly for memory, latency, and speed on target hardware.Architectural Bias Is a Function of Scale✦ from: At smaller, specialized model scales, biased/structured components like gating mechanisms improve performance, but as models scale up they benefit from becoming maximally unstructured (pure attention, pure matrix multiplication)—architectural bias is a function of scale, not a universal good.Attention Still Matters, But Gating Elsewhere Counts More✦ from: At smaller, specialized model scales, biased/structured components like gating mechanisms improve performance, but as models scale up they benefit from becoming maximally unstructured (pure attention, pure matrix multiplication)—architectural bias is a function of scale, not a universal good.Mobile Market Rivals AI Data Centers✦ from: The on-device compute market—smartphones and laptops combined—represents roughly a trillion dollars of underutilized 'dark compute,' with the mobile business alone being as large as the AI data center market.A Trillion Dollars of 'Dark Compute'✦ from: The on-device compute market—smartphones and laptops combined—represents roughly a trillion dollars of underutilized 'dark compute,' with the mobile business alone being as large as the AI data center market.Human Brain Beats AI on Intelligence-per-Watt✦ from: Current AI algorithms will not reach the intelligence-per-watt efficiency of the human brain, because human intelligence emerged from a diverse set of mechanisms shaped by biological evolution, not merely from next-token prediction.Current AI Can't Match Brain's Intelligence-per-Watt✦ from: With current AI algorithms, it will not be possible to reach the intelligence-per-watt efficiency of the human brain, because that efficiency is the product of a vastly long biological evolutionary process encoding multiple diverse learning mechanisms (in-context learning, reinforcement learning, mental simulation, Bayesian inference) beyond what next-token prediction alone can produce.AI Won't Match the Brain's Efficiency✦ from: Current AI algorithms will not reach the human brain's intelligence-per-watt efficiency, because the brain's efficiency is the product of an immensely long evolutionary process, not just something replicable through next-token prediction at smaller scale.Beyond Next-Token Prediction✦ from: Current AI algorithms will not reach the human brain's intelligence-per-watt efficiency, because the brain's efficiency is the product of an immensely long evolutionary process, not just something replicable through next-token prediction at smaller scale.Evolution's Multi-Mechanism Advantage✦ from: Current AI algorithms will not reach the intelligence-per-watt efficiency of the human brain, because human intelligence emerged from a diverse set of mechanisms shaped by biological evolution, not merely from next-token prediction.One Mechanism vs. Many Learning Algorithms✦ from: With current AI algorithms, it will not be possible to reach the intelligence-per-watt efficiency of the human brain, because that efficiency is the product of a vastly long biological evolutionary process encoding multiple diverse learning mechanisms (in-context learning, reinforcement learning, mental simulation, Bayesian inference) beyond what next-token prediction alone can produce.
Related episodes