ATRIUMsearch → argument graph
ClaimVideo · 0:00 — 1:30

Diffusion language models scale better than autoregressive models at inference time — they're cheaper to serve, faster, and yield more tokens per GPU, which lowers the price per token.

Stefano Ermon opens by arguing that diffusion LLMs beat autoregressive models on inference-time economics — lower price per token and lower watts per token — which is why Inception bet on the approach now that production economics matter most. ✦ AI generated

Stefano Ermon · The TWIML AI Podcast · 2026-03-26 · original ↗

starts at this moment · 0:00

what we're seeing with diffusion language models is that they scale better than auto-regressive models at inference time. They're cheaper to serve, they're faster, you get more tokens per GPU, which means that the price is actually lower.

verbatim transcript · starts at 0:00

Transcript · around this moment

0:00If you need to scale up these models and they are actually getting into production, the price per token or the watts needed per token becomes the key metric that you care about. And so, what we're seeing with diffusion language models is that they scale better than auto-regressive models at inference time. They're cheaper to serve, they're faster, you get more tokens per GPU, which means that the price is actually

0:22lower. And so, that's why we we felt like, yeah, [music] this is the time to to do it. And in fact, that's what we're seeing. >> [music] >> All right, everyone. Welcome to another episode of the Twilio AI podcast. I am your host, Sam Charrington. Today, I'm joined by Stefano Erman. Stefano is associate professor at Stanford University and the CEO of Inception. Before we get going, be sure to take a

0:58moment to hit that subscribe button wherever you're listening to today's show. Stefano, welcome back to the podcast. It has been a while. Yeah, thank you for hosting me again. Yeah, it's been a very long time since we last [laughter] chatted. Yeah, I think about eight years or so. Uh certainly lots has changed. Uh and we'll get into some of that, in particular, what you've been doing with

1:21diffusion models. But to get us started, why don't you tell us a little bit about what you've been up to with for the last eight years, maybe? Yeah, so I've been uh working still in the same space. So, I've been working in generative models, I guess, my my whole career, my whole life. Um now, what has changed is that the field really took off. I guess now it's called generative

1:43AI and everybody is paying attention to it and, you know, it's it's um become the thing that everybody is looking at and everybody's trying to, you know, get into. So, yeah, it's been it's been exciting to see the growth of the of the field and the capabilities of these models. When I started back in uh you know, 2014 or so, you know, we were barely able to model MNIST images

Around this claim