Diffusion language models scale better than autoregressive models at inference time — they're cheaper to serve, faster, and yield more tokens per GPU, which lowers the price per token.
Stefano Ermon opens by arguing that diffusion LLMs beat autoregressive models on inference-time economics — lower price per token and lower watts per token — which is why Inception bet on the approach now that production economics matter most.
transcript
Stefano Ermon: what we're seeing with diffusion language models is that they scale better than auto-regressive models at inference time. They're cheaper to serve, they're faster, you get more tokens per GPU, which means that the price is actually lower.
gives example · 1supports · 1