Claim◆Audio · 4:00 — 5:00
When AI models got smart enough to be useful in daily work, speed became critical — slow inference has no market, just like slow search or dial-up internet.
Feldman argues that once AI models crossed the threshold of being useful enough for daily use, speed became the decisive factor — and a market for slow inference simply does not exist, analogous to dial-up internet or slow search. ✦ AI generated
Andrew Feldman · No Priors · 2026-05-21 · original ↗
plays this moment only · 4:00 — 5:00
Elicited by
“Is this faster across the board or is this specific use cases?”
But we were fast when AI was a novelty. And when it's a novelty, nobody cares that you're fast because it's not being used. And so from about 2023 to the beginning of 25, sort of people pointed at AI, but nobody used it every day in their work. And once you use something every day in your work, it can't be slow. I mean, how long will you guys wait for a website to resolve? I'll have no attentions. Right, that's exactly right. That's exactly the way it is. I mean, how big is the market for slow search? It's 0. How big is the market for dial-up internet? It's 0. That's how big the market for slow inference will be. But we had to wait until it was smart enough to be useful. And that happened in 2025.
verbatim transcript · starts at 4:00
- ·From 2023 to early 2025, AI was a novelty — rarely used daily
- ·Once models became smart enough for daily work, speed became critical
- ·Slow inference has no market, just as slow search or dial-up have none
- ·Feldman: a useful tool that is slow simply won't be adopted
- ·Users' patience for latency mirrors short website wait times
- ·Fast inference was not rewarded until usefulness was proven
Around this claim
This moment responds to
explains mechanism → The standalone inference business (Fireworks, Together, etc.) is exploding because the open-weight model trend is exploding, and these companies are seeing massive growth with expanding gross margins.Rory · 20VCexplains mechanism → To achieve radical improvement over GPUs, you cannot build a derivative architecture — you must be fundamentally different, which is why Cerebras chose wafer-scale chips the size of a dinner plate.Andrew Feldman · No Priorsprovides context → Cerebras survived a brutal two-year period (2017–2019) where they spent $8M/month trying to build the wafer-scale chip that had never been successfully built in the 70-year history of computing.Andrew Feldman · No Priorsextends → Companies progress through a maturity cycle from relying entirely on per-token closed model APIs, to hyperscaler provisioned-throughput deals, to dedicated inference providers or in-house infrastructure as their AI product scales.Philip Kiely · The TWIML AI Podcast