MechanismAudio · 22:42 — 24:45
To be 20x better than Nvidia, you can't build a GPU — you need a fundamentally different architecture. Cerebras solved the memory-to-compute bottleneck by building a chip the size of a dinner plate with memory right next to compute, making it 15-18x faster than a GPU.
Andrew Feldman explains that Cerebras bet on dedicated silicon that couldn't look like a GPU. By building a wafer-scale chip with memory next to compute, they achieved 15-18x speedup over GPUs for OpenAI, which he argues is critical because users won't wait for slow AI. ✦ AI generated
Andrew Feldman · All-In Podcast · 2026-06-06 · original ↗
plays this moment only · 22:42 — 24:45
We made two bets. The first was dedicated silicon would be the answer. And the second was it couldn't look like a GPU. And our view as computer architects is if you want to be 20 times better than somebody, your architecture can't look like them. It can't. They have enjoyed and eaten all the low-hanging fruit. So if you build a GPU, the odds that you're better than Nvidia and our view are approximately 0. That led us to a fundamentally different architecture. The hard part here, the hard part is moving data from memory to compute. This is the fundamental problem in AI. And we solved it with a way that very few others had even attempted, which was to build a very big chip and to put memory right next to compute. And by building a big chip, a chip the size of a dinner plate, whereas most chips are the size of a postage stamp, we could use a different type of memory. And by using a different type of memory, a memory that was vastly faster, we opened up all sorts of opportunity. So when OpenAI uses us, we're 15 or 18 times faster than a GPU. That means your answers are delivered more quickly. It means your engagement with the AI is more enjoyable. It means you can use the AI to solve harder problems and not wait. And the way to think about this is sort of to ask yourself the counterfactual question. How big is the market for slow search today? Right? It's 0. How big is the market for dial-up? It's 0. How long do you wait for a website to resolve before you click away? Three seconds, 5 seconds? You will not wait for AI. We have to deliver it to you in real time.
verbatim transcript · starts at 22:42
Around this claim
This moment responds to
supports → SambaNova's dataflow architecture pushes memory bandwidth utilization to 70-80% of peak by orchestrating data movement in hardware, compared to the roughly 10-20% GPUs typically achieve because they synchronize data movement in software.Kunle Olukotun · The Cognitive Revolutionsupports → Standard GPUs typically use only 10-20% of their available memory-bandwidth and communication capability during inference, whereas SambaNova's dataflow architecture is designed to push effective utilization to 70-80% of peak.Kunle Olukotun · The Cognitive Revolutionexplains mechanism → Cerebras is the fastest at AI inference, 15-20x faster than GPUs, across the board — big models, small models, US models, Chinese models, trillion-parameter to 1-billion-parameter.Andrew Feldman · No Priorsextends → SambaNova's dataflow architecture pushes memory bandwidth utilization to 70-80% of peak by orchestrating data movement in hardware, compared to the roughly 10-20% GPUs typically achieve because they synchronize data movement in software.Kunle Olukotun · The Cognitive Revolutionextends → Standard GPUs typically use only 10-20% of their available memory-bandwidth and communication capability during inference, whereas SambaNova's dataflow architecture is designed to push effective utilization to 70-80% of peak.Kunle Olukotun · The Cognitive Revolutionexplains mechanism → Building a radically faster chip required a radically different architecture — wafer-scale integration, a chip the size of a dinner plate — which the industry dismissed as impossible until Cerebras proved it worked in 2019.Andrew Feldman · No Priors