ATRIUMsearch → argument graph
Audio · 2026-05-21 · 31m · 18 moments

The Story Behind Cerebras’ $63 Billion IPO with Founder and CEO Andrew Feldman

Companies in Silicon Valley from Nvidia to AMD are racing to fuel the AI revolution with postage stamp-sized AI chips. Meanwhile, a chip the size of a dinner plate just fueled a $63 billion IPO for Cerebras. Elad Gil and Sarah Guo sit down with Cerebras founder and CEO Andrew Feldman to discuss the company’s journey to making one of the largest tech go-publics in history. Andrew details the multi-year journey of pioneering wafer-scale AI computing, including surviving a brutal period of being ah ✦ AI generated

timeline · colored by role

01
Claim

Cerebras is the fastest at AI inference, 15-20x faster than GPUs, across the board — big models, small models, US models, Chinese models, trillion-parameter to 1-billion-parameter.

Cerebras's wafer-scale chips deliver 15-20x faster AI inference than GPUs across all model categories — a radical speed advantage that drove explosive demand once models became useful enough for daily work.

transcript

Andrew Feldman: And right now we're the fastest at inference, not by a little, but by a lot, 15, 18, 20x faster than GPUs. [...] Faster across the board. Big models, small models, US models, Chinese models. trillion parameter models, 1 billion parameter models across the board. [...] And once you use something every day in your work, it can't be slow. I mean, how big is the market for slow search? It's 0. How big is the market for dial-up internet? It's 0. That's how big the market for slow inference will be.

explains mechanism · 1

02
Mechanism

To achieve radical improvement over GPUs, you cannot build a derivative architecture — you must be fundamentally different, which is why Cerebras chose wafer-scale chips the size of a dinner plate.

Feldman explains that Cerebras's core architectural bet — wafer-scale chips — was necessary to achieve order-of-magnitude improvements over GPUs, and that critics called it impossible until they proved it worked in 2019.

transcript

Andrew Feldman: I think to be radically better, right? You can't build something that is a similar architecture, right? You're not going to get 15 or 20 times better than the GPU with a minor modification to their architecture. And that's probably true across the board, that if you're going to aspire to a radical improvement, your design has to be different. And from the beginning, we chose wafer scale, which means we build a 46,000 square millimeter chip, a chip the size of a dinner plate, whereas everybody else is building chips the size of postage stamps. They told us we were out of our mind, it would never work. They listed reasons why it was impossible. But in 2019, we proved it was possible.

explains mechanism · 1gives example · 1provides context · 1

03
Mechanism

Wafer-scale engineering was the right architectural bet because radical improvement requires radical architectural difference, not incremental modification of the GPU.

Feldman explains Cerebras's bet on wafer-scale chips: to be 15-20x faster than GPUs, you cannot make minor modifications — you need a fundamentally different architecture, which is why Cerebras built a dinner-plate-sized chip.

transcript

Andrew Feldman: I think to be radically better, right? You can't build something that is a similar architecture, right? You're not going to get 15 or 20 times better than the GPU with a minor modification to their architecture. And that's probably true across the board, that if you're going to aspire to a radical improvement, your design has to be different. And from the beginning, we chose wafer scale, which means we build a 46,000 square millimeter chip, a chip the size of a dinner plate, whereas everybody else is building chips the size of postage stamps. They told us we were out of our mind, it would never work.

explains mechanism · 1gives example · 2

04
Mechanism

Building a radically faster chip required a radically different architecture — wafer-scale integration, a chip the size of a dinner plate — which the industry dismissed as impossible until Cerebras proved it worked in 2019.

Cerebras bet on wafer-scale integration — building a 46,000 sq mm chip the size of a dinner plate — while the rest of the industry built postage-stamp-sized chips. The industry called it impossible, but Cerebras proved them wrong in 2019.

transcript

Andrew Feldman: And from the beginning, we chose wafer scale, which means we build a 46,000 square millimeter chip, a chip the size of a dinner plate, whereas everybody else is building chips the size of postage stamps. They told us we were out of our mind, it would never work. They listed reasons why it was impossible. But in 2019, we proved it was possible. We began delivering it and we improved on it and we improved on it.

explains mechanism · 1

05
Claim

When AI models got smart enough to be useful in daily work, speed became critical — slow inference has no market, just like slow search or dial-up internet.

Feldman argues that once AI models crossed the threshold of being useful enough for daily use, speed became the decisive factor — and a market for slow inference simply does not exist, analogous to dial-up internet or slow search.

transcript

Andrew Feldman: But we were fast when AI was a novelty. And when it's a novelty, nobody cares that you're fast because it's not being used. And so from about 2023 to the beginning of 25, sort of people pointed at AI, but nobody used it every day in their work. And once you use something every day in your work, it can't be slow. I mean, how long will you guys wait for a website to resolve? I'll have no attentions. Right, that's exactly right. That's exactly the way it is. I mean, how big is the market for slow search? It's 0. How big is the market for dial-up internet? It's 0. That's how big the market for slow inference will be. But we had to wait until it was smart enough to be useful. And that happened in 2025.

explains mechanism · 1

06
Claim

Speed is the fundamental value driver in AI inference — slow inference has no market, just like slow search or dial-up internet had no market.

Feldman argues that inference speed is not a nice-to-have but a market requirement: once AI is used daily, it cannot be slow, and the market for slow inference is zero.

transcript

Andrew Feldman: Once you use something every day in your work, it can't be slow. I mean, how long will you guys wait for a website to resolve? … How big is the market for slow search? It's 0. How big is the market for dial-up internet? It's 0. That's how big the market for slow inference will be.

07
Context

New compute workloads create windows of opportunity for new architectures — just as graphics created Nvidia and mobile created ARM, AI's emergence as a new workload demanded a dedicated, non-derivative architecture.

Feldman explains the strategic thesis behind Cerebras: every major new compute workload — graphics, mobile — created a new dominant architecture while incumbents like Intel and AMD got zero share. AI represented the same kind of opportunity, demanding a clean-sheet design, not a derivative of existing architectures.

transcript

Andrew Feldman: But when graphics emerged, you got the discrete GPU and you got Nvidia. And when the mobile compute hit, you got ARM. And it was interesting that not Intel, not AMD, not all sorts of people who you would have thought have been really well positioned to win in that business, they all got no share. And so we knew that this new workload would eat a lot of compute. It would require a new architecture, a dedicated architecture, and it ought to be very different. The architecture could not be a derivative of what's existing. Those were our big bets, and they were 100% contrarian. And they turned out to be dead right.

explains mechanism · 1

08
Anecdote

Cerebras survived a brutal two-year period (2017–2019) where they spent $8M/month trying to build the wafer-scale chip that had never been successfully built in the 70-year history of computing.

Feldman describes the technical near-death experience of Cerebras: spending $8M/month for two years while board meetings every six weeks reported failure, until finally yielding the wafer in summer 2019 — a moment the team watched in stunned silence.

transcript

Andrew Feldman: We had a period between about 2017, middle of 2017 and middle of 2019, where we couldn't build it. We were spending about 8 million a month. You're having board meetings every six weeks saying, I can't build it. No, it's still not working. And right, Oof is right. I mean, that's a huge amount of money and a huge amount of conviction your investors have. And each time we did a failure analysis, we got a little bit better at it. We got a little bit better at it. And then in the summer of 19, we yielded it and it began to work. And the first time we were sitting in a little makeshift office in downtown Los Altos in a building that was not designed for hardware guys. Yeah, we're staring at a computer, which is about as exciting as watching paint dry, and it's working, and we just couldn't speak for half an hour, right? It's like, nobody's been able to do this, and it's working, and we did this.

provides context · 1

09
Anecdote

Cerebras survived a brutal near-death period from mid-2017 to mid-2019 where they couldn't build the wafer-scale chip, burning $8 million a month with repeated failures, until a breakthrough in summer 2019.

From mid-2017 to mid-2019, Cerebras burned $8 million a month trying and failing to manufacture its wafer-scale chip — a problem nobody in the 70-year history of computing had solved, including legendary computer architect Gene Amdahl. The team finally achieved yield in summer 2019, an emotionally overwhelming moment.

transcript

Andrew Feldman: We had a period between about 2017, middle of 2017 and middle of 2019, where we couldn't build it. We were spending about 8 million a month. You're having board meetings every six weeks saying, I can't build it. No, it's still not working. [...] And then in the summer of 19, we yielded it and it began to work. And the first time we were sitting in a little makeshift office in downtown Los Altos in a building that was not designed for hardware guys. Yeah, we're staring at a computer, which is about as exciting as watching paint dry, and it's working, and we just couldn't speak for half an hour, right? It's like, nobody's been able to do this, and it's working, and we did this.

10
Context

After proving the technology worked, Cerebras faced a painful multi-year period where nobody cared about their speed advantage because AI wasn't yet used daily — the market only materialized in 2025 when models became smart enough to be useful.

Even after solving the hardest technical problem in the computer industry, Cerebras found almost no market — selling only a dozen units in the first generation. The market for fast inference didn't exist until 2025, when AI models became smart enough for daily use, triggering explosive demand from companies like OpenAI, Cursor, Cognition, and Lovable.

transcript

Andrew Feldman: We solved it and we solved this sort of the hardest problem in the computer industry and nobody cared. Nobody. It was like, the first Gen. we might have sold a dozen. The second Gen. we probably sold 300 and now we're still going to sell 10s of thousands in the third Gen. We had a two or three-year period where we were ahead of the market. And absolutely nobody cared that we were blisteringly fast. [...] And that happened in 2025. And that's why you got this sort of explosion of demand and companies like Cognition and Cursor and Lovable and just all these others that began ramping extraordinary.

11
Anecdote

Cerebras survived a brutal period of being ahead of the market — building something technically impossible that nobody wanted — by winning supercomputer customers and a sovereign strategic partner that bridged the chasm to mainstream demand.

Feldman describes the multi-year period after solving wafer-scale engineering where Cerebras was blisteringly fast but nobody cared, and how supercomputer labs and a $1B order from sovereign partner G42 provided the bridge to meet the later wave of demand from OpenAI and AWS.

transcript

Andrew Feldman: We solved it and we solved this sort of the hardest problem in the computer industry and nobody cared. Nobody. … I think there's a path that has been laid down by new computer architectures. And often you begin in the supercomputer world because those guys love speed and they don't care if your software is immature. … But then historically, there's this giant chasm because none of them provide the volume to get to mainstream. And we won a sovereign, a G42, and they became a strategic partner and close friends, and they placed a billion dollar order on us. And with that, we were able to sort of transform the company. We're able to change our supply chain. We're able to deploy equipment in big enough clusters that we could battle test at scale.

explains mechanism · 1supports · 1

12
Context

Cerebras bridged the chasm from supercomputer and niche customers to mainstream demand through a billion-dollar sovereign deal with G42, which let them battle-test at scale and transform their supply chain.

Feldman explains the classic hardware path: start in supercomputing where speed matters and software immaturity is tolerated, win oil & gas and pharma customers, then use a sovereign partner (G42) with a billion-dollar order to scale supply chain and battle-test at massive scale — so when OpenAI and AWS came calling, Cerebras was ready.

transcript

Andrew Feldman: I think there's a path that has been laid down by new computer architectures. And often you begin in the supercomputer world because those guys love speed and they don't care if your software is immature. And so we sort of ran the table there. We won at Argonne National Labs and at Lawrence Livermore and at Sandia and in Europe, at European Parallel Computing Center at LRZ. So we ran the table there. And then we won some guys in the oil and gas space and we won some guys in pharma, all of whom have long histories of using extraordinary amounts of compute. But then historically, there's this giant chasm. because none of them provide the volume to get to mainstream. And we won a sovereign, a G42, and they became a strategic partner and close friends, and they placed a billion dollar order on us. And with that, we were able to sort of transform the company. We're able to change our supply chain. We're able to deploy equipment in big enough clusters that we could battle test at scale. One of the challenges in hardware is your QA lab can't be as big as some of the customers you want to deploy to, right? I mean, you can't put $100 million in your QA lab worth of your own gear. And they worked with us and we began training models for them. We began doing inference for them. They've been an extraordinary partner.

extends · 2

13
Mechanism

The bridge from supercomputer niche to mainstream volume came through a $1 billion strategic partnership with G42, which gave Cerebras the scale to battle-test its hardware, transform its supply chain, and ultimately land deals with OpenAI and AWS.

Feldman describes how new computer architectures historically begin in supercomputing — where customers love speed and tolerate immature software — but face a chasm to mainstream volume. G42's billion-dollar order bridged that chasm, enabling Cerebras to scale manufacturing, battle-test at real deployment sizes, and ultimately land the $20B+ OpenAI deal and AWS partnership.

transcript

Andrew Feldman: I think there's a path that has been laid down by new computer architectures. And often you begin in the supercomputer world because those guys love speed and they don't care if your software is immature. [...] But then historically, there's this giant chasm. because none of them provide the volume to get to mainstream. And we won a sovereign, a G42, and they became a strategic partner and close friends, and they placed a billion dollar order on us. And with that, we were able to sort of transform the company. We're able to change our supply chain. We're able to deploy equipment in big enough clusters that we could battle test at scale.

14
Prediction

The hardest part of the software stack — the compiler — took about 10 years to build, and companies going public after a decade must guard against losing their fearless engineering culture.

Feldman recounts how his co-founder predicted the compiler would take 10 years — which turned out to be accurate — and reflects on the cultural danger as Cerebras grows toward thousands of employees: companies stop taking the risks that made them extraordinary, settling for ordinary execution instead.

transcript

Andrew Feldman: When we started the company, Sarah, one of my co-founders, I do remember. I know. We presented to you, one of my co-founders said, Andrew, it's going to take about 10 years to build a compiler. I said, no, that's crazy. That's big company talk. We can do it in five. It takes about 10 years. It takes a long time to build a compiler. It is an extraordinarily difficult piece of software. And now we've got a good software stack. ... I think we have to continue to sort of be fearless. I think one of the malaise of companies as they get to 1000 to 2000, 3000 people is they stop taking the type of risks that they were taking before, right? You move from being a fearless engineering culture to sort of being, what can we get in the timeframe of the next rev? And I think that's extraordinarily damaging. And we take such pride in doing fearless work. We want to hire people who do fearless work. We want to kind of sort of guard that culture that says we would much rather fail in pursuit of the extraordinary than succeed in the ordinary.

15
Data

The best engineers using AI coding agents have gone from being 10x engineers to being 100x engineers, but most people — including the CEO — are still limping along trying to figure out how to make it work for their roles.

Feldman notes that Cerebras's engineering spend on AI tokens went from $0 to $25-30K/month in eight months, and that a subset of engineers with the right mindset have become 100x engineers by governing multiple agents — but the majority of the workforce, himself included, are still struggling to integrate it effectively.

transcript

Andrew Feldman: I think it's not useful for everybody. I think that's truth. I think there are some people who have sort of the perfect mindset for it. … They are running 8 or 10 agents, 7 by 24. … And they've gone from being sort of 10x guys to being 100x guys. I think the rest of us, myself included, we're sort of limping along. We're trying to figure out how we can make it work for our different jobs, for being the CEO, for being the CFO, for being accountants, for being in marketing.

rebuts · 1

16
Claim

The biggest risk for a growing company is losing its fearless engineering culture — the willingness to fail in pursuit of the extraordinary rather than succeed in the ordinary.

Feldman identifies the greatest internal threat as the transition from a fearless engineering culture to a risk-averse quarterly-delivery culture, and commits to guarding against that by hiring for fearlessness and spending significant time on recruiting.

transcript

Andrew Feldman: I think one of the malaise of companies as they get to 1000 to 2000, 3000 people is they stop taking the type of risks that they were taking before. You move from being a fearless engineering culture to sort of being, what can we get in the timeframe of the next rev? And I think that's extraordinarily damaging. … We would much rather fail in pursuit of the extraordinary than succeed in the ordinary. That is a horrible thing to do.

extends · 1

17
Anecdote

Being CEO is extraordinarily lonely, and you need to genuinely love the journey of competing as a David against Goliath — doing this for money is a horrible reason because there are far easier ways to make it.

Feldman offers advice to founders in the long wait for product-market fit: being a leader is lonely, you need the chip-on-your-shoulder defiance of being told it can't be done, and you must love building — because competing against Nvidia as a David is not the easiest path to wealth.

transcript

Andrew Feldman: Being CEO is an extraordinarily lonely thing. And you're building a business, you're building a business. You guys know this, that being a leader is lonely and it's not easy. And people don't like to say that, especially for those of us who like to solve problems, specifically the problems everyone else says can't be solved. You sort of, you gain fire from that chip on your shoulder, right? When they say it can't be solved, you say in your head, you can't solve it. Right. That's exactly right. ... The other thing is you have to love the journey, right? This things we do are too hard if you don't like the building, right? That you do this for the money is a horrible thing. There are way easier ways to make money than trying to create something extraordinary and compete with somebody as strong as Nvidia. That is not the easiest path. You got to love being a David, right? I'm a professional David. This is my fifth startup. I compete against Goliath. That is what I do for a living. And I think to myself that every dollar, every $1,000,000, every billion we sell, if it wasn't for our brains, their muscle would have taken it in a heartbeat. And you got to love that. And if you don't love that, it's a very long road.

18
Context

Going public is an exchange of professional investors for a different class of investors at a lower cost of capital, but for most companies it remains the path to legitimacy and high valuation — only a handful of AI companies can raise public-market-sized rounds privately.

Feldman explains the IPO decision as a trade-off: exchanging sophisticated VCs for retail investors in return for lower cost of capital and legitimacy, noting that only OpenAI, Anthropic, and Databricks have been exceptions to the traditional path of going public.

transcript

Andrew Feldman: First, sort of going public is exchanging some professional investors, venture capitalists who specialize in technology investing for a different class of investors. And in so doing, reducing your cost of capital a little bit. … I think your question is complicated by the fact that there have been, for the first time in history, four or five companies that can raise huge amounts of money without going public. That this was never a thing before OpenAI and Anthropic and maybe Databricks. … I think for the rest of the world, if you want super high valuations, if you want the legitimacy that comes with it, historically, large companies like doing business with other public companies in the US.

Highlight slides
Cerebras Claims 15–20x Inference Speed Advantage Over GPUs✦ from: Cerebras is the fastest at AI inference, 15-20x faster than GPUs, across the board — big models, small models, US models, Chinese models, trillion-parameter to 1-billion-parameter.Speed Is the New Market Requirement✦ from: Cerebras is the fastest at AI inference, 15-20x faster than GPUs, across the board — big models, small models, US models, Chinese models, trillion-parameter to 1-billion-parameter.Wafer-Scale Bet vs. Industry Orthodoxy✦ from: Building a radically faster chip required a radically different architecture — wafer-scale integration, a chip the size of a dinner plate — which the industry dismissed as impossible until Cerebras proved it worked in 2019.The Path from Impossible to Delivery✦ from: Building a radically faster chip required a radically different architecture — wafer-scale integration, a chip the size of a dinner plate — which the industry dismissed as impossible until Cerebras proved it worked in 2019.AI crossed the usefulness threshold in 2025✦ from: When AI models got smart enough to be useful in daily work, speed became critical — slow inference has no market, just like slow search or dial-up internet.Speed is the gating factor for daily-use AI✦ from: When AI models got smart enough to be useful in daily work, speed became critical — slow inference has no market, just like slow search or dial-up internet.Speed Is the Market Requirement for AI Inference✦ from: Speed is the fundamental value driver in AI inference — slow inference has no market, just like slow search or dial-up internet had no market.Three Analogies, One Answer: Zero✦ from: Speed is the fundamental value driver in AI inference — slow inference has no market, just like slow search or dial-up internet had no market.The Wafer-Scale Death March✦ from: Cerebras survived a brutal two-year period (2017–2019) where they spent $8M/month trying to build the wafer-scale chip that had never been successfully built in the 70-year history of computing.Summer 2019: First Yield✦ from: Cerebras survived a brutal two-year period (2017–2019) where they spent $8M/month trying to build the wafer-scale chip that had never been successfully built in the 70-year history of computing.Ahead of the Market, Solving a Problem Nobody Cared About✦ from: Cerebras survived a brutal period of being ahead of the market — building something technically impossible that nobody wanted — by winning supercomputer customers and a sovereign strategic partner that bridged the chasm to mainstream demand.The Sovereign Partner That Bridged the Chasm✦ from: Cerebras survived a brutal period of being ahead of the market — building something technically impossible that nobody wanted — by winning supercomputer customers and a sovereign strategic partner that bridged the chasm to mainstream demand.Bridging the Chasm to Mainstream Demand✦ from: Cerebras survived a brutal period of being ahead of the market — building something technically impossible that nobody wanted — by winning supercomputer customers and a sovereign strategic partner that bridged the chasm to mainstream demand.
Related episodes