ATRIUMsearch → argument graph
Audio · 2026-05-01 · 43m · 12 moments

Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud

Baseten CEO and co-founder Tuhin Srivastava sits down with Sarah Guo and Elad Gil to discuss the rapid growth of AI inference demand, Baseten’s 30x growth, and why inference is becoming the strategic “last market.” Tuhin Srivastava argues the application layer will persist because companies with unique user signals can encode value into workflows and post-train specialized models, citing examples like Abridge and support workflows. The conversation covers GPU capacity constraints, Baseten’s mult ✦ AI generated

timeline · colored by role

01
Claim

The application layer will persist because companies that gather unique user signals can encode that value into workflows and post-train specialized models on those signals.

Tuhin argues that the independent application layer survives because companies like Abridge gather unique user signals (e.g., clinician workflows) that frontier labs cannot access. Those signals let them build deep integrations and post-train models, creating durable moats.

transcript

Tuhin Srivastava: I think the application layer will exist for a number of reasons. One is because, you know, I think this idea that what is valuable to a company, is, the user signal that they can gather, that only they can gather. And to the extent that is encoded in a model, I think a lot of their business will be at risk, but to the extent that it is encoded in workflows, that is where they will be able to develop modes. So I think a good example of that is say a company like Abridge where the clinicians edits of the notes and what they do with those notes after the fact and the thing that happens in inside the EMR 3 steps down and that becomes a workflow that only... my argument would be here is that actually, you know, it's very, very hard for a frontier model company to be able to compete at that because they just don't have access to that user signal. And what will happen over time is the folks who have access to that user signal can start to post-train models on that reward signal and start to get long horizon agentic models running that. And I think to the extent that is possible and that signal is differentiated and unique and is somewhat rare to get access to, there will be an application layer.

extends · 1rebuts · 1supports · 3

02
Claim

The application layer will persist because companies with unique user signals can encode that value into workflows and post-train specialized models, creating moats that frontier model companies cannot easily cross.

Tuhin argues the application layer survives because unique user signals—encoded into workflows and used for post-training specialized models—create defensible moats. He cites Abridge's deep clinical workflow integration as an example of something a frontier lab cannot easily replicate.

transcript

Tuhin Srivastava: I think the application layer will exist for a number of reasons. One is because, you know, I think this idea that what is valuable to a company, is, the user signal that they can gather, that only they can gather. And to the extent that is encoded in a model, I think a lot of their business will be at risk, but to the extent that it is encoded in workflows, that is where they will be able to develop modes. So I think a good example of that is say a company like a bridge where the clinicians edits of the notes and what they do with those notes after the fact and the thing that happens in inside the EMR 3 steps down and that becomes a workflow that only... And you know, they've basically, you know, got this very, very deep integration into hospitals, into clinician workflows. And my argument would be here is that actually, you know, it's very, very hard for a frontier model company to be able to EWA at that because they just don't have access to that user signal. And what will happen over time is the folks who have access to that user signal can start to post-train models on that reward signal and start to get long horizon agentic models running that. And I think to the extent that is possible and that signal is differentiated and unique and is somewhat rare to get access to, there will be an application layer.

extends · 1

03
Data

99% of inference volume by count still comes from traditional enterprise workloads that haven't adopted AI yet, meaning the vast majority of the market is still ahead of us.

Tuhin estimates that 99% of inference count today comes from non-AI-traditional enterprise workloads, implying the AI application layer has barely scratched the surface of the total addressable market.

transcript

Tuhin Srivastava: The answer is just that it's crazy that the answer is still, I think, if you look by inference count, it'd be ninety-nine percent the fall. I think that kind of represents the scope of the opportunity here. is that the majority of the market hasn't come online and added AI into this. Yeah, most of enterprise adoption is well ahead of us. And I think that's one of the very exciting things about AI. Because there's just so much still to come and people are underestimating that, I think.

provides context · 1supports · 1

04
Data

99% of the inference market by count is still traditional enterprise — the vast majority of the market hasn't come online yet, and enterprise adoption is well ahead of us.

When asked about the split between AI-native application companies and enterprise in-house adoption, Tuhin says that by inference count, 99% is still traditional enterprise — the bulk of the market has not come online yet.

transcript

Tuhin Srivastava: It is crazy. The answer is just that it's crazy that the answer is still, I think, if you look by inference count, it'd be ninety-nine percent the fall. I think that kind of represents the scope of the opportunity here. is that the majority of the market hasn't come online and added AI into this. Yeah, most of enterprise adoption is well ahead of us.

supports · 1

05
Claim

Customers optimizing for capability first, not cost — they use whatever frontier model is best because the economic value is in delivering capability, and they optimize for cost afterward.

Tuhin explains that the fastest-growing AI companies are capability-first: they pick the best model (from GPT to DeepSeek to open source) because that's where the economic value is unlocked, then optimize for cost later.

transcript

Tuhin Srivastava: Yeah, I think customers, at least the customers we are serving, are very, and these are like the fastest growing AI companies in the world that are very forward thinking. They want to use the best model and they are optimizing. I think there is there's a subset of tasks which I think is small today, where people really start to start with cost. But everyone comes from capability first, because that's really where the economic growth is being unlocked, where the value is being delivered, and then they optimize.

extends · 1provides context · 1supports · 1

06
Claim

Customers start from capability first, not cost, because that is where economic value is unlocked, then optimize later.

Tuhin observes that the fastest-growing AI companies optimize for the best model capability first, using everything from GPT to DeepSeek to specialized models like Orpheus, and only later optimize for cost.

transcript

Tuhin Srivastava: I think customers, at least the customers we are serving, are very, and these are like the fastest growing AI companies in the world that are very forward thinking. They want to use the best model and they are optimizing. I think there is there's a subset of tasks which I think is small today, where people really start to start with cost. But everyone comes from capability first, because that's really where the economic growth is being unlocked, where the value is being delivered, and then they optimize.

supports · 1

07
Data

95% of tokens served on Baseten are from custom models where customers make modifications with their own data — no one just runs vanilla open source weights.

Tuhin reveals that 95% of Baseten's inference tokens come from their dedicated (custom model) business, where customers are making modifications to models with their own data — either for quality or performance. No one is running vanilla open source weights.

transcript

Tuhin Srivastava: It is all custom. It's basically. 95% plus. And I think that's really cool, to be honest. And look, we have two businesses. We have three businesses. We have three businesses right now. ... I'd say 95% of the tokens today are on the first business. And almost all of them There's probably a, for almost all of them, the customer is making some modifications to the model with their own data specialized for the use case. And I think what's even more important is they might be compiling it in different ways. No one is just running the vanilla open source weights. Like you might be customizing it for quality, but you also might be customizing it for performance.

explains mechanism · 1gives example · 2supports · 2

08
Data

95%+ of tokens served on Baseten come from custom models where customers modify weights with their own data, and no one is just running vanilla open source weights.

Tuhin reveals that 95%+ of Baseten's inference tokens come from dedicated custom model inference, where customers modify the model with their own data and compile it for performance—no one just runs vanilla open source weights.

transcript

Tuhin Srivastava: It is all custom. It's basically. Okay. So like 95% plus. 95%. And I think that's really cool, to be honest. And look, we have two businesses. We have three businesses. We have three businesses right now. So We have like dedicated inference, which is basically custom model inference. Your SLA is your SLA. Then we have shared inference, which is shared inference endpoint, shared SLAs. And then we have a training business. I'd say 95% of the tokens today are on the first business. And almost all of them There's probably a, for almost all of them, the customer is making some modifications to the model with their own data specialized for the use case. And I think what's even more important is they might be compiling it in different ways. No one is just running the vanilla open source weights. Like you might be customizing it for quality, but you also might be customizing it for performance.

explains mechanism · 1gives example · 2supports · 1

09
Context

The GPU supply crunch is worse than people realize — there is very little slack compute, even for us across 18 clouds and 90 clusters, and we run at mid-90s utilization with a daily standup on capacity management.

Tuhin says the capacity crunch is worse than the narrative suggests. Baseten runs clusters at mid-90% utilization across 18 clouds and 90 clusters globally, and still has a daily 4 PM meeting to manage capacity. He also notes that many newer GPU suppliers lack operational experience running inference workloads, compounding the crunch.

transcript

Tuhin Srivastava: I think, you know, there's so much narrative around the supply crunch. And no matter, like as much as we hear about it, I don't think people realize how bad it really is. Like there is, there's very, very little slack compute available. we run pretty large clusters ourselves, and we run them in uncomfortably high utilization. what I'm saying, we're like mid-90s utilization most of the time. There is, we have made, we have, we sit in 18 different clouds now. We have 90 clusters around the world across 18 different clouds, and like, Initially, we started, we built this technology to be able to kind of create one runtime fabric that spans all these different clouds... That same technology, like just our ability to get compute wherever humanly possible, has been really, really helpful in our ability to get supply. ... I think the second part, which people don't really, the two, the second part that people don't really understand is that there are also a lot of suppliers right now that it's kind of grifty. You know, like I think, you know, they haven't run they haven't run data centers before. they don't understand SLAs, especially for inference. ... there's probably like a dozen good, like, clouds, and I'd probably like put like three or four of them in like the gold tier. And I think that just means that supply, not only are we supply crunch, we're supplier and operationally crunched onto people who can run these data centers as well.

supports · 1

10
Data

The GPU supply crunch is worse than people realize—there is very little slack compute, and even Baseten runs at mid-90s utilization across 90 clusters in 18 clouds, with a standing 4 PM meeting to manage capacity.

Tuhin describes the severity of GPU supply constraints: Baseten runs clusters at mid-90% utilization, has a daily company-wide meeting to manage capacity, and notes that many new suppliers lack data center experience, creating both a supply crunch and an operational crunch.

transcript

Tuhin Srivastava: I think, you know, there's so much narrative around the supply crunch. And no matter, like as much as we hear about it, I don't think people realize how bad it really is. Like there is, there's very, very little slack compute available. we run pretty large clusters ourselves, and we run them in uncomfortably high utilization. what I'm saying, we're like mid-90s utilization most of the time. There is, we have made, we have, we sit in 18 different clouds now. We have 90 clusters around the world across 18 different clouds, and like, Initially, we started, we built this technology to be able to kind of create one runtime fabric that spans all these different clouds... Even for us, it is hard for us to grow. We have a, we have a, I think it's, yeah, I'll say it. We have a, a 4:00 PM standing meeting for the company where we basically like, how do we like, how do we, how do we manage capacity for the demand right now? I think the second part, which people don't really, the two, the second part that people don't really understand is that there are also a lot of suppliers right now that it's kind of grifty. You know, like I think, you know, they haven't run they haven't run data centers before. they don't understand SLAs, especially for inference. And so, like, even when there is capacity available, there's a lot of, like, there's probably, we run a lot more than this and we have redundancy, so it's fine. But if you, know, there's probably like a dozen good, like, clouds, and I'd probably like put like three or four of them in like the gold tier. And I think that just means that supply, not only are we supply crunch, we're supplier and operationally crunched onto people who can run these data centers as well.

11
Data

Term lengths for GPU capacity have jumped dramatically — a 1024 B200 cluster from a good cloud now requires a 3-5 year contract with 20-30% TCV prepay, making cost of capital strategically critical.

Tuhin describes how the market for GPU capacity has tightened: getting a 1024 B200 cluster now requires 3-5 year contracts with 20-30% total contract value prepaid. This shifts the competitive dynamics, favoring companies with lower cost of capital, which is a factor in Baseten's thinking about going public sooner.

transcript

Tuhin Srivastava: So if you wanted 1000, 1024 B2 hundreds, which is, from a good cloud. Right now you're not getting that less than a three to five year contract. Right now with a probably a 20 to 30% TCV prepay. So like actually what becomes important when acquiring capacity is you need to have enough demand to supply it to serve, but then you also need a low cost of capital, which is actually changing the dynamic pretty significantly. ... Yeah, I think you need, like, I think the, and I think there was demand for that, but I think, you know, the pull, the, it also, you know, one of our One of the realizations that we had recently, and we're software people, and so we don't think like this all the time, is that our business has very interesting working capital requirements. And I think even, and that as a result of that, it has very interesting financing.

12
Mechanism

Securing large GPU capacity now requires 3-5 year contracts with 20-30% prepayment, fundamentally changing the working capital dynamics and creating pressure to go public sooner to access lower-cost capital.

Tuhin explains that getting a large block of Blackwell GPUs from a reputable cloud now requires a 3-5 year contract with 20-30% prepayment, which means inference companies need both proven demand and low-cost capital—pushing companies like Baseten toward an IPO sooner.

transcript

Tuhin Srivastava: So if you wanted 1000, 1024 B2 hundreds, which is, from a good cloud. Right now you're not getting that less than a three to five year contract. Right now with a probably a 20 to 30% TCV prepay. So like actually what becomes important when acquiring capacity is you need to have enough demand to supply it to serve, but then you also need a low cost of capital, which is actually changing the dynamic pretty significantly. Does that impact how you think about going public as a company? Because arguably. Yeah. I think you'd go sooner. Yeah, exactly. Yeah, I think you need, like, I think the, and I think there was demand for that, but I think, you know, the pull, the, it also, you know, one of our One of the realizations that we had recently, and we're software people, and so we don't think like this all the time, is that our business has very interesting working capital requirements. And I think even, and that as a result of that, it has very interesting financing.

Highlight slides
The Application Layer Survives via Unique User Signals✦ from: The application layer will persist because companies with unique user signals can encode that value into workflows and post-train specialized models, creating moats that frontier model companies cannot easily cross.Unique user signals create durable moats at the application layer✦ from: The application layer will persist because companies that gather unique user signals can encode that value into workflows and post-train specialized models on those signals.Abridge: workflow integration vs. model access✦ from: The application layer will persist because companies that gather unique user signals can encode that value into workflows and post-train specialized models on those signals.Workflow Integration > Model Access✦ from: The application layer will persist because companies with unique user signals can encode that value into workflows and post-train specialized models, creating moats that frontier model companies cannot easily cross.Long-horizon agentic models reinforce the moat✦ from: The application layer will persist because companies with unique user signals can encode that value into workflows and post-train specialized models, creating moats that frontier model companies cannot easily cross.95% of inference tokens are from custom models✦ from: 95% of tokens served on Baseten are from custom models where customers make modifications with their own data — no one just runs vanilla open source weights.Customization drivers: quality and performance✦ from: 95% of tokens served on Baseten are from custom models where customers make modifications with their own data — no one just runs vanilla open source weights.Custom Models Dominate Baseten Inference✦ from: 95%+ of tokens served on Baseten come from custom models where customers modify weights with their own data, and no one is just running vanilla open source weights.Three Businesses, One Dominant✦ from: 95%+ of tokens served on Baseten come from custom models where customers modify weights with their own data, and no one is just running vanilla open source weights.GPU supply crunch is worse than the narrative suggests✦ from: The GPU supply crunch is worse than people realize—there is very little slack compute, and even Baseten runs at mid-90s utilization across 90 clusters in 18 clouds, with a standing 4 PM meeting to manage capacity.Supply crunch is also an operational crunch✦ from: The GPU supply crunch is worse than people realize—there is very little slack compute, and even Baseten runs at mid-90s utilization across 90 clusters in 18 clouds, with a standing 4 PM meeting to manage capacity.Supply Crunch + Supplier Crunch✦ from: The GPU supply crunch is worse than people realize—there is very little slack compute, and even Baseten runs at mid-90s utilization across 90 clusters in 18 clouds, with a standing 4 PM meeting to manage capacity.
Related episodes