99% of inference volume by count still comes from traditional enterprise workloads that haven't adopted AI yet, meaning the vast majority of the market is still ahead of us.
Tuhin estimates that 99% of inference count today comes from non-AI-traditional enterprise workloads, implying the AI application layer has barely scratched the surface of the total addressable market. ✦ AI generated
Tuhin Srivastava · No Priors · 2026-05-01 · original ↗
plays this moment only · 4:58 — 5:34
“What proportion of the market today do you think is these new application companies versus enterprises just adopting AI? And how do you think that looks in a couple of years?”
The answer is just that it's crazy that the answer is still, I think, if you look by inference count, it'd be ninety-nine percent the fall. I think that kind of represents the scope of the opportunity here. is that the majority of the market hasn't come online and added AI into this. Yeah, most of enterprise adoption is well ahead of us. And I think that's one of the very exciting things about AI. Because there's just so much still to come and people are underestimating that, I think.
verbatim transcript · starts at 4:58
(00:00:06) Hi, listeners. (00:00:07) Today, Elad and I are here with Tuhan Srivastava, the founder and CEO of Baseten, the AI Inference Cloud. (00:00:13) We're here to talk about capacity constraints for AI compute, why inference is the last market, how the workload is changing, the open source and perhaps multi-chip future, and what 30x scale in a year looks like. (00:00:28) Tuhan, welcome back. (00:00:29) Good to see you. (00:00:30) Thanks for having me. (00:00:31) All right, you are in one of the craziest markets, AI in France. (00:00:37) It's very important. (00:00:38) There's a lot going on. (00:00:39) You guys have grown 30x over the last year. (00:00:42) And I think I can say you're expecting to do more than a billion dollars in revenue this year. (00:00:47) What's going on? (00:00:48) Tell us about scale. (00:00:49) Yeah, no, it's been nuts. (00:00:52) I think what's happened over the last, honestly, 24 months, but just kind of keeps getting bigger and bigger, is that (00:00:59) I think everyone is realizing that you can put AI everywhere. (00:01:05) You have all these great options available from closed source, open source models. (00:01:10) The open source models have crossed some sort of chasm in terms of their baseline capability. (00:01:16) And then I think RL techniques and post-raining for specialized models has (00:01:25) become mainstream enough and there's enough examples of it working, the customers realizing they can kind of own their inference more and more. (00:01:35) And what that's meant for us is more, the long tail models coming true, customers in-housing a lot of that intelligence themselves. (00:01:44) And as the application layer just gets bigger and bigger and bigger, and that's growing, when we are just someone indexed on that and we've been around (00:01:54) to be able to collect the demand. (00:01:55) There's an existential question in here that I think everybody is continually asking of, does the independent application layer get to exist at all versus the labs? (00:02:04) Like, how do you? (00:02:06) You have to believe this. (00:02:07) Why do you believe it? (00:02:08) Yeah, look, I think it'd be, it'd be a sad thing if it didn't exist in general. (00:02:11) And I think that's like my, but you know, sadness is fine. (00:02:15) I said all the time. (00:02:17) Yeah, sadness is fine. (00:02:19) But that's not the reason why I think the application layer will exist. (00:02:22) I think the application layer will exist for a number of reasons. (00:02:25) One is because, you know, I think this idea that what is valuable to a company, (00:02:36) is, the user signal that they can gather, that only they can gather. (00:02:42) And to the extent that is encoded in a model, I think a lot of their business will be at risk, but to the extent that it is encoded in workflows, that is where they will be able to develop modes. (00:02:56) So I think a good example of that is (00:02:58) say a company like a bridge where the clinicians edits of the notes and what they do with those notes after the fact and the thing that happens in inside the EMR 3 steps down and that becomes a workflow that only. (00:03:12) Can you explain what a bridge does? (00:03:13) Sorry, a bridge is a ambient scribe that is used by physicians in (00:03:22) almost all hospitals in the US. (00:03:24) I think a large investor, great ship's amazing, great company, great team, great product. (00:03:31) And you know, they've basically, you know, got this very, very deep integration into hospitals, into clinician workflows. (00:03:40) And my argument would be here is that actually, you know, it's very, very hard for a frontier model company to be able to EWA at that because they just don't have access to that user signal. (00:03:51) And what will happen over time is the folks who have access to that user signal can start to post-train models on that reward signal and start to get long horizon agentic models running that. (00:04:02) And I think to the extent that is possible and that signal is differentiated and unique and is somewhat rare to get access to, there will be an application layer. (00:04:16) And I think, you know, it's like support companies is another example of that where, you know, a support (00:04:21) a support task isn't one-shotted. (00:04:24) Usually at a company like Basetime, when a ticket comes in, there's like, what, like 1, 2, 10, 20 actions that get taken. (00:04:30) And that is where, someone can develop a specialized model. (00:04:34) So there's almost 2 versions of this then. (00:04:36) There's new companies like Abridge or Decagon or some of these other things that you mentioned that are doing these new types of applications that are using AI and they sell it to customers. (00:04:44) The other is enterprises building things in-house or building their own models. (00:04:48) What proportion of the market today do you think is (00:04:52) these new application companies versus enterprises just adopting AI? (00:04:56) And how do you think that looks in a couple of years? (00:04:58) Yeah, I think that's a, that's a, I think you asked me the same question two years ago. (00:05:02) I had to be repetitive. (00:05:04) It is crazy. (00:05:06) The answer is just that it's crazy that the answer is still, I think, if you look by inference count, it'd be ninety-nine percent the fall. (00:05:16) I think that kind of represents the scope of the opportunity here. (00:05:20) is that the majority of the market hasn't come online and added AI into this. (00:05:26) Yeah, most of enterprise adoption is well ahead of us. (00:05:28) And I think that's one of the very exciting things about AI. (00:05:30) Yeah. (00:05:30) Because there's just so much still to come and people are underestimating that, I think. (00:05:34) 100%. (00:05:34) But what's cool is that we're seeing the transition happen. (00:05:37) Right before it was like, hey, are they using AI tools? (00:05:41) I don't think that was immediately obvious two years ago. (00:05:43) I think that's obvious now that yes, they are. (00:05:45) they using closed source model APIs? (00:05:48) I think they're starting to get there. (00:05:49) I think once you do that and then you kind of see what is possible, then comes the whole custom model adoption. (00:05:55) I think that is all that is ahead of us today. (00:05:58) So if the majority of your customer base today is, as you described, the former like application companies, AI natives, the fast growing, I mean, some of them are at considerable scale now, like the Abridge, Cursor, Open Evidences of the world. (00:06:13) What, you know, (00:06:15) What do they teach you? (00:06:17) What does that push the company to do? (00:06:19) How do you think about serving them versus evolving for the enterprise? (00:06:22) Yeah, I think firstly, like you just learn a lot by building with the company's greatest scale, doing the most interesting things. (00:06:31) We think of it as two ways. (00:06:34) Like I think there's like the most obvious way, which is just build for the highest scale, you know, most (00:06:42) the customers that will push you the most from technologically and everything kind of will fall into play. (00:06:47) I think the Stripe evolution as a company showed that which like Stripe now serves so many enterprises, but 12 years ago that wasn't the case, but they just built for the frontier and kind of went with them. (00:06:59) And the second way we think about that, this is to just think about building for companies that are serving enterprises. (00:07:05) So yes, we don't serve the enterprises, but our customers serve enterprises. (00:07:08) A bridge serves as OpenEvidence, Decagon, (00:07:13) all these writer, gamma, all these companies serve enterprises in mass. (00:07:19) And what we actually get is like a translation of the requirements from them, which is like, they're like, hey, we need this sort of data retention. (00:07:26) We need this type, this web models need to be deployed. (00:07:30) This is the types of GPUs or the latencies they're okay with. (00:07:33) There's the model requirements from like a transparency perspective that they care about. (00:07:37) And so I think that is actually the more nuanced answer. (00:07:39) It's that (00:07:40) If you listen to what their needs are, we actually get a full translation of what the enterprise will require. (00:07:45) Like I would say that by serving companies like Abridge and Open Evidence, we're probably pretty well suited to go serve the healthcare system, given that they are selling, and latent health, given that they are selling to them. (00:07:56) How much of a shift are you seeing in terms of the types of open source models that are being used? (00:07:59) And so I think we've seen an evolution where two, three years ago, (00:08:04) I think the main thing was kind of Mistral and then a few other things. (00:08:07) And then Meta kind of came along with Llama and then it kind of really shifted in terms of the most performant models or of Chinese origin in different ways. (00:08:14) Do you see that sort of mix reflected in terms of what's being used by our customers? (00:08:17) Yeah, I think customers, at least the customers we are serving, are very, and these are like the fastest growing AI companies in the world that are very forward thinking. (00:08:26) They want to use the best model and they are optimizing. (00:08:29) I think there is (00:08:31) there's a subset of tasks which I think is small today, where people really start to start with cost. (00:08:40) But everyone comes from capability first, because that's really where the economic growth is being unlocked, where the value is being delivered, and then they optimize. (00:08:47) And I think that's like actually being, you know, and so with that in mind, you know, you name, like, you name it, everything from GPT, OSS, all the way to Moonshot, or to Deepseeks, to (00:09:01) Canopy, Orpheus, which is like really good text to speech models. (00:09:08) Customers generally want to use whatever's at the frontier. (00:09:12) And I think the difference has just been, I think we have a lot more visibility into how to run these and how to run these really well. (00:09:21) And secondly, that they're good now. (00:09:22) There have been a number of different concerns raised about the use of Chinese models, in particular security, or is there something embedded in the models or, you know, Trojan horses or other things? (00:09:32) A, do you think there's any real concern there? (00:09:34) And B, people often talk about how there should be like US counterweights to this. (00:09:39) From A geopolitical perspective, do you think that's something that's legitimate or something we should be worried about? (00:09:42) Or how do you think about the sort of origins of these models versus their uses? (00:09:47) Yeah, look, I think these models, firstly, are fantastic. (00:09:51) They're amazing. (00:09:52) We work with these teams. (00:09:53) They're truly awesome. (00:09:55) I'd say, look, I don't (00:09:58) It is hard for me, it's hard for me to see, and I could be wrong, but if I network bound these models that they're not magically, going to be able to cross those network boundaries and so data is that, and I don't, and we, I've never seen any real evidence except from some very early models that I think people picked up on very quickly that there is some agenda or bias built into this. (00:10:26) I do think that (00:10:29) to some extent is I think there is importance to the US that we develop our own models. (00:10:36) I think that would be a massive loss if that there are five companies, five different labs in China that are creating open source models and we're struggling to get one set up. (00:10:47) So it's necessary. (00:10:50) I also think it's inevitable. (00:10:51) And you know, like the deepseek moment a year ago, I remember (00:10:59) someone saying to me now, I thought it was like very well said, which is like, and the world's changed a lot, but they said, hey, we should kind of just forget that this is a Chinese model. (00:11:07) We should just act like this came from meta and build with that in mind. (00:11:14) It's like, you know, I think you're kind of missing the forest from the trees. (00:11:17) Like there's two scenarios, right? (00:11:19) Either America does not ever come up with good open source models. (00:11:22) I think there's probably a fundamental problem there, or we will get there, and we need to be ready. (00:11:28) Yeah, that makes sense. (00:11:29) It's interesting because, like you, I think it's very important for the US to have a strong open source footprint here. (00:11:37) At least for now, it looks like effectively the Chinese government is subsidizing at least a large subset of these models. (00:11:42) And that subsidy or surplus is effectively just being passed on to US enterprises who are adopting these models. (00:11:47) In other words, it's a way for the Chinese government to effectively subsidize US enterprise in an indirect manner. (00:11:52) And I think that's a little bit lost right now. (00:11:55) But it's always interesting to weigh that against some of the other concerns that are raised. (00:11:58) I appreciate your comments on this. (00:12:00) And I think the concern also just there just becomes, it's like, what happened if we aren't able to, if it is fun, I think if you think about the economics here, which is deepseek by most, deep seek's a very good model, and you can argue whether it's at the absolute frontier or not, but let's go back three months and it's there. (00:12:22) And so think about everything, and we're doing a whole other thing three months ago. (00:12:25) And so let's just think about that. (00:12:26) Well, you know, if it (00:12:29) You can run Deepseek, probably 20% of the cost of running anthropic models in production with comparable, better latency, probably better reliability. (00:12:41) If we don't have access to that intelligence in that form, I think it's just a massive loss. (00:12:46) And as a country, we won't be able to innovate as fast because the cost of intelligence going down in control of intelligence, what we have seen just needs more intelligence. (00:12:54) Intelligence being embedded in more places. (00:12:56) Yeah, an important note here that we didn't mention explicitly is that the state-of-the-art models, the ones that are most far ahead on the frontier, are actually still the closed source, anthropic, OpenAI, Google, et cetera. (00:13:07) What has been (00:13:09) Actually, maybe you can just characterize like workload a little bit, like how of tokens being served on base 10, like how many of them are from custom models of some kind versus like vanilla open source today? (00:13:21) It is all custom. (00:13:23) It's basically. (00:13:24) Okay. (00:13:25) So like 95% plus. (00:13:26) 95%. (00:13:26) And I think that's really cool, to be honest. (00:13:29) And look, we have two businesses. (00:13:33) We have three businesses. (00:13:34) We have three businesses right now. (00:13:36) And like we'll help you count. (00:13:37) No, So we have like dedicated inference, which is basically custom model inference. (00:13:42) Your SLA is your SLA. (00:13:43) Then we have shared inference, which is shared inference endpoint, shared SLAs. (00:13:47) And then we have a training business. (00:13:52) I'd say 95% of the tokens today are on the first business. (00:13:58) And almost all of them (00:14:02) There's probably a, for almost all of them, the customer is making some modifications to the model with their own data specialized for the use case. (00:14:10) And I think what's even more important is they might be compiling it in different ways. (00:14:14) No one is just running the vanilla open source weights. (00:14:18) Like you might be customizing it for quality, but you also might be customizing it for performance. (00:14:23) You made an acquisition of a research team a few months ago. (00:14:27) You've mentioned post-training customization. (00:14:31) What was the rationale behind the acquisition? (00:14:33) What is that team doing today? (00:14:34) Yeah, so the rationale around the acquisition was, you know, we are infrastructure and product people. (00:14:42) We are product people and now are really good infrastructure people. (00:14:46) And we didn't have much of A research. (00:14:53) capability ourselves. (00:14:54) And what we saw was the market moving heavily, like that we could accelerate the market itself with post-training resources, either productized or onto even just as resources for that market. (00:15:10) So PaaSD was a company that was a base 10 customer. (00:15:14) So there were post-training models and running them on base 10. (00:15:18) And I think what they realized was (00:15:24) that they would eventually need to become an inference company. (00:15:28) And what we realized was like, hey, we really needed that expertise because it is, it represents a way for us to get closer to the customer earlier and be able to support them all. (00:15:40) And it just made sense as a, like pairing them together. (00:15:43) And just as I said in the opening statement here, which is, as more and more post-train models (00:15:50) have come up. (00:15:51) We've realized that the demand for people to either for software loops to do post-training or for post-training expertise is very high, and we're really, really investing in that. (00:16:07) There are also a bunch of Australians, you know, I like to think that we had a bit of alpha there, but yeah, that's been fantastic. (00:16:15) They're working with (00:16:16) all sorts of customers. (00:16:18) And it's also very interesting when you start, we were doing a lot of research on the performance side and less so on the post-training side. (00:16:29) It's interesting as we've started to do a lot more research on the post-training side, you start to see how linked inference and post-training are. (00:16:36) And like, you know, even when you think about stuff like quantization, (00:16:40) and when you should do that. (00:16:41) And how training, how you train the model affects how you need to quantize for inference and how paired these problems are has become very apparent. (00:16:54) And more and more relating the post-training inference are kind of both sides of the same problem. (00:16:58) So because inference ideally will beget more post-training where inference creates data, you do evals, you can now post-train on that reward function. (00:17:07) that you found with those evals and hopefully just set up the entire look. (00:17:11) Plenty of folks from Ant and OpenAI, Sam, Greg, et cetera, have said in recent months that like inference is super strategic, inference talent is strategic, capacity is strategic. (00:17:23) So between that and post-training, these are very difficult to gather like capabilities. (00:17:32) I imagine that lots of your customers go to you guys for advice on like how to do this. (00:17:37) this progression of moving to custom models? (00:17:39) Like, what do you tell people about the life cycle and when they should invest in that? (00:17:43) Yeah, I think it's, hey, go find, go prove to yourself with the best-in-class model that you have something worth optimizing. (00:17:50) And I think, you know, a lot of, you know, if a customer comes to us, (00:17:57) It was that meme which was like, it was like 2 years ago. (00:18:00) It feels like no GPUs pre-product market fit. (00:18:02) It's like no post-training pre-product market fit is what I know. (00:18:06) It's what I'd say. (00:18:07) So people that you're working with here are very, very at scale first. (00:18:10) Yeah, they have a user signal that they know how to optimize and they've shown that they can, you know, they can serve customer value. (00:18:17) and that value, and that they have something special around that value. (00:18:20) And once you have that value, it's like, okay, now how can I do that better, faster, and cheaper? (00:18:24) With the idea being that, hey, if you need to be very good at customer support, you maybe don't need to be that good at coding, and a specialized model might be a better fit for that problem, and you can do it better, faster, cheaper. (00:18:36) What about the capacity side? (00:18:37) You started with unifying capacity across all the clouds and neo clouds. (00:18:43) How do you think about this when everybody keeps talking about a supply crunch and a multi-year supply crunch? (00:18:48) I think, you know, there's so much narrative around the supply crunch. (00:18:53) And no matter, like as much as we hear about it, I don't think people realize how bad it really is. (00:19:01) Like there is, (00:19:03) there's very, very little slack compute available. (00:19:07) we run pretty large clusters ourselves, and we run them in uncomfortably high utilization. (00:19:14) what I'm saying, we're like mid-90s utilization most of the time. (00:19:19) There is, we have made, we have, we sit in 18 different clouds now. (00:19:27) We have 90 clusters around the world across 18 different clouds, and like, (00:19:31) Initially, we started, we built this technology to be able to kind of create one runtime fabric that spans all these different clouds and try to abstract that away from our customers as a way to think about reliability, latency, failover, all these things that we think are going to be very important for very mission-critical use cases. (00:19:49) That same technology, like just our ability to get compute wherever humanly possible, has been really, really helpful in our ability to get supply. (00:19:59) And what I mean by that is (00:20:01) we can be introduced to a new provider in a different country and have it up and running with the whole base 10 inference stack. (00:20:10) As part of the fabric. (00:20:11) Part of fabric in half a day, half, maybe less. (00:20:18) Even, and that gives us enormous flexibility. (00:20:23) Even for us, it is hard for us to grow. (00:20:26) We have a, we have a, I think it's, yeah, I'll say it. (00:20:30) We have a, (00:20:31) a 4:00 PM standing meeting for the company where we basically like, how do we like, how do we, how do we manage capacity for the demand right now? (00:20:41) I think the second part, which people don't really, the two, the second part that people don't really understand is that there are also a lot of suppliers right now that it's kind of grifty. (00:20:57) You know, like I think, you know, they haven't run (00:21:01) they haven't run data centers before. (00:21:04) they don't understand SLAs, especially for inference. (00:21:08) And so, like, even when there is capacity available, there's a lot of, like, there's probably, we run a lot more than this and we have redundancy, so it's fine. (00:21:18) But if you, know, there's probably like a dozen good, like, clouds, and I'd probably like put like three or four of them in like the gold tier. (00:21:29) And I think that just means that supply, not only are we supply crunch, we're supplier and operationally crunched onto people who can run these data centers as well. (00:21:40) How far ahead can you actually buy capacity right now? (00:21:42) In other words, like, is there any slack in the market if you buy two years ahead or five years? (00:21:48) You mean like actually like contract length or actually like, hey, I want this in January? (00:21:55) Either one, yeah. (00:21:56) I mean, it's more the, I want this in January 28, or at least I have some visibility into my future supply. (00:22:00) Yeah. (00:22:02) You could buy that, but you got to also remember how quickly the market is, how quickly the market is moving. (00:22:09) And like, you know, that gets balanced somewhat off like the fact that the H100 is such a great chip. (00:22:15) And like, and then, (00:22:18) It's crazy. (00:22:18) If it's four years, 4 1/2 years old, the price is going up still. (00:22:21) Maybe as a useful of nine years. (00:22:23) So, that's good. (00:22:25) But at the same time, at the same time, yes, you can do that. (00:22:32) But you're making a lot, like you're making a lot of bets as part of that. (00:22:37) And then in terms of, I think that's the big thing that's changed over the last six months is that the term length that people want has just gone up. (00:22:45) So if you wanted (00:22:49) 1000, 1024 B2 hundreds, which is, from a good cloud. (00:22:58) Right now you're not getting that less than a three to five year contract. (00:23:01) Right now with a probably a 20 to 30% TCV prepay. (00:23:06) So like actually what becomes important when acquiring capacity is you need to have (00:23:12) enough demand to supply it to serve, but then you also need a low cost of capital, which is actually changing the dynamic pretty significantly. (00:23:20) Does that impact how you think about going public as a company? (00:23:23) Because arguably. (00:23:25) Yeah. (00:23:25) I think you'd go sooner. (00:23:26) Yeah, exactly. (00:23:27) Yeah, I think you need, like, I think the, and I think there was demand for that, but I think, you know, the pull, the, it also, you know, one of our (00:23:38) One of the realizations that we had recently, and we're software people, and so we don't think like this all the time, is that our business has very interesting working capital requirements. (00:23:52) And I think even, and that as a result of that, it has very interesting financing.