Data◆Audio · 13:21 — 14:23
95%+ of tokens served on Baseten come from custom models where customers modify weights with their own data, and no one is just running vanilla open source weights.
Tuhin reveals that 95%+ of Baseten's inference tokens come from dedicated custom model inference, where customers modify the model with their own data and compile it for performance—no one just runs vanilla open source weights. ✦ AI generated
Tuhin Srivastava · No Priors · 2026-05-01 · original ↗
plays this moment only · 13:21 — 14:23
Elicited by
“Actually, maybe you can just characterize like workload a little bit, like how of tokens being served on base 10, like how many of them are from custom models of some kind versus like vanilla open source today?”
It is all custom. It's basically. Okay. So like 95% plus. 95%. And I think that's really cool, to be honest. And look, we have two businesses. We have three businesses. We have three businesses right now. So We have like dedicated inference, which is basically custom model inference. Your SLA is your SLA. Then we have shared inference, which is shared inference endpoint, shared SLAs. And then we have a training business. I'd say 95% of the tokens today are on the first business. And almost all of them There's probably a, for almost all of them, the customer is making some modifications to the model with their own data specialized for the use case. And I think what's even more important is they might be compiling it in different ways. No one is just running the vanilla open source weights. Like you might be customizing it for quality, but you also might be customizing it for performance.
verbatim transcript · starts at 13:21
- ·95%+ of tokens served come from dedicated custom model inference
- ·Customers modify model weights with their own data
- ·Models are compiled differently for quality or performance
- ·Dedicated inference: custom models, customer-defined SLAs
- ·Shared inference: shared endpoints, shared SLAs
- ·Training business rounds out the portfolio
Around this claim
In practice · 2
Frontier AI models capture the vast majority of economic value even though open-source models may account for the majority of tokens consumed worldwide.Gavin Baker · BG2 Pod · conf 75%Frontier AI models are capturing the vast majority of economic value even though open-source models may account for the majority of raw tokens consumed.Gavin Baker · BG2 Pod · conf 75%
This moment responds to
supports → The application layer will persist because companies that gather unique user signals can encode that value into workflows and post-train specialized models on those signals.Tuhin Srivastava · No Priorsrebuts → Frontier AI models capture the vast majority of economic value even though open-source models may account for the majority of tokens consumed worldwide.Gavin Baker · BG2 Podrebuts → Frontier AI models are capturing the vast majority of economic value even though open-source models may account for the majority of raw tokens consumed.Gavin Baker · BG2 Pod