Data◆Audio · 13:20 — 14:15
95% of tokens served on Baseten are from custom models where customers make modifications with their own data — no one just runs vanilla open source weights.
Tuhin reveals that 95% of Baseten's inference tokens come from their dedicated (custom model) business, where customers are making modifications to models with their own data — either for quality or performance. No one is running vanilla open source weights. ✦ AI generated
Tuhin Srivastava · No Priors · 2026-05-01 · original ↗
plays this moment only · 13:20 — 14:15
Elicited by
“Actually, maybe you can just characterize like workload a little bit, like how of tokens being served on base 10, like how many of them are from custom models of some kind versus like vanilla open source today?”
It is all custom. It's basically. 95% plus. And I think that's really cool, to be honest. And look, we have two businesses. We have three businesses. We have three businesses right now. ... I'd say 95% of the tokens today are on the first business. And almost all of them There's probably a, for almost all of them, the customer is making some modifications to the model with their own data specialized for the use case. And I think what's even more important is they might be compiling it in different ways. No one is just running the vanilla open source weights. Like you might be customizing it for quality, but you also might be customizing it for performance.
verbatim transcript · starts at 13:20
- ·95%+ of Baseten tokens come from the dedicated model business
- ·Nearly all customers modify models with their own data
- ·No one runs vanilla open source weights
- ·Customers customize for quality improvements
- ·Customers also customize for performance gains
- ·Models may be compiled differently per use case
Around this claim
Evidence · 2
The application layer will persist because companies that gather unique user signals can encode that value into workflows and post-train specialized models on those signals.Tuhin Srivastava · No Priors · conf 85%99% of the inference market by count is still traditional enterprise — the vast majority of the market hasn't come online yet, and enterprise adoption is well ahead of us.Tuhin Srivastava · No Priors · conf 70%
In practice · 2
Frontier AI models capture the vast majority of economic value even though open-source models may account for the majority of tokens consumed worldwide.Gavin Baker · BG2 Pod · conf 75%Frontier AI models are capturing the vast majority of economic value even though open-source models may account for the majority of raw tokens consumed.Gavin Baker · BG2 Pod · conf 75%