ATRIUMsearch → argument graph
DataAudio · 13:21 — 14:23

95%+ of tokens served on Baseten come from custom models where customers modify weights with their own data, and no one is just running vanilla open source weights.

Tuhin reveals that 95%+ of Baseten's inference tokens come from dedicated custom model inference, where customers modify the model with their own data and compile it for performance—no one just runs vanilla open source weights. ✦ AI generated

Tuhin Srivastava · No Priors · 2026-05-01 · original ↗

plays this moment only · 13:21 — 14:23

Elicited by

Actually, maybe you can just characterize like workload a little bit, like how of tokens being served on base 10, like how many of them are from custom models of some kind versus like vanilla open source today?

It is all custom. It's basically. Okay. So like 95% plus. 95%. And I think that's really cool, to be honest. And look, we have two businesses. We have three businesses. We have three businesses right now. So We have like dedicated inference, which is basically custom model inference. Your SLA is your SLA. Then we have shared inference, which is shared inference endpoint, shared SLAs. And then we have a training business. I'd say 95% of the tokens today are on the first business. And almost all of them There's probably a, for almost all of them, the customer is making some modifications to the model with their own data specialized for the use case. And I think what's even more important is they might be compiling it in different ways. No one is just running the vanilla open source weights. Like you might be customizing it for quality, but you also might be customizing it for performance.

verbatim transcript · starts at 13:21

Around this claim