90% of Decagon's workflow runs on open-source models because smaller, fine-tuned models deliver lower latency and higher performance on the specific tasks of an agent's conversation, while actually outperforming large frontier models on those specialized tasks.
Jesse explains Decagon's migration from frontier models to 90% open-source: as the company scaled and launched voice agents, latency became critical, and frontier labs' small models couldn't be controlled and fine-tuned to their needs. Fine-tuned smaller models prove as good or better on individual agent subtasks. ✦ AI generated
Jesse · a16z Podcast · 2026-07-31 · original ↗
starts at this moment · 2:12
“then thinking machines um Kimk 3 you know some of some very interesting open source models came out sort of right after um and there's this really interesting debate going on around what does it mean to own your destiny when it comes to AI especially in the enterprise”
So today 90% of our workflow is on open source and um you know again the main reason was for latency to really optimize our voice agents and um I think we've just over the last year we've seen tremendous improvement in like how how it sounds how it feels and but still also like keeping the accuracy high um and then the remaining 10% of course we're still using the uh the closed source models and the frontier models for a lot of you know new new projects or new products
verbatim transcript · starts at 2:12
1:55>> sounds good um so I'm gonna talk about our journey first uh just to make it very concrete for people so when we started the any uh the goal was just to get something working, right? So if when you get something when you're the goal is to get something working, of course you're just going to use the frontier models because you want to get something out there and have actually deliver
2:12value. And so we were using opening anthropic at that time they were kind of like one uping each other in terms of how the how the models performed. And then at some point as we got to larger scale and we started working with larger and larger companies and they had you know millions of customers and then we we also have uh we also launched our voice agent right so a big factor became
2:30latency so it wasn't just like can you deliver good responses you have to deliver them really fast and uh the only way to get latency down um but also kind of you make our agent operate the way we want it to is to use smaller models and when [clears throat] you want to go to smaller models unfortunately the the Frontier labs, they do have small models, but you you can't really control
2:51them in the way the way that you want and most small models out of the box are not going to be good enough at the task that we want them to do. So, you have to fine-tune them, you have to change them. And so, that's when we started looking at open source. So, this was about year plus ago. And um it worked really well because if you think about it in the
3:10agent, right? So, in our agent, our agent's job is to have conversations. So, it needs to do a lot of things at once, right? And like one one the first step it might do is like hm what topic is this person talking about or and something else it might do is oh is this person a bad actor that's coming in and trying to mess things up. There's all
3:27these like tasks it has to do. Each individual task doesn't need all of the intelligence of a big model. So you know all the frontier models are obviously very smart but they can do a bunch of different things. Like they can do math they can do coding. Like you just need them to be good at that one task. And so that's why you can use a smaller model
3:44and if you fine-tune it to be really good at that task and can be just as good or better than the big models, right? So that was that was step one for us. You know, about a year ago, we were like, okay, let's start using these open source models. Uh we we we took the small ones and then um you know, that's why we have now a uh a research team and
4:00it's a very expensive team, but it's we we have it because, you know, we need people that are really good at taking these open source models and and tuning them and and so on. So today 90% of our workflow is on open source and um you know again the main reason was for latency to really optimize our voice agents and um I think we've just over the last year we've seen tremendous
- ·Today 90% of Decagon's workflow runs on open-source models
- ·Main reason was latency to optimize voice agents
- ·Remaining 10% still uses closed-source and frontier models
- ·Smaller fine-tuned models deliver lower latency
- ·They outperform large frontier models on specialized agent tasks
- ·Frontier small models can't be controlled or fine-tuned for needs