Fine-tuned smaller open-source models outperform large frontier models on specific tasks while being cheaper and faster.
Decagon found that the smart/costly trade-off is false—fine-tuned smaller models excel at specific tasks while gaining latency and cost advantages. ✦ AI generated
Jesse · a16z Podcast · 2026-07-31 · original ↗
starts at this moment · 4:58
“I actually think that is a false trade-off, right? Because what we've seen in practice...”
I actually think that is a false trade-off, right? Because what we've seen in practice is even if you have a quote dumber model, you can get it, and we've seen this in practice, you can get it to higher performance on that specific task. So when we fine-tune smaller, dumber models, it's that they're just not as general purpose, but on the specific task we want them to do, they actually outperform the large, smart, state-of-the-art models, right? So we end up getting all three things. It is better at the toss. It is cheaper and it is faster.
verbatim transcript · starts at 4:58
4:58these latency advantages. I want to push a tiny bit on that point actually because um often times you know when you see these debates being had on Twitter the the trade-off tends to be oh do we want uh you know the smartest model that is very expensive or can we like dumb it down a little bit and get it cheaper. I actually think that is a false
5:19trade-off, >> right? Because what we've seen in practice is even if you have a quote dumber model, you can get it, and we've seen this in practice, you can get it to higher performance on that specific task. So when we fine-tune smaller, dumber models, it's that they're just not as general purpose, but on the specific task we want them to do, they actually outperform the large, smart,
5:44state-of-the-art models, >> right? So we end up getting all three things. It is better at the toss. It is cheaper and it is faster. >> And so do you feel like today like you need the most Frontier models for really anything at Decagon because your performance is very good already. So >> we we do and we often need them we often end up needing them for auxiliary tasks,
6:04right? Where when you have uh when you sort of have auxiliary models to our sort of primary conversational flow, right? you have an agent and it's helping a customer with their rebooking or it's helping them with a process in healthcare then these are like well-defined pots. So we have smart bos models to do that but we've for instance recently launched do autopilot right which is our agent that improves the
- ·Fine-tuned smaller models outperform frontier models on specific tasks
- ·Smaller models are not as general-purpose, but excel at narrow tasks
- ·You get better performance, lower cost, and lower latency simultaneously
- ·Better at the specific task than large state-of-the-art models
- ·Cheaper to run than large frontier models
- ·Faster inference than large frontier models