ATRIUMsearch → argument graph
ClaimVideo · 5:00 — 5:47

The trade-off between smart expensive models and dumb cheap models is false: a fine-tuned smaller model is better at the task, cheaper, and faster than a large state-of-the-art model on that specific task.

Jesse argues the common Twitter debate frames a false trade-off. In practice, fine-tuning smaller, dumber models yields all three wins on the specific task: better performance, cheaper, and faster. ✦ AI generated

Jesse · a16z Podcast · 2026-07-31 · original ↗

starts at this moment · 5:00

Elicited by

I want to push a tiny bit on that point because often times you seen these debates being had on Twitter the trade-off tends to be oh do we want the smartest model that is very expensive or can we dumb it down a little bit and get it cheaper

often times you know when you see these debates being had on Twitter the the trade-off tends to be oh do we want uh you know the smartest model that is very expensive or can we like dumb it down a little bit and get it cheaper. I actually think that is a false trade-off... even if you have a quote dumber model, you can get it to higher performance on that specific task. So when we fine-tune smaller, dumber models, it's that they're just not as general purpose, but on the specific task we want them to do, they actually outperform the large, smart, state-of-the-art models... So we end up getting all three things. It is better at the toss. It is cheaper and it is faster.

verbatim transcript · starts at 5:00

Transcript · around this moment

4:40can kind of evaluate along three dimensions. It's you know cost intelligence and latency and depending on what you need you want to kind of be at the limit of those three and sometimes you can trade off right so in our case we knew that we actually pull back on intelligence because we all had to do was that one task but now we get these latency advantages. I want to push

5:00a tiny bit on that point actually because um often times you know when you see these debates being had on Twitter the the trade-off tends to be oh do we want uh you know the smartest model that is very expensive or can we like dumb it down a little bit and get it cheaper. I actually think that is a false trade-off, >> right? Because what we've seen in

5:22practice is even if you have a quote dumber model, you can get it, and we've seen this in practice, you can get it to higher performance on that specific task. So when we fine-tune smaller, dumber models, it's that they're just not as general purpose, but on the specific task we want them to do, they actually outperform the large, smart, state-of-the-art models, >> right? So we end up getting all three

5:47things. It is better at the toss. It is cheaper and it is faster. >> And so do you feel like today like you need the most Frontier models for really anything at Decagon because your performance is very good already. So >> we we do and we often need them we often end up needing them for auxiliary tasks, right? Where when you have uh when you sort of have auxiliary models to our

6:09sort of primary conversational flow, right? you have an agent and it's helping a customer with their rebooking or it's helping them with a process in healthcare then these are like well-defined pots. So we have smart bos models to do that but we've for instance recently launched do autopilot right which is our agent that improves the core conversational agent. Now, for something like autopilot, it is doing a

Around this claim