Token costs will fall roughly 10x over the next three years, and that price collapse will drive roughly 100x growth in AI usage.
Lin Qiao predicts a 10x drop in inference token costs over three years, driven by supply-chain and competitive pressure, which she says will in turn unlock a 100x explosion in AI usage. ✦ AI generated
Lin Qiao · 20VC · 2026-07-20 · original ↗
starts at this moment · 50:15
“When you think forward a year or two how what percent of developer salaries do you think we'll spend? Is it less because these tools will get cheaper or is it more because they'll get better and better?”
in the long term two to three years it should change and that cost will compress uh so overall I can imagine no 10x X uh cost reduction in the next three years and this 10x cost reduction will drive a 100x usage.
verbatim transcript · starts at 50:15
50:15to get much better situation will get much better it I don't think it probably in the next one year or a year a half the situation will not change but in the long term two to three years it should change and that cost will compress uh so overall I can imagine no 10x X uh cost reduction in the next three years and this 10x cost reduction will drive a
50:39100x usage. You said there about kind of uh token efficiency um and how you enable your customers to be much more efficient. With that efficiency, you do charge more. You know, when I when I did the research when it compared to competitors, I got like together's price king. And I don't mean this disparagingly, but like they're cheaper. If you want cheap, you go there and respectfully, if you want better quality
51:05product, you go to you. But it is more expensive. Do you think that's a fair assessment and a fair analogy? >> I think we're probably not comparing apples to apple in the sense that uh again goes back to our business. Majority of our traffic is um is customized model um and we optimize for quality number one. Always quality. quality as in model quality uh towards your applications,
51:35your specific business, your use case and so on. The second is um when we deliver those model in inference is also quality. Um and we care quality so much we do extreme things um for example during training time there's a very hard thing to achieve is called zero KOD. It's a little bit technical. The idea here is >> zero KOD. KLD KLD is a measure of uh of
52:01quality. Um and uh what it means is between the training system and the inference system when model move over uh we have bit equivalence uh so as in the numeric are fully the same we do not lose a bit of accuracy. Uh that's really hard to achieve. But the reason we push that, we deliver that. Um and the reason we push that is because we know um our primary business
- ·Inference token costs projected to fall over next 3 years
- ·Driven by supply-chain gains and competitive pressure
- ·Compression, not just gradual decline, expected
- ·10x cost drop expected to unlock 100x usage growth
- ·Lower prices remove barrier to broader AI adoption
- ·Usage growth outpaces cost decline by 10x