Even as nominal model context windows have grown from about 10,000 tokens to 1 million tokens, the effective context window—where models actually perform reliably—has been stuck around 80,000-120,000 tokens for roughly two years and isn't likely to change soon.
Sidhant points out that despite headline context windows expanding to a million tokens, models' effective working context has plateaued at roughly 80-120k tokens for about two years, a limitation he doesn't expect to change soon. ✦ AI generated
Sidhant Pardeshi · The TWIML AI Podcast · 2026-03-10 · original ↗
starts at this moment · 24:39
“So, multi-agentic, where does that come in?”
the effective you you've gone from 10,000 tokens to 1 million tokens and you've gone from, uh, you know, maybe 10,000 to 200k tokens and then 1 million, but we've been stuck at 80k to 100k tokens, um, 80k to 120, I would say, with the latest models, since 2 years.
verbatim transcript · starts at 24:39
24:39you know, maybe 10,000 to 200k tokens and then 1 million, but we've been stuck at 80k to 100k tokens, um, 80k to 120, I would say, with the latest models, since 2 years. So, even though you're getting a new model every 3 months, the effective context window is not changing and it's taken a while for us to go from 10k to 200k to 1 million, right? Because
25:00you have uh physics constraints in these. You have um you know, the amount of compute capacity, you have power. Um you have, you know, how much you can scale for all these model providers. So, um they're they're always trying to find out, you know, better solutions for that. Uh but that's not getting solved in the next 3 months, 6 months, or even I would say in my opinion that's that's
25:23not changing drastically even in the next 3 years. Um so, these are very important considerations. So, what happens if you have um multi-agent capabilities, the ability to recruit multiple agents, and we've seen two techniques uh that have been applied. One is the concept of having sub-agents. So, you have one um you know, orchestrator or leader model that is going to recruit multiple sub-agents. I've seen this used in Cloud
25:49Code, for example. And uh then you can do searches in parallel, right? If you're finding four different things, just run four agents. You know, see which one comes back like like throw four darts, see which one sticks. Um you can do that kind of stuff. Um uh and then or you can do like you can parallelize tasks, right? Give one to a front-end, give one to a back-end agent,
26:07and uh get more work done. Um so, you can do that kind of stuff. Um the advantage you have is of course speed, right? Of course maybe um the effect of intelligence because you're doing multiple things in parallel. And you also have um a a significant, I would say, but not sufficient, a significant improvement in the amount of context you're using with the hydration because it's no longer having to make
26:28all these searches and traverse the code, right? Uh it's getting the result from different agents. Um so, that's more effective than this this guy just having to do it himself. But then you still have a bottleneck, and the bottleneck is this leader agent, right? Uh because everyone's going to report back, so you can't like you can't run hundreds of agents because then you'll run you're you're going to you go