ATRIUMsearch → argument graph
MechanismVideo · 4:07 — 5:34

OpenRouter did not expect that a separate ecosystem of specialized inference providers would emerge to host open weight models, rather than hyperscalers dominating that layer.

OpenRouter's founding thesis assumed hyperscalers like Google, Amazon, and Azure would monopolize open weight model hosting, but specialized inference providers like Fireworks and Together proved far faster and more capable at serving these models. ✦ AI generated

Alex Atallah · 20VC · 2026-08-10 · original ↗

starts at this moment · 4:07

Elicited by

When you go back to the founding thesis of the company, what has happened in the ecosystem in the model landscape that you did not expect to happen?

one thing that we did not expect was that a an ecosystem of companies would emerge to host and serve the open weight models. Um like early on, it wasn't clear that that that market wasn't going to be a monopoly where like just, you know, the three hyperscalers serve all the open weight models and uh and start and startups don't, you know, they're they're really far behind. In reality, like you know, how often do you hear people running, you know, GLM on a hyperscaler? Never. Like they're using the the inference providers like Fireworks and Together.

verbatim transcript · starts at 4:07

Transcript · around this moment

3:48I just like OpenSea just kind of like drilled that into me in a way where [clears throat] I could like take it productively to open router. >> Can I ask you, when you go back to the founding thesis of the company, what has happened in the ecosystem in the model landscape that you did not expect to happen? >> Um okay, well, one thing that we did not

4:07expect was that a an ecosystem of companies would emerge to host and serve the open weight models. Um like early on, it wasn't clear that that that market wasn't going to be a monopoly where like just, you know, the three hyperscalers serve all the open weight models and uh and start and startups don't, you know, they're they're really far behind. In reality, like you know, how often do you hear people

4:39running, you know, GLM on a hyperscaler? Never. Like they're using the the inference providers like Fireworks and Together. And um there's like go big list that we that we see doing the best job of hosting all of the the open weight models. And um in the early days, we um we had I think we called it provider one and and provider fallback. We didn't like show which providers were actually doing

5:11the hosting cuz I I mean we weren't really a marketplace. We were kind of a an like we were an exploration tool for like finding and discovering new LLMs. And we wanted to get we wanted to build like a marketplace of model labs, but like the inference provider layer, we weren't sure it would actually be a marketplace. And it turned out that those companies were doing a way better

5:34job than the hyperscalers, were way faster to host the models and figure out these edge cases to hosting them. And um and uptime was just going to be a a constant problem. It wasn't going to like magically get solved by the supply side of the market. >> A lot of people suggest that that inference provider layer is a commoditizable element or layer that will be removed or see margin reduction

5:55competed out over time. What would you say to that theory? >> Right now we're in a massively supply constrained market where um but and it's likely going to be supply constrained for a while where [snorts] all the inference providers are are short, pretty much constantly short. Uh and and you're like, "Okay, so GPUs are are really really beneficial." And like, "Why doesn't Google or Amazon or Azure

Around this claim