ATRIUMsearch → argument graph
Video · 2026-08-10 · 1h 8m · 18 moments

OpenRouter CEO: Why Chinese Open Models Are Beating the US | Why Enterprises Fear OpenAI & Anthropic

✦ AI generated

timeline · colored by role

01
Mechanism

OpenRouter did not expect that a separate ecosystem of specialized inference providers would emerge to host open weight models, rather than hyperscalers dominating that layer.

OpenRouter's founding thesis assumed hyperscalers like Google, Amazon, and Azure would monopolize open weight model hosting, but specialized inference providers like Fireworks and Together proved far faster and more capable at serving these models.

transcript

Alex Atallah: one thing that we did not expect was that a an ecosystem of companies would emerge to host and serve the open weight models. Um like early on, it wasn't clear that that that market wasn't going to be a monopoly where like just, you know, the three hyperscalers serve all the open weight models and uh and start and startups don't, you know, they're they're really far behind. In reality, like you know, how often do you hear people running, you know, GLM on a hyperscaler? Never. Like they're using the the inference providers like Fireworks and Together.

explains mechanism · 1gives example · 1provides context · 2

02
Mechanism

Nvidia actively prevents hyperscaler concentration in GPU allocation, which sustains the independent inference provider ecosystem and benefits both Nvidia and end users.

The inference provider layer resists commoditization partly because Nvidia's top priority is avoiding customer concentration — they want many separate GPU allocations across the market, which keeps specialized providers alive and drives innovation in model serving.

transcript

Alex Atallah: Why doesn't Google or Amazon or Azure run around and like buy up all the GPUs and take all these inference providers out of business? Well, the people making the GPUs don't want that. Like one of Nvidia's top priorities is not having customer concentration. They want lots of customers to all have like separate like allocations of GPUs. Um they want the the heterogeneity of the market. They want like competition on the compute layer. Um and I and this is good for the ecosystem. Like like users also want this. This like this is it's good for Nvidia and it's good for end users as well. It like allows these inference providers to kind of like come up with new innovations on like how to serve the models better.

03
Claim

A multi-model future is inevitable, and companies that specialize in one proprietary model will still need to use other models created by the ecosystem to stay competitive.

Alex Atallah argues that no single model will win the entire AI market, and companies are incentivized to use a diverse ecosystem of models to improve productivity, reduce costs, and stay on the state-of-the-art frontier.

transcript

Alex Atallah: I think our mission from the very beginning has been to increase neurodiversity in AI for the whole ecosystem. And we really believe that like a multi-model future is inevitable. And when you start, let's say like let's say there's one model that like, you know, hypothetically, let's say you're right. Let's say there's one model that fulfills all of your desires, um either within your company or like as a consumer. Um and every you know more and more people start using that model. And then um someone decides, you know what? I'm going to like create a neurodivergent model. I'm going to create a model that's like a little bit different that like talks a little differently, that has ideas that the first model like could never have come up with cuz it's like completely different data um that's being used to train it. Then it kind of creates inevitable demand to use both models. Like creativity is not a um it's not verifiable. There's there's like you can't really put an easy number on on creative ideas. And when you use two models together, um you're more likely to get creative ideas than if you just use one.

04
Claim

A multi-model future is inevitable because using multiple models together produces more creative ideas than any single model, and companies are always incentivized to experiment with the ecosystem's new models to improve productivity and reduce costs.

Alex Atallah argues that consolidation on one model makes no sense — creativity is non-verifiable and using models trained on different data yields more novel ideas. Even when companies build proprietary models, the ecosystem constantly generates new data and models that incentivize cross-model experimentation.

transcript

Alex Atallah: I think our mission from the very beginning has been to increase neurodiversity in AI for the whole ecosystem. And we really believe that like a multi-model future is inevitable. And when you start, let's say like let's say there's one model that like, you know, hypothetically, let's say you're right. Let's say there's one model that fulfills all of your desires, um either within your company or like as a consumer. Um and every you know more and more people start using that model. And then um someone decides, you know what? I'm going to like create a neurodivergent model. I'm going to create a model that's like a little bit different that like talks a little differently, that has ideas that the first model like could never have come up with cuz it's like completely different data um that's being used to train it. Then it kind of creates inevitable demand to use both models. Like creativity is not a um it's not verifiable. There's there's like you can't really put an easy number on on creative ideas. And when you use two models together, um you're more likely to get creative ideas than if you just use one. It's it's just a fact if if that other model was trained in different way on a different data set, like or has like made a big update. So, um consolidation on one just seems like it just doesn't make any sense to me.

explains mechanism · 1rebuts · 1supports · 1

05
Claim

Companies building routers because it's fashionable puts them months behind fully focused players like OpenRouter, reducing user leverage and access to the full model ecosystem.

Alex Atallah argues that many companies are building LLM routers or gateways as a trend, but their lack of focus puts them far behind dedicated companies like OpenRouter. This limits users' leverage by not providing access to the full market of models and customizations.

transcript

Alex Atallah: I think a lot Yeah, a lot of companies are making routers because it's fashionable. Um I think they're you know, they're seeing growth happen here and uh or they're they're making gateways at least. Um I First, I think there's two There's two issues with that. First, it immediately puts you in the mindset of copying instead of like, you know, winning something... Um and and immediately kind of like puts that gateway like many many months behind the companies that are fully focused on it. Like I am 100% focused on building the best router and gateway and and LLM marketplace and it shows in our product... Um the other problem is that it uh it reduces the leverage of all of your users. So, um like I really deeply believe in giving users and developers more leverage... you're you're kind of like being cut out. You're cutting out all your employees at your company of things that they they need.

06
Data

Token price reductions follow a Jevons paradox pattern, where a 10x price drop leads to a >10x increase in usage, as demonstrated by OpenAI's GPT 5.6 Luna model.

Alex Atallah uses the example of OpenAI's GPT 5.6 Luna model to illustrate the Jevons paradox in AI: a 10x price reduction led to a 13x surge in usage, meaning lower prices expand the total market rather than shrink the revenue pie.

transcript

Alex Atallah: Well, so a lot of people talk about the Jevons paradox that, you know, when when uh prices go down by 10x, the usage increases by more than 10x... For example, uh GPT 5.6 Luna on OpenRouter. Um OpenAI cut prices by 5x and then in coordination with us by another 2x. So in total, prices the price of Luna has dropped 10x on OpenRouter over the last 2 weeks. And guess how much usage has grown? 13x. So it's a close to perfect Jevons paradox story where you drop prices 10x and usage grows by more than 10x. Um just a bit more. And uh and the also the the usage is pretty stable.

07
Data

Token price reductions follow the Jevons paradox: when prices drop 10x, usage increases by more than 10x, so overall platform revenue grows rather than shrinks.

OpenRouter's real data shows that when OpenAI cut the price of GPT 5.6 Luna by 10x, usage grew 13x—confirming the Jevons paradox and suggesting that token price declines are net positive for inference platforms.

transcript

Alex Atallah: A lot of people talk about the Jevons paradox that, you know, when when uh prices go down by 10x, the usage increases by more than 10x. Um but like no one has really done a great job modeling it. There We do have a lot of spot stories that confirm it. For example, uh GPT 5.6 Luna on OpenRouter. Um OpenAI cut prices by 5x and then in coordination with us by another 2x. So in total, prices the price of Luna has dropped 10x on OpenRouter over the last 2 weeks. And guess how much usage has grown? 13x. So it's a close to perfect Jevons paradox story where you drop prices 10x and usage grows by more than 10x. Um just a bit more. And uh and the also the the usage is pretty stable. Like it like grew, you know, it like flattened out but at at 13x and then you know, it's been kind of like growing at the same rate that it was growing before it hit the the 13x multiple. Um so that's pretty interesting and it's a pretty like uh low variable like there are there are there are few other confounding variables in the story

08
Claim

America is very behind China in open weight model development.

Alex Atallah states bluntly that America is 'very, very behind' China in open weight model development, though he notes some US labs like Poolside and Thinking Machines are beginning to catch up.

transcript

Alex Atallah: We should. We're behind. I America is very, very behind still. Um I think things are picking up. I and I think I think um you know, we we have Poolside, we have Thinking Machines, we have RC.

rebuts · 1

09
Claim

America is very behind China in open-weight AI models, and the US needs better pathways to get compute to emerging labs to remain competitive.

Alex Atallah acknowledges that the US is significantly behind China on open-weight models, pointing to DeepSeek, Kimi, and GLM as evidence, and suggests that distilling Chinese models and improving compute access for neo labs are key strategies to catch up.

transcript

Alex Atallah: We should. We're behind. I America is very, very behind still. Um I think things are picking up. I and I think I think um you know, we we have Poolside, we have Thinking Machines, we have RC. ... The other thing I would try to figure out is is the compute question. Like compute is just a huge advantage that I think we still have relative to China and these neo labs need a shot. And there like there needs to be an easier way to like get compute to the right talent.

10
Fact

The US is significantly behind China in the development of competitive open-weight AI models, though new American 'neo-labs' are beginning to emerge.

Alex Atallah asserts that America is 'very, very behind' China in the open-weight model space, highlighting GLM 5.2 as a major step. He acknowledges emerging US labs like Poolside and Thinking Machines but notes the significant gap remains.

transcript

Alex Atallah: We should. We're behind. I America is very, very behind still. Um I think things are picking up. I and I think I think um you know, we we have Poolside, we have Thinking Machines, we have RC... GLM 5.2 was a really big big step for open weight models. Kimi was kind of like moonshot getting up to that step. That's a little bit how I see it.

11
Claim

Enterprises are more nervous about frontier model providers like OpenAI and Anthropic than about Chinese models, due to confusion around data policies and the inability to run frontier models on their own infrastructure.

Alex Atallah observes a counterintuitive dynamic: companies are more nervous about US frontier labs than Chinese models, primarily because data policies around where prompts are stored and reviewed create uncertainty, and frontier models cannot be run on a company's own infrastructure.

transcript

Alex Atallah: I think they're they're more nervous about frontier models usually. Part because there's just like a a much there's much more confusion around the data policy about what's like actually happening to the prompts they're sending and um where they're being stored and how they're being looked at. Um and you can't run them on your own machine or in a provider of your choice. Uh and so that just immediately creates all of this uncertainty in a lot of enterprises.

rebuts · 2

12
Claim

Enterprises are more nervous about frontier (US) model providers than Chinese models due to greater confusion over data policies, storage, and inability to run models on their own infrastructure.

Alex Atallah states that when speaking to customers, enterprises are typically more concerned about using frontier models from US labs than about using Chinese models. The key reasons are opaque data policies, uncertainty about data storage and access, and the inability to host models on their own infrastructure.

transcript

Alex Atallah: I think they're they're more nervous about frontier models usually. Part because there's just like a a much there's much more confusion around the data policy about what's like actually happening to the prompts they're sending and um where they're being stored and how they're being looked at. Um and you can't run them on your own machine or in a provider of your choice. Uh and so that just immediately creates all of this uncertainty in a lot of enterprises.

explains mechanism · 1

13
Claim

Enterprises are more nervous about US frontier model providers than Chinese models because of uncertainty around data policies, where prompts are stored, and the inability to run models on their own infrastructure.

Alex Atallah explains that companies fear frontier labs like OpenAI and Anthropic partly because those labs have incentives to compete directly with thin wrapper startups, and partly because data handling policies create uncertainty that enterprises cannot resolve without on-premise options.

transcript

Alex Atallah: I think they're they're more nervous about frontier models usually. Part because there's just like a a much there's much more confusion around the data policy about what's like actually happening to the prompts they're sending and um where they're being stored and how they're being looked at. Um and you can't run them on your own machine or in a provider of your choice. Uh and so that just immediately creates all of this uncertainty in a lot of enterprises.

14
Claim

AI harnesses will persist as a valuable layer because they are more composable than traditional apps, allow developers to own user relationships, and are easier for AI to orchestrate due to their Unix-based architecture.

Alex Atallah argues that harnesses (AI coding tools like Cursor) are more composable than traditional apps, more inspectable for developers, and Unix-based architecture makes them easier for AI to orchestrate—giving non-model-lab developers a way to own user relationships.

transcript

Alex Atallah: The nice thing about the harnesses compared to the apps is that they're more composable. Like I can have a harness call another harness. I can have a harness spin up another harness in a sandbox in the cloud. ... it's much more reliable and and deterministic and sort of easy for users to grok with a harness because um the harnesses are Unix based. They all have and and the models are so well trained on on Unix, on bash commands ... Very very very very few unknown unknowns when you're composing around a harness. So I just think it just gives developers more flexibility um and flexibility that they can inspect.

15
Mechanism

In the age of AI, an employee's cost to the company should be a dynamic number based on their model usage efficiency, not a static salary, and performance should be assessed on a 'productivity vs. cost' quadrant.

Alex Atallah proposes a new framework for managing employee costs where an employee's total cost is dynamic, based on how efficiently they use expensive and cheap AI models. He suggests assessing performance on a quadrant matrix of productivity versus cost-effectiveness.

transcript

Alex Atallah: Really, your your employees all cost totally dynamic different amounts now... Your employees should like figure out which tools and models to use that are best for their tasks and then we should figure out how much you're costing be like due to the choices that you make as an employee... your your cost as an employee is a going to be a dynamic number and it's going to be, you know, dependent on how much that employee is like effectively using, you know, expensive and cheap models to do their job... I would I I advise companies to kind of like still do their normal management work... but also line it up with how much their employees cost, and then kind of come up with, you know, a quadrant of of celebration... And then a quadrant of concern.

16
Prediction

In the age of AI, employee costs should become dynamic numbers based on how efficiently workers choose and use AI models, rather than remaining static salaries — and managers should evaluate employees on both productivity and cost-effectiveness.

Alex Atallah proposes that as AI inference becomes a major operational expense, the concept of employee cost should shift from a static salary to a dynamic number reflecting how efficiently each employee uses expensive and cheap models — creating a quadrant of cost-effectiveness and productivity that managers can act on.

transcript

Alex Atallah: Really, your your employees all cost totally dynamic different amounts now and uh I think a lot of like companies are putting it on them to do routing. And I think in the future there's a good chance that it will like get pushed downwards to the employee level. Your employees should like figure out which tools and models to use that are best for their tasks and then we should figure out how much you're costing be like due to the choices that you make as an employee [snorts] and you know, your your cost as an employee is a going to be a dynamic number and it's going to be, you know, dependent on how much that employee is like effectively using, you know, expensive and cheap models to do their job. And uh and then I would I I advise companies to kind of like still do their normal management work, like have their managers kind of assess how how effective and productive employees are, but also line it up with how much their employees cost, and then kind of come up with, you know, a quadrant of of celebration. Like these employees are like doing a good job and they're pretty price effective or cost effective. And then a quadrant of concern, and these employees are kind of maybe doing a so-so job and whoa, they are not cost effective at all.

17
Prediction

The most exciting near-term AI applications are in previously inference-bottlenecked fields like rare disease research and crowdsourced solutions for urban/rural quality of life improvements.

When asked what excites him most, Alex Atallah highlights two areas: using AI to tackle rare disease research (previously limited by inference compute) and crowdsourcing ideas for tangible improvements to urban and rural life, leveraging global 'brilliant ideas' that were previously unfeasible.

transcript

Alex Atallah: One is rare disease research, which I think is one of those things that has been intelligence bottleneck or really just the inference bottleneck. Like it involves like trying out lots of ideas and seeing if they work. Um the other is crowdsourcing productive urban life improvements... like imagine if somebody had like if someone was curious about finding every lead pipe in America or every lead pipe in the UK... now you can use AI to do that... I'm excited about sort of very like broad kind of urban um like or urban or or rural like quality of life improvements that we'll be able to make.

explains mechanism · 2

18
Claim

In the age of AI, employee cost should be a dynamic number based on their model usage choices rather than a static salary, creating a new quadrant for evaluating employee cost-effectiveness.

Alex Atallah proposes that companies should track employee cost as a dynamic metric reflecting their AI model choices, and evaluate workers on a quadrant of productivity versus cost-effectiveness—something he believes is under-discussed in the AI transformation.

transcript

Alex Atallah: Really, your your employees all cost totally dynamic different amounts now and uh I think a lot of like companies are putting it on them to do routing. And I think in the future there's a good chance that it will like get pushed downwards to the employee level. Your employees should like figure out which tools and models to use that are best for their tasks and then we should figure out how much you're costing be like due to the choices that you make as an employee ... your your cost as an employee is a going to be a dynamic number and it's going to be, you know, dependent on how much that employee is like effectively using, you know, expensive and cheap models to do their job.

Highlight slides
Why a Multi-Model Future Is Inevitable✦ from: A multi-model future is inevitable because using multiple models together produces more creative ideas than any single model, and companies are always incentivized to experiment with the ecosystem's new models to improve productivity and reduce costs.The Multi-Model Future Is Inevitable✦ from: A multi-model future is inevitable, and companies that specialize in one proprietary model will still need to use other models created by the ecosystem to stay competitive.Consolidation on One Model Makes No Sense✦ from: A multi-model future is inevitable because using multiple models together produces more creative ideas than any single model, and companies are always incentivized to experiment with the ecosystem's new models to improve productivity and reduce costs.Why Two Models Beat One✦ from: A multi-model future is inevitable, and companies that specialize in one proprietary model will still need to use other models created by the ecosystem to stay competitive.Why Dedicated Players Win✦ from: Companies building routers because it's fashionable puts them months behind fully focused players like OpenRouter, reducing user leverage and access to the full model ecosystem.Cost to Users✦ from: Companies building routers because it's fashionable puts them months behind fully focused players like OpenRouter, reducing user leverage and access to the full model ecosystem.Jevons Paradox in Token Pricing✦ from: Token price reductions follow the Jevons paradox: when prices drop 10x, usage increases by more than 10x, so overall platform revenue grows rather than shrinks.GPT 5.6 Luna Case Study✦ from: Token price reductions follow the Jevons paradox: when prices drop 10x, usage increases by more than 10x, so overall platform revenue grows rather than shrinks.US Behind China in Open-Weight AI Models✦ from: America is very behind China in open-weight AI models, and the US needs better pathways to get compute to emerging labs to remain competitive.Compute Access: The Competitive Lever✦ from: America is very behind China in open-weight AI models, and the US needs better pathways to get compute to emerging labs to remain competitive.The Open-Weight Gap✦ from: The US is significantly behind China in the development of competitive open-weight AI models, though new American 'neo-labs' are beginning to emerge.US 'Neo-Labs' Begin to Emerge✦ from: The US is significantly behind China in the development of competitive open-weight AI models, though new American 'neo-labs' are beginning to emerge.Core Claim: Enterprises Nervous about Frontier Models✦ from: Enterprises are more nervous about frontier model providers like OpenAI and Anthropic than about Chinese models, due to confusion around data policies and the inability to run frontier models on their own infrastructure.Key Reasons for Enterprise Nervousness✦ from: Enterprises are more nervous about frontier model providers like OpenAI and Anthropic than about Chinese models, due to confusion around data policies and the inability to run frontier models on their own infrastructure.
Related episodes