ATRIUMsearch → argument graph
Video · 2026-08-03 · 1h 10m · 6 moments

Arena CEO: There Will be a $100BN US Open-Source Model & Data is a Trillion Dollar Market

✦ AI generated

timeline · colored by role

01
Claim

Chinese labs like Kimi are no longer merely distilling American models; Kimi K3 beat the best closed-source American models on important subsets of tasks, violating the persistent American narrative that China only keeps up through distillation.

Anastasios calls Kimi's open-source beat over closed-source American models a big moment because it breaks the narrative that Chinese labs only succeed by distilling American models. Even if they still use distillation as a sub-step, their performance above American labs shows they are doing something more.

transcript

Anastasios: It was a pretty big moment. And the reason I'd say it was a big moment is because it violates a narrative that has been persistent in the United States, which is that the Chinese are just distilling American models. And that's the only way that they're able to, you know, keep up. When really what happened is that Kimmy actually beat all American models, including Fable, in some subset of tasks. That doesn't mean that they're not distilling. They may still be using distillation as a substep in their training procedure, but it does mean that distillation is only part of the story and that there's something that those labs are doing above and beyond distillation that's bringing the performance up above what the American labs are currently doing.

02
Claim

Enterprises will increasingly want AI sovereignty — owning their whole AI supply chain — because in the age of AI software is no longer a moat, while data and network effects are, and companies will not want to hand their data to a third party that may one day compete.

Anastasios argues AI sovereignty is inevitable because businesses need a moat: software can be produced instantaneously, but data moats and network effects endure. Companies can take open-source models, fine-tune them on their own data, and own their stack to avoid cost, sovereignty, and supply-chain concerns.

transcript

Anastasios: Enterprises are going to want to own their own intelligence. They're going to want so-called AI sovereignty, which is a fancy word for meaning that you own your whole supply chain of AI. That means you can take an open-source model and you can fine-tune it on your own company's data and own your stack end to end basically outside of the compute hosting. ... In the age of AI, software is no longer really a moat because it can be produced instantaneously, right? Or let's project out five years. That's what's going to happen. And so what modes exist? Network effects exist and data modes exist. And if you can take your data mode and turn it into a self-improving product that is a way for businesses to remain sustainable in the age of AI. ... People are going to care about sovereignty. People are going to care about cost. People are going to care about self-improving. and they're not necessarily going to want to give their data to an external third party service that might even be competing with them one day.

explains mechanism · 1

03
Prediction

Given the US regulatory environment, a great American open-source competitor will arise — at least one multi-hundred-billion or trillion-dollar American company focused on American-first open source.

Anastasios predicts a huge American open-source company will emerge because, while Chinese models may play a role now, the US regulatory environment favors a domestic champion. The durable business model combines open-source models as a lead-generation tool with deployed engineering for AI modernization.

transcript

Anastasios: I think not. I think that the Chinese models will be potentially part of the story for now. Um, but that given the regulatory environment in the US, it's probably more likely in the long run that we see a great American open-source competitor arise. And this is why I've been a strong proponent, for example, of thinking machines. I believe that we're going to have at least one massive, you know, multiundred billion if not trillion dollar American company focused on American first open source.

provides context · 1

04
Anecdote

We are heading into an era of unprecedented cyber attacks, and AI-generated fake job applicants who pass technical interviews are already infiltrating American businesses.

Anastasios describes AI-generated fake people applying for engineering roles, passing technical interviews with top engineers, then turning out not to exist. He frames this as widespread, not just at Arena, and a harbinger of an insane era of cyber attacks.

transcript

Anastasios: It's going to be so [ ] insane what happens with like the cyber attacks because here's what we see at Arena. We see another dude on the other side of the interview. They come in, they're like, 'Hey, I want to be an infrastructure engineer at Arena, which is a great job that we're hiring for you.' ... some guy looks perfectly normal. They're getting, you know, they're passing all of our technical interviews. They're like such an amazing blah blah blah. And then what happens at the end of it? You try to hire them and it's vaporware. Person doesn't [ ] exist. I'm not kidding. I am not kidding you. I don't know whether this is corporate espionage or cyber attacks or, you know, nation states, but people are trying to get into all of the American businesses. And we're not the only ones. This is happening everywhere. fake people applying to companies.

05
Prediction

Data is a durable scaling complement to AI models, making the data market at least a $100 billion industry by 2030, if not a trillion.

Anastasios argues there are two types of hypergrowth markets, one being 'scaling complements' like data, whose demand scales with model growth. Because data only becomes irrelevant when AGI is achieved, it is more durable than GPUs, and he projects the market at $100B+ by 2030.

transcript

Anastasios: So, I have a thesis on hyperrowth. There's two types of hyperrowth markets that we see today. Market A is what I call scaling compliments and these are goods that are complimentary goods to the scaling of AI models. And I mean that in the economic sense. A complimentary good is good A and B are the good A is a complement to good B if the demand for good B drives demand for good A. So if I have a car, gas is a complimentary good to cars. The more cars are sold, the more gas is sold. And so data is one of these scaling compliments because the bigger models scale, the more data you need. And that's a scaling law question. And so the more models you get and the bigger that they're getting, the more they're proliferating. The more businesses are training their own models, the more data you are going to need. And it's a very fundamental need. People forget this. They think about data as a commodity. It's really not. It's actually less so of a commodity than even GPUs because in order for data to become uh irrelevant, humans need to become irrelevant and that means that we've achieved AGI. So data is a very durable need and companies are spending on it usually with within frontier labs at about 10 to 20% about the amount that they're spending on GPUs. And so if you believe in the GPU market accelerating, if you believe in the scaling of models, if you believe this is going to be a big industry that keeps accelerating and growing, then absolutely you should believe in the data market. I believe it's going to be at least hundred billion dollars by 2030, if not a trillion.

06
Claim

Evaluation is the single biggest bottleneck to deploying AI because people don't understand how to define value, and every business in the world will unambiguously need it.

Anastasios asserts that evaluation — defining and measuring value — is the largest obstacle to AI deployment, making it an unambiguous necessity for every business, since cost is easy to define but performance depends on each business and use case.

transcript

Anastasios: I think every business in the world is going to need evaluation unambiguously and that is the single biggest bottleneck to deploying AI because people don't understand how to define value. all this co all this like stuff around cost for value. It's like how do you define value? It's easy to cook costs. I can tell you to go use you know Gemini Flash and that's going to be like way more efficient in terms of token spend.

Highlight slides
Related episodes