ATRIUMsearch → argument graph
MechanismVideo · 6:19 — 7:49

Chinese open-source AI labs are compounding progress by distilling and building on each other's models, which could let them surpass the best US proprietary models by late 2025.

Sunny Madra explains that China's cluster of open-weight model makers can remix and distill each other's work instead of training in silos, accelerating both frontier capability and smaller 'turbo' model quality. ✦ AI generated

Sunny Madra · BG2 Pod · 2025-07-31 · original ↗

starts at this moment · 6:19

Elicited by

What is your theory of the case? Why is China, you know, doing so well in open source? And should US model companies like OpenAI and Anthropic be concerned?

Instead of working in silos and instead of having to create giant training clusters separately they can basically take each other's work, build on top of it, almost consider it like a remix of someone's model — K2 sort of a well-known remix of what Deepseek had done — and now we're starting to see that happen really fast.

verbatim transcript · starts at 6:19

Transcript · around this moment

6:19point around using copyrighted work and he says, you know, he used a great example. If you read a book and you use it, you're not violating the copyright there. And so, he addressed that concern and that was one of the major things that, you know, a lot of people didn't talk about, but I think it's important for the model makers. And so the Chinese just have been able to work around that

6:38because of, you know, their position on IP. And what we're really seeing here, and I think, you know, Bill teed it up even better off off of my tweet, which is, um, you know, they're able to compound. So what you're seeing very quickly is both the open- source nature, the open weights nature, uh, allow them to basically compound on each other. So instead of working in silos and instead

6:58of having to create giant training clusters um separately they can basically take each other's work build on top of it almost consider it like a remix of someone's model K2 sort of a well-known remix of what what Deepseek had done and now we're starting to see that happen really fast and we're seeing two dimensions of it going quickly. One we're seeing the leading edge models getting quicker and we're then seeing

7:21them distill down smaller you know turbo models really really fast as well. There was a a release today of a quen you know 30 billion parameter model which is performing as good as GPT40. So think about that right and GPT40 was you know world class not that long ago. So those are the reasons that we're really seeing an acceleration right now. >> You know it I want to dig into this

7:43model development in particular. You know a year ago we were talking about these models being these you know stochastic parrots um you know and we really had to compress the entire internet. So you go back to GPT4, you know, and and and you're compressing, you know, the entire internet. But now we really don't need to do it because we've trained them to use tools like the internet, right? The they're true

8:05reasoning engines. When I ask a question today, you know, it doesn't just spit out an answer immediately. It goes and uses the tool and it searches the internet. Um, so if you don't have to compress all this wikipedia information, I don't know, take take a subject like World War II, you just need to know how to go out and use the internet to find the information to summarize it in real

Around this claim