Distillation is a common industry practice, not IP theft, and the real issue is that closed labs like Anthropic could stop it with KYC but choose not to because it would slow their growth.
Chamath and Freeberg explain that distillation—training on another model's outputs—is a standard practice across every industry, that Anthropic could stop it with KYC but won't, and that it's hypocritical for Anthropic to call it IP theft while they themselves train on all the world's output without permission. ✦ AI generated
Chamath Palihapitiya · All-In Podcast · 2026-07-24 · original ↗
starts at this moment · 8:08
“Can you give people an idea of what distillation is and why it's so important here?”
Distillation is when you fire up a model and you ask it a question and you observe it and you take its output and you use that in training of your own model. Now multiply that behavior by tens of millions and what you exfiltrate is essentially trillions of questions and answers. And Sax is right. If you really care about distillation, you implement KYC. You force people to make an account, not just with a username and a password, but with some form of identification, maybe with a bounded credit card. There's all kinds of steps that you can take that would frankly slow things down in terms of revenue traction, but would solve the distillation problem on its face. So, that isn't really a thing. It's a bit of a red herring. The other thing on distillation is everybody has at some point distilled. The question is who is distilling from whom? And it looks like that funny meme where there's like nine Spider-Man all pointing at each other. That's what this is because Anthropic has distilled from all of these publishers. They just pay a $1.5 billion fine. Apparently, Open AI distilled from the New York Times. There's still an ongoing lawsuit. The Chinese labs have distilled from anthropic.
verbatim transcript · starts at 8:08
7:48you, Daario, this pod, other places talking about, hey, open source is ready. This is the moment. Get control, AI sovereignty, etc. They must be calling you up and saying, hey, okay, we're ready. Like, how do we get these things on prem? How do we do it? So what's the what's happening on those calls you're doing with the enterprise and then can you maybe give people an idea of what distillation is just and
8:11and why it's so important here. >> Let's start with the second thing. Distillation is when you fire up a model and you ask it a question and you observe it and you take its output and you use that in training of your own model. Now multiply that behavior by tens of millions and what you exfiltrate is essentially trillions of questions and answers. And Sax is right. If you really care
8:44about distillation, you implement KYC. You force people to make an account, not just with a username and a password, but with some form of identification, maybe with a bounded credit card. [snorts] There's all kinds of steps that you can take that would frankly slow things down in terms of revenue traction, but would solve the distillation problem on its face. So, that isn't really a thing. It's a bit of
9:14a red herring. The other thing on distillation is everybody has at some point distilled. The question is who is distilling from whom? And it looks like that funny meme where there's like nine Spider-Man all pointing at each other. That's what this is because Anthropic has distilled from all of these publishers. They just pay a $ 1.5 billion fine. Apparently, Open AI distilled from the New York Times.
9:41There's still an ongoing lawsuit. the Chinese labs have distilled from anthropic. >> It's wholesale stealing everywhere. Yeah. >> Well, I don't want to call it stealing because it's not clear who actually owns the copyright in the first place. But here should be the important observation. These models are getting commoditized much faster than anybody thought. And how do we know this? Because there is no meaningful sustained advantage
10:09once a model publishes their performance criteria. What you see is literally within weeks other models some open some closed some open weight who are able to match and in some cases exceed the performance. So I think what's happening here is a handful of American companies have realized whoa this value that we are seeing today may not be sustainable in a 5 and 10 year period and when you
- ·Querying a model and using its outputs to train your own
- ·Exfiltrates trillions of questions and answers at scale
- ·Calling it IP theft is a red herring
- ·Anthropic could stop distillation with KYC identification
- ·They choose not to because it would slow revenue traction
- ·Anthropic itself trained on publisher content, paid $1.5B fine
- ·"Nine Spider-Man pointing at each other" — all labs do it
- ·OpenAI — NYT lawsuit ongoing
- ·Chinese labs distilled from Anthropic