ATRIUMsearch → argument graph
MechanismVideo · 4:17 — 6:25

Distillation — using outputs from one AI model to tweak or fine-tune another — is extremely hard if not impossible to detect or prove, even though it typically violates terms of service.

CJ explains the process of distillation, notes that it typically violates providers' terms of service, and concludes that determining whether one model was distilled from another is essentially impossible. ✦ AI generated

CJ · Syntax · 2026-07-30 · original ↗

starts at this moment · 4:17

And this is where this idea of distillation comes in. It's essentially where you can prompt a model to see how it responds, then save that response in its own file. So these AI labs create these data sets of questions and answers and some of those questions and answers come from an existing model and there are actually a lot of open data sets online that are essentially this distillation of these models because the models really are just a black box. They predict words but there's no way for us to really know what's going on inside of them. So in order to figure out their capabilities, we essentially prompt them over and over again and then we save all of the good responses in a file that include this is the question that was asked and then this was the response that I got back. We can then further tweak and fine-tune the numbers that are already inside of our model using that data. And this process is known as distillation. And some AI labs are accusing other AI labs of distilling their models. Now, distillation typically violates the terms of service of a given provider of an AI model. And in their terms, they typically say something like, 'You should not be using the output of our models to create, tweak, or fine-tune your own models.' And that is happening. But what's extremely hard to do, if not impossible, is to determine if some distilled work from some model was used to create, tweak, or fine-tune some other model.

verbatim transcript · starts at 4:17

Transcript · around this moment

3:58to anything. It it can't access the internet. It is literally just a file that when run through some other program predicts words. Now, the process of creating these models and analyzing all of this data is where a lot of the news is coming from. Because when you slurp up all the data from the internet, there's going to be a lot of copyrighted works in there as well as works from

4:17news organizations or individual blogs or musicians or artists. and they didn't necessarily agree up front that their data could be used to create this model file. And these models were not only trained on all of the data that exists on the internet, but also trained on more curated or synthetic data. And this is where this idea of distillation comes in. [music] It's essentially where you can prompt a model to see how it

4:41responds, then save that response in its own file. So these AI labs create these data sets of questions and answers and some of those questions and answers come from an existing model and there are actually a lot of open data sets online that are essentially this distillation of these models [music] because the models really are just a black box. They predict words but there's no way for us

5:02to really know what's going on inside of them. So in order to figure out their capabilities, we essentially prompt them over and over again and then we save all of the good responses in a file that include this is the question that was asked and then this was the response that I got back. We can then further tweak and fine-tune the numbers that are already inside of our model using that

5:20data. And this process is known as distillation. And some AI labs are accusing other AI labs of distilling their models. Now, distillation typically violates the terms of service of a given provider of an AI model. And in their terms, they typically say something like, "You should not be using the output of our models to create, tweak, or fine-tune your own models." And that is happening. But what's

5:43extremely hard to do, if not impossible, is to determine if some distilled work from some model was used to create, tweak, or fine-tune some other model. You can of course poke and prod and compare the outputs of one to to another to see that oh they respond in a similar way. [music] And the owners of these AI labs can watch their network traffic to see who's making requests to these

6:06models to maybe pin it on some other company. That's the extent of it really. Why all of this stuff is so cloudy and why it's so hard to judge is these models are black boxes. There's no way for us to figure out what's actually inside of them. So that's who's creating the models and how they create them. Now, let's talk about how you get access to these models because you can get

6:25access to them in various ways. And one of the first main ways that people get access to these models is via what I'm going to call a restaurant or a closed model. And so, essentially, when you go to a restaurant, the restaurant owns that entire experience. You order from their menu, the kitchen makes it, they have chefs that they trust, you then eat your food, pay, tip your waiter, but

Around this claim