Anthropic's position on training data is breathtakingly hypocritical: it claims the right to train on all the world's copyrighted output for free even if the creator objects, while forbidding others from training on Anthropic's own output — even though the courts have ruled that LLM-generated output is not copyrightable because it was not created by a human.
On the story that Anthropic is bulk-buying and shredding rare books for training data, Sacks holds his fair-use position but attacks the double standard: Anthropic trains on the world's output for free yet prohibits training on Anthropic's output even if paid for, despite court rulings that LLM output cannot be copyrighted. ✦ AI generated
David Sacks · All-In Podcast · 2026-07-31 · original ↗
starts at this moment · 70:41
“But what do you think just about the books being destroyed and you know being used in this way? It obviously has made people a little emotional about it.”
Is breathtaking hypocrisy for anthropic to maintain that it is entitled to train on all the world's output for free even if the creator objects. But the one type of output that you're not allowed to train on is their output even if you pay for it. That is their current position. So you know what I'm saying is that you know if you want to train on anthropics output that cannot be considered IP theft under fair use especially given the fact that the courts have ruled that LLM generated output is not copyrightable because it was not created by a human. That is the current position of the courts is that LLM output cannot be copyrighted.
verbatim transcript · starts at 70:41
70:41on all the world's output for free even if the creator objects. But the one type of output that you're not allowed to train on is their output even if you pay for it. That is their current position. So you know what I'm saying is that you know if you want to train on anthropics output that cannot be considered IP theft under fair use especially given the fact that the courts have ruled that
71:08LLM generated output is not copyrightable because it was not created by a human. That is the current position of the courts is that LLM output cannot be copyrighted. So there's no IP theft here. You can make the argument and I think it's probably true that if a competitor creates massive numbers of fake accounts on your service, >> you're breaking the terms of service. That's a definitely a break in the terms
71:28of service and it's probably a deceptive business practice and there may be other things you can do but I don't think you can >> depending on the jurisdiction by the way because in Philippines, Israel, India, they have different rules about like breaking the terms of service which LinkedIn found out >> when people started scraping their data. Chimoth any any thoughts here before we move on to socialism corner everybody's
71:50favorite new feature here on the oil and pod don't cut to the books. uh keeping the books intact. >> Why do you cut the books in the library? It's very hard to read them. There's a no spine. >> I think uh I think it's not the kind and I don't like to cut the books. >> I mean, you take your time, you move the page like a Google does. It's a little