Anthropic Downloaded 4 Million Books. Its Settlement Covers 500,000.
Anthropic downloaded about 7 million files from Library Genesis and Pirate Library Mirror, two pirate book sites. Roughly 3 million were duplicates. Of the 4 million unique works left, its $1.5 billion settlement covers about 500,000.
The rest fell out on paperwork. To be covered, a book needed an ISBN or an Amazon identifier and a U.S. copyright registration filed within five years of publication. That registration also had to come before Anthropic downloaded the book, or within three months of the book’s first publication. About 2.5 million of the unique works were in languages other than English, and most of those were never registered with the U.S. Copyright Office at all.
There is a reason for those deadlines. Registering on time is what lets a copyright owner ask for statutory damages, which is money a court can award without making the owner prove what a particular copy cost them. Without that, proving harm from a book sitting in a training set is nearly impossible, and a case this big gets too expensive to bring. The paperwork is what makes the lawsuit possible. It also decides who the lawsuit can include.
Authors whose books fall outside all of that get no payment. They also give up nothing, since the deal only settles claims over books on its list. They keep the right to sue, which is worth whatever they can afford to spend on a lawyer.
The list of covered books is public and I looked up my own work. None of it is there, which tells me less than it sounds like it does. I publish magazine essays, a newsletter and bylined reporting, so a settlement about books was never going to reach me. What I still do not know is whether anything I have written has passed through a training set. Nothing exists that would tell me. That question comes before any question about payment, and for most working writers it is the one that stays open.
Lawmakers have proposed ways to answer it. The CLEAR Act, introduced in February by Sens. Adam Schiff and John Curtis, would make AI developers file a summary of the copyrighted works in their training data with the Copyright Office before release, in a database anyone can search. The TRAIN Act would let a copyright owner who has reason to believe their work was used subpoena the training records. Neither has become law.
Both start from the same place. The CLEAR Act counts a work as covered only if it is registered with the Copyright Office, so only owners who registered could enforce it. It does not copy the settlement’s ISBN rule or its five-year deadline, so it would reach a wider group than this case did. The barrier is lower. It sits in the same place. The TRAIN Act does not write registration into its text. But the owner has to start the process, and the records stay confidential, so it gives one writer an answer rather than making a public record.
This is what accountability looks like when a copyright lawsuit is the only tool on hand. Judge William Alsup, who handled the case before retiring, ruled in June 2025 that training a model on lawfully purchased books is fair use. He left the downloading itself for trial. When he let the case go forward as a group lawsuit, he did it on the downloading alone. The deal covers only what Anthropic did before Aug. 25, 2025. It does not touch anything a model produces, and it says plainly that it is not a license to train on copyrighted work. Anthropic has to destroy the downloaded files and says neither dataset was in a model it shipped. Because the company settled rather than going to trial, no appeals court will review any of it.
For the books that do qualify, private contracts finish what the court started. The settlement splits each award 50-50 between author and publisher by default, which covers trade books and university press books. Textbook contracts get no automatic split at all, because they vary too much for one rule to fit them. Those authors have to state the share they believe they are owed and produce the contract if anyone disputes it. The Authors Guild reported that some education publishers are claiming shares based on their royalty rates, so an author earning a 10% royalty can face a publisher claiming 90% of the payment. Authors can dispute that, and nobody gets paid on the book until it is resolved. When claimants cannot agree, a court-appointed referee called a Special Master decides, and there is no appeal.
I have seen this from the other end of it. I spent months on research for The Black Policy Institute about AI-powered misinformation in Kenya and Nigeria. What kept coming back was that the money to build these systems and the attention to govern them collect in the same few places. Internal Facebook documents disclosed by a whistleblower in 2021 showed the company putting 87% of its misinformation budget into English-language content while English speakers made up 9% of its users. Moderation budgets and copyright registration work nothing alike. They still point at the same few places, and the same people keep ending up outside.
Licensing deals are next, and they are harder to get into, since a deal takes a lawyer and something to bargain with. Even then, the size of the check settles less than it seems to. Researchers led by Phoenix Perry at Goldsmiths posted a paper, ahead of peer review, arguing that a creator can consent to training and still have little say over the model that results, or over who adapts it, sells it, or runs it next. The lawsuit settled how the books were obtained. A license will settle a price. Neither one settles who governs the model.
A $1.5 billion payout is real money, and it is the largest anyone has gotten out of an AI company. What the number cannot tell you is how many writers have no way to learn whether their work was used, and no affordable way to act on it if it was. Every fix on the table, this settlement and both bills, starts by counting who qualifies. The writers who have never been counted stay where they are.