Anthropic cut up a million books to train Claude. A court said that part was legal
What the Bartz v. Anthropic record shows: sliced spines, recycled paper, a fair use win, and a $1.5 billion bill for piracy.
Reports of the book destruction often leave out what the court actually considered. Anthropic bought millions of used print books, had the bindings stripped off and the pages cut to size, ran the loose sheets through document scanners, and sent the paper to be recycled. A federal judge then looked at that operation and ruled it lawful.
The $1.5 billion the company later agreed to pay had nothing to do with the cutting. It was for roughly seven million books downloaded from pirate libraries years before anyone at Anthropic thought about buying a warehouse of secondhand paperbacks.
What follows is the case, entry by entry. Quotations come from Judge William Alsup's order of 23 June 2025 in Bartz v. Anthropic PBC, case 3:24-cv-05417, Northern District of California, unless another source is named. Each entry closes with a verdict line.
The timeline
Date | What happened |
|---|---|
2021 to 2022 | Anthropic downloads Books3, LibGen and PiLiMi. The order puts the total at over seven million pirated copies |
February 2024 | Anthropic hires Tom Turvey, formerly head of partnerships for Google's book scanning project, and tasks him with obtaining "all the books in the world" |
Early 2024 | Bulk buying of used print books begins, reported as Project Panama |
August 2024 | Authors Andrea Bartz, Charles Graeber and Kirk Wallace Johnson file a class action |
23 June 2025 | Alsup rules: training is fair use, print to digital conversion is fair use, the pirated library is not |
September 2025 | Anthropic agrees to pay $1.5 billion, about $3,000 per work plus interest |
July 2026 | Final approval granted. Reported as the largest known settlement in a US copyright case |
Entry 1: What actually happened to the paper
Anthropic's team emailed book distributors and retailers about bulk buying their print stock for a "research library". The order records that the company "spent many millions of dollars to purchase millions of print books, often in used condition. Then, its service providers stripped the books from their bindings, cut their pages to size, and scanned the books into digital form, discarding the paper originals."
Press reporting fills in the last step. Coverage of Project Panama describes hydraulic cutters for spine removal, high speed document scanners, and the loose paper going to recycling. Each book came out the other side as a PDF of page images with machine readable text, including front and back cover scans for softcovers.
Verdict: the books were dismantled and the paper recycled. It was an industrial scanning line run by contractors.
Entry 2: Where the books came from
The print stock came out of the ordinary secondhand trade. Reporting on Project Panama names online secondhand retailers including Better World Books, World of Books and Zoom Books. In August 2026 the Irish Times reported booksellers across Europe fielding orders they could not explain, including what a Galway bookseller called a "bananas" order for 5,000 titles, and Berlin shops handling mass orders worth €10,000 to €50,000 that arrived through automated systems.
Anthropic told the paper that sourcing books is "a widely used approach for training large language models across the AI industry", and denied buying rare or antiquarian books to destroy them. Booksellers quoted in the piece stayed sceptical, suspecting third party providers were doing the buying on the company's behalf. Treat the European sourcing chain as reported rather than established.
Verdict: the print supply came from the ordinary secondhand trade, bought at retail scale, which is exactly why the court treated those copies as owned.
Entry 3: Why training on books was held to be fair use
Alsup was blunt about the first fair use factor. The "purpose and character" of using works to train a large language model, he wrote, was transformative, "spectacularly so", and later in the order, "quintessentially transformative". The comparison he reached for was a reader who becomes a writer.
The authors argued that training would flood the market with competing work. The order assumes that is true and still rejects it as a copyright injury: "Authors' complaint is no different than it would be if they complained that training schoolchildren to write well would result in an explosion of competing works. This is not the kind of competitive or creative displacement that concerns the Copyright Act."
One limit matters. The authors did not allege that Claude reproduced their books in its output, and the order says the record shows the opposite. Alsup explicitly left the door open: "Authors remain free to bring that case in the future should such facts develop."
Verdict: training on lawfully held books is fair use in this court, on this record, and the ruling is narrower than the headlines suggested.
Entry 4: Why cutting up the books was also fair use
This is the part that reads as counterintuitive and turns out to be the most conventional. Anthropic already owned the print copies. Converting them to digital did not add a copy to the world.
The order leans on Sony Betamax, the Texaco microfilm reasoning and Google Books, then settles it in two short sentences: "The print original was destroyed. One replaced the other." Nothing was shared, sold or shown outside the company. Alsup placed storage and searchability outside the creative work itself: they are physical properties of the frame around it, or informational properties about it.
The third fair use factor, how much was copied, also favoured Anthropic here. Copying the whole book was precisely what the purpose required. In the order's phrase, "There was no surplus copying."
Verdict: destructive scanning of a book you own is a format change, and format changes have been surviving fair use analysis since the videocassette recorder.
Entry 5: Why the pirated library was not fair use
Between 2021 and 2022 Anthropic downloaded 196,640 books from Books3, at least five million copies from Library Genesis and at least two million from the Pirate Library Mirror. The order's total: "over seven million copies of books."
The company's own internal messages did not help. It had become "not so gung ho about" training on pirated books "for legal reasons", and kept them anyway, retaining copies even after deciding they would never be used for training again. Alsup called that acquisition and retention its own use, requiring its own justification, and found none offered "except for Anthropic's pocketbook and convenience". Downloading a pirated copy, he wrote, is "inherently, irredeemably infringing even if the pirated copies are immediately used for the transformative use and immediately discarded".
Buying the book later did not fix it: "That Anthropic later bought a copy of a book it earlier stole off the internet will not absolve it of liability for the theft but it may affect the extent of statutory damages."
Verdict: acquisition was the whole case. How you get the book decided this lawsuit, and what you do with it afterwards did not.
Entry 6: What the settlement actually is
Facing a damages trial with statutory exposure on every pirated work, Anthropic settled in September 2025 for $1.5 billion, amounting to about $3,000 per book plus interest. Dividing one figure by the other implies a class of roughly 500,000 works. Treat that as our arithmetic, because no source we could reach publishes the count directly.
Final approval came in July 2026. Reuters reported the approval on 20 July 2026, naming Judge Araceli Martínez-Olguín and calling it the largest known settlement in a US copyright case. Some authors and publishers opted out and are continuing separate suits. Reporting also indicates the settlement obliged Anthropic to destroy its pirated copies, which we could not confirm from a primary filing.
Verdict: the largest copyright settlement in US history traces back to a torrent client running four years before the first spine was cut.
Entry 7: What this says about how any model gets trained
Three things generalise beyond this one company.
First, provenance is now a line item. The order draws a hard boundary between a copy you are entitled to hold and one you are not, and that boundary survived even where the downstream use was found to be "spectacularly" transformative.
Second, licensing was a road not taken. The order notes that Anthropic's own hire had opened conversations with publishers in spring 2024 and let them wither, and that another major technology company reached an agreement with a major publisher around the same time. The order does not name it, and neither will we.
Third, the output question is still open everywhere. This case was decided on inputs. Alsup was careful to say that a case about infringing outputs would be a different case.
Verdict: the industry's legal risk sits in the acquisition log, and every serious lab now knows it.
What it changes for you
Very little, day to day, which is worth saying plainly. Claude is trained the way it is trained, and the models you reach through MultiChats come from 18 providers with their own sourcing histories and their own pending litigation. Nothing here gives any of us a clean bill of health on training data, and nobody who tells you otherwise has read the order.
The distinction is between scanning purchased books and acquiring pirated copies. Anthropic ran a book scanning line that destroyed every copy it processed, and won that argument in court. It ran a torrent client four years earlier and lost that one for $1.5 billion.
Sources
Order on Fair Use, Bartz v. Anthropic PBC, No. C 24-05417 WHA (N.D. Cal., 23 June 2025), full text: https://storage.courtlistener.com/recap/gov.uscourts.cand.434709/gov.uscourts.cand.434709.231.0_3.pdf
Bartz v. Anthropic, Wikipedia (settlement figures, final approval, opt outs): https://en.wikipedia.org/wiki/Bartz_v._Anthropic
Project Panama, Wikipedia (retailers, hydraulic cutters, paper recycling): https://en.wikipedia.org/wiki/Project_Panama
"A 'bananas' order for 5,000 obscure titles from a bookshop in Galway fuels suspicion", The Irish Times, 10 August 2026: https://www.irishtimes.com/world/europe/2026/08/10/a-mysterious-buying-spree-is-unsettling-europes-booksellers/