They bought the books. They cut off the bindings.
A book can become a searchable file by losing the physical form that made it a book. Anthropic’s court record documents that conversion at scale. The questions concern what was purchased, what was copied, what was discarded and what legal permission can leave ethically unresolved.
The June 2025 order in Bartz v. Anthropic documents millions of purchased print books converted through the destructive process summarized above. [1] Destruction here means destruction of physical copies, not disappearance of every copy of a title.
The order found training and purchased-copy conversion fair use on this record, while refusing to excuse pirated library acquisitions on that basis. [1] Those distinctions prevent an accurate account of one practice becoming a blanket verdict on all AI data collection.
In July 2026, plaintiffs’ counsel reported final approval of a $1.5 billion settlement resolving unlawful book-acquisition claims. [4] That settlement concerns the acquisition dispute. It should not be presented as a judgment that cutting up purchased books was itself unlawful.
The invention examined here is a change in what counts as the valuable object. For a reader, the book is something to open, lend, annotate and return to. For a data pipeline, the desired object is usable text. Binding, paper and the possibility of another reader using that particular copy become costs or obstacles.
Scanning can make a work easier to search. That benefit is real. But calling the process digitization does not tell us what happened to its physical source. The word can describe a photograph of a page that remains on a shelf, or a process after which the original copy no longer exists.
The article’s objection is to allowing the digital result to erase the accounting for the material loss. A readable file may preserve the words while ending the life of that copy as a circulating object. Whether that trade is justified requires more than announcing that the file is useful.
Four stages need separate names: acquiring the source, making a digital copy, retaining it in a collection and selecting material for training. “Used for AI” can compress all four into a single act. It then becomes hard to ask whether a particular source was purchased, pirated, licensed or ever selected for a particular model.
US copyright law separates ownership of a material object from ownership of copyright in the work embodied in it. Section 202 makes that distinction explicit. [2] Buying a volume therefore does not itself purchase the author’s copyright. Whether a particular copying use is permitted requires a further legal basis.
Section 107 sets out a contextual fair-use inquiry involving purpose, the nature of the work, the amount used and market effect. [3] A slogan such as “we bought it” cannot do all that legal work. Nor can “it trained an AI” settle every question about acquisition, copying and output.
As an ethical proposal, each stage should leave a trace: edition and source, acquisition basis, scanning method, retention rules and training selection. That record would let an author or auditor ask a concrete question about a concrete work. It would also prevent a company’s general description of its pipeline from becoming an answer to every individual claim.
The court’s purchased-copy analysis emphasized replacement of one copy by another for storage and search, with no evidence the digital copies were shared or sold outside the company. [1] That rationale belongs in the record. It limits the claim that this decision authorized unlimited distribution of scans.
A fair-use finding answers a legal question in a particular dispute. The separate ethical questions are still available: Was destructive scanning necessary? Could a publisher have supplied a digital source? What would preserving the physical copies have cost? Was the work’s creator consulted? These are questions for evidence, not assumptions that an alternative was available for every title.
A technical study by A. Feder Cooper and colleagues tested extraction of memorized book text from thirteen open-weight language models. It found substantial extraction possible for some books, with results varying greatly by book and model. [5] This is evidence against assuming all models either retain nothing or reproduce everything. It is not a finding about Anthropic’s models, nor proof that the scans in this case produced a particular output.
Text in a source file, text selected for training and text recoverable from a model are different evidentiary objects. An honest account keeps them separate even when the distinctions complicate a convenient defence or accusation.
The court record documents purchased physical copies being destroyed during conversion. It does not establish that rare or irreplaceable editions were destroyed, that every scanned title trained a model, or that AI companies universally use this method. Those claims would require additional evidence.
The ethical argument reaches beyond whether a legal exception applies. A commercial system can benefit from accumulated writing while reducing the author’s contribution to an input. A purchased copy may satisfy one part of the transaction without settling whether the creator should have a voice in a new use. Legal permission, consent and recognition answer different questions.
Preservation also needs a declared purpose. Preservation for whom, with whose access, for how long and in what form? A private searchable collection can be useful to its owner. Public cultural preservation would require a different account of access and stewardship. Calling either one a library does not supply that account.
The book’s destruction is visible and literal. The larger interrogation concerns the decision that made its physical survival expendable, and the evidence needed to judge that decision rather than merely admire the resulting machine.
- For each scanned title, what was acquired, what was discarded and what evidence shows whether it entered a model’s training data?
- Why was destructive conversion chosen, which alternatives were examined and who had a say in the decision?
- When the result is called a library or preservation, who can access it, who controls its future uses and where is the creator in that arrangement?
The library became searchable. The books became disposable.
XORCIS.AI · Forensic satire
Sources
- Bartz v. Anthropic, order on fair use, 23 June 2025, docket 231: process and distinct copying uses. Primary court document hosted by Justia.
- US Copyright Office, Title 17, Section 202: material-copy ownership and copyright ownership are distinct.
- US Copyright Office, Title 17, Section 107: fair-use factors. Read 1 October 2026.
- Lieff Cabraser, 21 July 2026: plaintiffs’ counsel’s account of final settlement approval on 20 July. Settlement administrator’s current notice. Settlement is distinguished from the earlier fair-use ruling.
- Cooper et al., Extracting memorized pieces of (copyrighted) books from open-weight language models: research first submitted in 2025; current version read 1 October 2026. Different models from those at issue in Bartz.
Documented destruction of purchased copies is separated from piracy, training, memorization and the article’s ethical argument. The case makes no claim that unique editions were destroyed or that a particular scan caused a particular model output. US findings are not presented as a universal rule for other jurisdictions.