The court filings in Bartz v. Anthropic PBC describe an operation the company codenamed “Project Panama” — a systematic effort to acquire physical books, strip their bindings, cut their pages, scan the contents into a training dataset for Claude, and discard the originals. The company’s stated rationale for the codename, per the exhibits, was that it did not want the project publicly known. These details come from reporting by Kathryn James in the Guardian, drawing on exhibits and filings before Judge William Alsup in the Northern District of California.

It is true that Alsup found the training itself — feeding digitized text into a large language model to improve next-token prediction — to be transformative use within the meaning of fair use. The training ruling addresses a narrow question: whether the act of ingestion for model development constitutes infringement. It does not address the procurement pathway that brought those texts into Anthropic’s possession, the destruction of the physical objects, or the competitive effects on the markets for the original works. The fair-use question the court answered and the question the case raises about how those books were obtained are not the same question, and the gap between them is where Anthropic’s operation lived.

The procurement pathway, as the court exhibits describe it, followed a sequence. Anthropic first attempted to acquire texts from pirated sources. When that path proved legally untenable — the company subsequently settled related claims, with the settlement reported at approximately $1.5 billion, a figure drawn from media reporting on the settlement rather than a publicly filed judgment — Anthropic turned to purchasing physical books and destroying them to extract their text. The company’s co-founder and CEO had earlier described the alternative — negotiating copyright permission — as a “legal/practice/business slog.” The exhibits show Anthropic becoming “not so gung ho” about pirated sources “for legal reasons.” The destructive-scanning path occupied the space between those two retreats: apparently designed to be less legally risky than piracy and less costly than licensing.

This is where the engineering-substance distinction matters. Anthropic frames the project as machine-learning research. The court framed the ingestion as transformative use. The internal memo, per the exhibits, framed it as something the company preferred not to have publicly known. What the court exhibits describe — warehouses of books, vendors stripping bindings, staff managing the workflow — is a physical supply chain for extracting text from copyrighted works without the consent of the authors who wrote them. The technical process is straightforward: physical book to scanner to OCR text to training corpus. The political economy is the question that matters: the company chose to acquire, transport, warehouse, slice, scan, and destroy physical books rather than negotiate with the people who wrote them, and fair-use doctrine as currently interpreted addressed only whether the ingestion was permissible, not whether the procurement and destruction should have been.

Doctorow and Giblin’s account of how intermediaries extract value from creators — positioning between the work and the audience, taking what they need, and leaving the creator holding a right that the law has been interpreted to let the intermediary walk through — describes this architecture from the creator’s side. The gap between the ingestion ruling and the procurement pathway is not incidental to that architecture. It is the architecture. If fair use permits the purchase and destruction of copyrighted works to produce training data — if the ingestion is “transformative” regardless of how the text arrived — then every library, every archive, every secondhand bookshop becomes a procurement source for any company willing to buy, destroy, and digitize at scale. The authors do not need to be consulted. The publishers do not need to be notified. The texts just need to exist in physical form, and the law as interpreted in this ruling permits the extraction while declining to examine the pathway. The value of the work flows upward to the company that consumed it. The cost — the destroyed books, the uncompensated labour, the cultural objects reduced to data and then to waste — flows downward to the people who produced and preserved them.

The U.S. Copyright Office has been examining the intersection of copyright and generative AI since 2023, but what is not currently under review is the physical-destruction pathway this case exposed. The court’s fair-use ruling addresses ingestion; the question of whether purchasing and destroying copyrighted books to build commercial training data should be lawful is a question that remains open. The gap will not close on its own.