Judge rules AI training on copyrighted material is not, by itself, infringement
Anthropic arranged the work through a logistics manager, vendors, staff and a warehouse. According to the court’s decision, the vendors “stripped the books from their bindings, cut their pages to size, and scanned the books into digital form – discarding the paper originals.” Case exhibits showed warehouses of books stacked and labelled on shelves, with staff moving between them.
An internal memo introduced as court exhibit 21 described an initiative code-named “Project Panama.” The memo advised discretion: “Why use a codename? … [B]ecause we don’t want it to be known that we are working on this.” It answered “What is Project Panama?” by stating: “Project Panama is our effort to destructively scan all the books in the world.”
The project was driven by a data need. To train Claude, Anthropic sought a large, high-quality language dataset, preferably one created before 2022. The court’s decision noted that Anthropic believed books’ “well-curated facts, well-organized analyses, and captivating fictional narratives” would help “Claude write as accurately and as compellingly as Authors.”
Anthropic’s co-founder and CEO had characterized securing copyright permission for existing e-books as a “legal/practice/business slog.” The company first chose to use pirated sources instead, a decision that informed its $1.5bn out-of-court settlement with authors. The court decision described Anthropic as having become “not so gung ho about” training on pirated books “for legal reasons,” after which the company turned to destructive scanning.
Judge Alsup, applying the fair use doctrine in US copyright law, ruled that using copyrighted books to train an AI model did not, in and of itself, constitute infringement. He observed that one version replaced the other and added: “There is no evidence that the new, digital copy was shown, shared, or sold outside the company.” Anthropic has stated its intention to create a “forever” research library.
Kathryn James, the rare book librarian at the Lillian Goldman Law Library at Yale University, wrote in The Guardian that the case raises unresolved questions about the cultural status of printed works. She noted that the US has few legal provisions governing the treatment of books as cultural objects and wrote that destructive scanning is “both legal and less regulated than strip mining.” James argued that the broader risk is that creators cede “the means of production of our large language lives” and “turn from creators to consumers,” handing the generative promise of their work to language model proprietors.
Reflecting on the significance of the works involved, Judge Alsup wrote: “For centuries, we have read and re-read books. We have admired, memorized, and internalized their sweeping themes, their substantive points, and their stylistic solutions to recurring writing problems.”