Why is Anthropic destroying books? | Kathryn James

← Back to the feed

Why is Anthropic destroying books? | Kathryn James

The Guardian · 2 hours ago

Court documents from the Bartz v Anthropic PBC case have revealed that the AI company ran a secret initiative, internally codenamed "Project Panama," to destructively scan millions of physical books to build training data for its Claude language model. Rather than negotiate copyright permissions with authors and publishers, Anthropic apparently found it simpler to buy printed books, cut their spines, scan them and dispose of the originals, exploiting the "transformative use" provision of US fair use law. The case raises broader concerns about how AI firms treat the written works they rely on, and whether legal technicalities are being used to bypass proper engagement with creators.

Internal memos cited in the case show Anthropic acknowledged wanting to keep the project secret, and that the company had earlier used pirated book sources before shifting to destructive scanning for "legal reasons" — the piracy issue separately led to a $1.5bn out-of-court settlement with authors. Judge William Alsup ruled in late July that training an LLM on copyrighted material does not itself constitute infringement, likening it to how a person learns to write by reading. The article's author argues this approach reflects a troubling trend of AI companies treating human-authored works as raw material to consume rather than engaging with the people who created them.

  • Anthropic secretly scanned and destroyed physical books to train Claude AI.
  • Court case revealed codenamed "Project Panama" scanning initiative.
  • Judge ruled AI training on copyrighted books isn't inherently infringement.

New here? Start with this

Anthropic, the company behind the Claude AI system, is at the centre of a US copyright lawsuit brought by authors, known as Bartz v Anthropic PBC. Court filings from the case describe an internal effort called "Project Panama," in which the company bought physical books, scanned them and discarded the originals so it could use the text to train its AI models. Anthropic had previously relied on books obtained through piracy, a separate issue that led to a $1.5bn settlement with authors.

The case matters because it tests how copyright law applies to AI training, an area with little established precedent. A US judge, William Alsup, has already ruled that training an AI on copyrighted material is not itself illegal, comparing it to a person learning by reading. That ruling leaves open broader questions about whether AI companies should be seeking permission from, or compensating, the authors whose work underpins their products.

More widely, the story feeds into an ongoing debate about how technology firms source the material used to build AI systems, and what obligations, if any, they have towards the writers, publishers and other creators whose work is involved.

AI Art Business Companies Culture Technology

Read the full article at the source →