Legal status of AI training on copyrighted books remains unsettled

← Back to the feed

Legal status of AI training on copyrighted books remains unsettled

TechCrunch · 3 hours ago

Legal experts say the question of whether AI firms can legally train models on copyrighted books remains unsettled, despite a landmark ruling last year. Judge William Alsup ordered Anthropic to pay a $1.5 billion settlement to authors, but ruled that the AI training itself was lawful, comparing an AI's ingestion of text to a writer studying literature; Anthropic was instead penalised for sourcing the books from pirated shadow libraries. The case matters because it sets an early precedent for how courts weigh authors' rights against AI companies' use of their work, at a time when US copyright law has not been updated since 1976.

Attorneys Cathy Gellis and Jason Henderson told TechCrunch that the law remains inconsistent across cases, largely because judges are applying decades-old fair use principles to unprecedented questions about AI. Fair use permits copyrighted material to be used without permission if it is sufficiently "transformative", weighed against factors including the purpose of use and its effect on the market. Gellis noted the Anthropic ruling favours AI firms, given the settlement is small relative to the company's projected $200 billion annual revenue by 2028, while Henderson said courts tend to rule against AI companies when their tools are found to directly compete with the copyrighted works used to train them.

  • Courts remain divided on the legality of AI training on copyrighted books
  • Anthropic's AI training was ruled lawful; piracy of source books was not
  • Outdated 1976 copyright law struggles to address new AI questions

New here? Start with this

Anthropic, the company behind the AI model Claude, was recently ordered to pay $1.5 billion to a group of authors after being sued for using their books to train its AI systems. The judge in the case, William Alsup, made an important distinction: he ruled that training an AI on copyrighted text is not itself illegal, but Anthropic broke the law by obtaining the books from pirate websites rather than buying or licensing them properly.

This matters because AI companies increasingly rely on vast amounts of text, including books, articles and other creative work, to build their models, and authors and publishers have been pushing back over the use of their material without payment or permission. US copyright law was last substantially updated in 1976, long before AI existed, so courts are now having to apply old rules, particularly the concept of "fair use", to a technology nobody anticipated. Fair use allows copyrighted material to be used without permission in some circumstances, depending on factors such as how transformative the new use is and whether it harms the market for the original work.

Legal specialists say this single ruling has not settled the wider question of whether AI training on copyrighted material is generally lawful, since different judges are reaching different conclusions in similar cases. Because there is no updated law written specifically for AI, and because rulings so far have gone in different directions, the legal position facing AI companies, authors and publishers remains uncertain.

Both sides, in good faith

The strongest fair case each way — we don't pick a winner.

The case for

Those who favour treating AI training as lawful fair use argue that machine learning on text mirrors how human writers absorb influences from books they have read, extracting patterns and style rather than reproducing protected expression verbatim. They contend that requiring licences for every training text would make building large language models practically and financially impossible, stifling innovation and concentrating AI development in the hands of only the wealthiest firms. On this view, copyright law's purpose is to encourage the creation and dissemination of new works, and transformative uses like training, which produce a fundamentally different kind of output, serve that same public interest.

The case against

Those who believe authors deserve stronger protection argue that AI firms are commercially exploiting decades of creative labour without consent or fair compensation, often by sourcing works from pirated libraries precisely because obtaining licences would be costly or slow. They contend that when AI-generated text can substitute for the very books used to train it, this directly undermines authors' livelihoods and the market for their work, which is exactly the harm copyright law was designed to prevent. On this view, a company projected to earn $200 billion annually should not be permitted to build its product on unlicensed creative work merely because a settlement represents a small fraction of that revenue.

AI Technology

Read the full article at the source →

Originally published by TechCrunch as “Is it legal to train AI models on copyrighted books? It’s complicated”.