Tuesday, August 25, 2026

Independent technology reporting and practical analysis

Manila ·
ARTIFICIAL INTELLIGENCE

Independent reporting, useful context, and practical analysis.

Back to Technomalist
Artificial Intelligence / explainer

Can AI Companies Legally Train Models on Copyrighted Books?

Court decisions increasingly distinguish between training itself, how books were obtained and whether the resulting product competes with the original work.

Featured image for Can AI Companies Legally Train Models on Copyrighted Books?
Featured image for Can AI Companies Legally Train Models on Copyrighted Books?

There is no universal yes-or-no answer to whether an artificial intelligence company may train a model on copyrighted books. Recent cases show that courts examine how material was obtained, what the training is meant to accomplish and whether the resulting product harms the market for the original work.

Training and acquiring the dataset are separate questions

One major case involving Anthropic produced a mixed result. A federal judge found that using books to train a language model could qualify as fair use because the model was intended to create something different. But the court treated Anthropic's acquisition and storage of pirated copies as a separate copyright problem.

That distinction is important. A company might make a persuasive fair-use argument about the analytical process while still face liability for downloading books from unauthorized shadow libraries.

Anthropic later reached a $1.5 billion settlement covering hundreds of thousands of books. The settlement resolved claims over the pirated library, but it did not create a blanket rule permitting every form of AI training.

Fair use depends on context

U.S. courts weigh four factors in a fair-use analysis: the purpose and character of the use, the nature of the copyrighted work, how much was copied and the effect on the market for the original.

A separate dispute involving Thomson Reuters and Ross Intelligence went the other way. The court found that copying legal research content to build a competing product was not sufficiently transformative. Direct competition and market harm can therefore weaken an AI developer's fair-use defense.

The U.S. Copyright Office has similarly warned that some training uses may qualify as fair use while others may not. The source of the training material, the purpose of the model and the behavior of its outputs all matter.

What remains unsettled

Dozens of lawsuits continue across multiple jurisdictions, and different countries apply different copyright exceptions. Licensing agreements are also evolving faster than legislation.

The safest conclusion is that AI training is not automatically legal simply because it is called training—and it is not automatically infringement simply because copyrighted material was analyzed. Each system depends on facts that courts are still working through. This article is a general explanation, not legal advice.

See an error? Read our corrections policy or email [email protected].

MORE FROM TECHNOMALIST

Continue reading

View all
Telecommunications network equipment representing online safety controls
Internet

PLDT Moves to Block Over 100 Sites in Child Online-Safety Push

Remote business meeting shown across multiple computer screens
Software

Microsoft Teams Adds Automatic Blocking for External Meeting Bots

Identity security dashboard representing a blocked cyberattack
Cybersecurity

ReliaQuest Says Device Trust Stopped ShinyHunters Data Theft