← Back to StoryBreak

StoryBreak

Technology

Is it legal to train AI models on copyrighted books? The answer remains unsettled

Recent court rulings have offered partial guidance on whether AI companies may use copyrighted books for training, but the legality still depends on how the books were obtained, how they were used and whether the resulting system harms the market for the originals.

Published 8/24/2026, 12:54:08 AM

Detailed view of a server rack with a focus on technology and data storage.
Photo by panumas nikhomkhai / Pexels
The legal question surrounding AI training on copyrighted books remains unresolved in the United States, despite a series of major court rulings and settlements that have begun to define the boundaries. At the center of the dispute is fair use, a doctrine that can permit limited use of copyrighted material without the copyright holder’s permission. Courts typically weigh the purpose of the use, the nature of the work, the amount copied and the effect on the market for the original. AI companies argue that training is a transformative process. Rather than publishing a digital copy of a book, they say, a language model analyzes patterns in large collections of text and uses what it learns to generate new responses. Technology companies have cited that argument in defending their use of books, articles and other copyrighted works. Authors and publishers counter that building a commercial model can require making complete copies of their books, often without payment or permission. They also argue that AI-generated writing, summaries and competing products could reduce demand for human-authored work and weaken the market for future licensing. A closely watched 2025 ruling in the Northern District of California offered a split result in a case brought by authors against Anthropic, the company behind Claude. The court found that using books to train large language models could qualify as fair use in the circumstances before it. But the court separately concluded that Anthropic’s acquisition and retention of millions of pirated books was not protected by fair use. That distinction has become one of the most important points in the debate. A company may have a stronger legal position when it lawfully acquires books, even if it makes complete digital copies for a transformative purpose. Obtaining the same material from pirate libraries can create a separate infringement problem, regardless of what happens during model training. The Anthropic litigation later produced a $1.5 billion settlement covering more than 482,000 books. A federal judge approved the settlement in July 2026, with authors and publishers receiving payments for works connected to the case. The settlement resolved claims against Anthropic, but it did not establish a universal rule for every AI model or every copyrighted book. Other cases remain active. In May, five publishing companies and author Scott Turow sued Meta in federal court in Manhattan, alleging that the company used millions of copyrighted books and journal articles to train its Llama system. Meta has said that training on copyrighted material can qualify as fair use and vowed to defend itself. The U.S. Copyright Office is also studying the issue. Its artificial-intelligence initiative has examined the use of copyrighted materials in training and published a pre-publication version of a report on generative-AI training in May 2025. The office has said a final version is expected, leaving policymakers and courts to continue working through questions that existing copyright law was not written specifically to answer. For now, there is no simple rule that training on copyrighted books is always legal—or always illegal. The outcome may turn on whether the material was lawfully obtained, whether the copies were retained, the design and purpose of the model, the nature of its outputs and the effect on authors’ and publishers’ markets. That uncertainty is likely to persist until more courts issue decisions on the merits or Congress creates a system specifically governing AI training licenses and compensation.