top of page

Anthropic’s $1.5 Billion Copyright Settlement: What the Case Means for AI Training and Fair Use

  • Writer: panagos kennedy
    panagos kennedy
  • 6 days ago
  • 6 min read

Anthropic, the developer of the Claude artificial intelligence models, has received final approval of a $1.5 billion copyright settlement arising from its acquisition of hundreds of thousands of copyrighted books.



The size of the settlement in Bartz v. Anthropic PBC makes the case remarkable. But for businesses developing or using artificial intelligence, the more important issue is why Anthropic faced that liability even after winning a significant ruling that its use of copyrighted books to train Claude constituted fair use.


The court drew a distinction with potentially broad consequences for AI copyright litigation:


Using copyrighted material to train an AI model may qualify as fair use under some circumstances. Acquiring copyrighted material unlawfully is a different question.

That distinction makes Bartz one of the most useful AI copyright decisions to date for companies evaluating training data, third-party datasets, licensing practices, and other sources of information used in artificial intelligence systems.


What Happened in the Anthropic Copyright Lawsuit?

Authors Andrea Bartz, Charles Graeber, and Kirk Wallace Johnson sued Anthropic, alleging copyright infringement arising from copies of their books that Anthropic had acquired and used in connection with the development of its Claude large language models.


Anthropic had obtained books through more than one route.


It purchased physical books, removed their bindings, scanned them, and retained digital copies. It also downloaded millions of books from online pirate libraries as it assembled a large central digital library.


The distinction between those sources eventually became critical.


In June 2025, U.S. District Judge William Alsup ruled on Anthropic's fair-use defense. The court separated three different types of copying: copies used to train Anthropic's large language models, digital replacements made from books Anthropic had purchased, and books obtained from pirate libraries.


The results were very different.


Was Anthropic’s Use of Books to Train Claude Fair Use?

On the facts before it, the court said yes.


The court granted summary judgment to Anthropic on its use of copies of the authors' books to train Claude and its predecessor models. It regarded that use as highly transformative for purposes of the Copyright Act's fair-use analysis.


The ruling therefore represents an important victory for the argument that using copyrighted works as inputs in large-language-model training can constitute fair use.

The court also reached a favorable result for Anthropic regarding physical books that it had lawfully purchased and then digitized. Anthropic destroyed the physical copy as part of the scanning process and retained the digital version as a replacement. The court treated that one-for-one format conversion as fair use for reasons separate from the AI-training analysis.


But neither holding protected Anthropic's pirate-library downloads.


Why Did Anthropic Still Face Copyright Liability?

Before purchasing many books, Anthropic had downloaded more than seven million books from pirate sources and retained them in a central digital library.

The court rejected Anthropic's attempt to treat those copies as protected simply because some of them might later be used for a transformative purpose such as AI training.


That is the central lesson of the case.


The court treated acquiring and maintaining the pirated library as a separate act from using particular copies to train an AI model. Anthropic had retained some of those works even after deciding they would not be used for training.


For those pirated library copies, the fair-use analysis came out differently. The court denied Anthropic summary judgment and contemplated a trial addressing the pirated copies and resulting actual or statutory damages, including potential willfulness.

The court also made clear that subsequently buying a legitimate copy would not necessarily eliminate liability associated with an earlier pirated copy, although it could affect damages.


Instead of proceeding to that trial, the parties settled.


How Large Is the Anthropic Copyright Settlement?

On July 20, 2026, the U.S. District Court for the Northern District of California entered final approval of the class-action settlement and dismissed the case with prejudice.

The settlement provides $1.5 billion in monetary relief and covers a definitive list of 482,460 copyrighted works.


The financial magnitude explains why the case has received so much attention. But businesses should be careful about drawing the wrong lesson from the settlement.

Anthropic did not pay $1.5 billion because a court determined that AI training, standing alone, infringes copyright.


In fact, Anthropic had prevailed on that part of the fair-use dispute.


The claims creating the remaining exposure concerned how copyrighted material had been acquired and retained.


What Does Bartz v. Anthropic Say About AI Training and Fair Use?

It would be tempting to reduce the decision to either of two propositions:

“AI training is fair use.” Or: “Using copyrighted material for AI creates massive copyright liability.”


Neither accurately captures the case.


The more precise conclusion is that copyright analysis can depend upon the particular act of copying being challenged. The use of a copyrighted work as training material may present one fair-use question. Making or obtaining the copy that enters the training-data pipeline may present another. Maintaining that material afterward for additional uses may present still another.


That distinction is particularly important because modern AI systems often depend upon complicated data supply chains.


A model developer may obtain information from a commercial data vendor. A business may license an AI platform developed by another company. Developers may collect material from public websites, repositories, customers, contractors, employees, or third-party datasets.


The fact that the eventual use of that information may be transformative does not necessarily answer whether the company had the right to make or acquire the underlying copy.


Why Data Provenance Is Becoming an AI Copyright Issue

For businesses, one of the most important concepts emerging from the Anthropic copyright case is data provenance: knowing where information came from and being able to document the legal basis on which it was obtained.


Companies developing or acquiring AI technology should increasingly consider questions such as:


  • Where did the training or reference data originate?

  • Was it purchased, licensed, scraped, downloaded, supplied by a vendor, or collected internally?

  • What rights did the source have to provide it?

  • What restrictions accompany the data?

  • What representations and warranties did a vendor provide?

  • Does the agreement allocate copyright risk or provide indemnification?

  • Can the company reconstruct the provenance of the dataset if challenged several years later?

  • Is information retained after the purpose for which it was originally acquired has ended?


Those questions are not merely technical or operational issues. Bartz demonstrates how they can become central questions in copyright litigation.


What Does the Anthropic Settlement Mean for Companies Using AI?

The immediate implications extend beyond companies building foundational AI models.

A business incorporating artificial intelligence into products or services may receive datasets from vendors, acquire pretrained models, commission custom AI systems, or integrate third-party AI tools into existing products.


In each situation, copyright diligence should address more than whether the finished technology produces infringing output.


Companies should also consider the inputs.


That may include reviewing dataset licenses, understanding collection methods, negotiating appropriate representations and warranties, allocating infringement risk in vendor agreements, documenting internally created datasets, and establishing retention practices for copyrighted material.


Similar questions can arise during acquisitions and investments involving AI companies. A buyer conducting intellectual property diligence may need to understand not merely who owns a company's software and models, but how the underlying training data was assembled.


A technically impressive AI asset may carry substantial undisclosed copyright exposure if its provenance cannot be established.


Does Bartz Mean All AI Training Is Fair Use?

No. The Anthropic decision is significant, but it does not establish a nationwide rule that all AI training involving copyrighted works constitutes fair use.


Fair use under Section 107 of the Copyright Act is a fact-specific inquiry. Different copyrighted works, methods of acquisition, model architectures, uses, outputs, licensing markets, or evidence of market harm may produce different results.


The Bartz ruling was also a federal district court decision rather than a controlling Supreme Court or appellate decision.


And because the claims involving Anthropic's pirated library were settled, there will be no trial resolving damages on those claims in this case and no appellate decision reviewing that dispute.


The settlement therefore leaves important AI copyright questions unresolved.


The Larger Lesson

The most important lesson from Bartz v. Anthropic may be narrower—and more useful—than a broad pronouncement about whether artificial intelligence and copyright law are compatible.


A potentially lawful downstream use does not necessarily cure an unlawful upstream act of copying.


For companies working with AI, that means copyright risk analysis should begin before information reaches the model.


Companies developing, purchasing, licensing, or investing in artificial intelligence systems should understand where important datasets came from, what rights accompany them, how those rights are documented, and whether information is later being used or retained beyond its original purpose.


The Anthropic case turned the distinction between use and acquisition into a $1.5 billion issue. For businesses working with copyrighted data, that is a distinction worth understanding.


Comments


© 2026 Panagos Kennedy PLLC. All Rights Reserved. | Disclaimer

bottom of page