During the early years of generative artificial intelligence, technology companies moved faster than the laws designed to regulate them. Models were trained on extraordinary quantities of books, news articles, images, recordings and other works, often without their authors knowing they had been incorporated into those databases.

The judicial approval of the $1.5 billion settlement between Anthropic and thousands of authors represents the first major payout by an artificial intelligence company for the use of pirated books. The company will be required to pay approximately $3,000 for each of the more than 482,000 works included in the case. Around 91% of the books have already been claimed by writers or publishers.

The settlement is extraordinary in its scale, but its importance extends beyond the money — which for Anthropic is not particularly significant either. The case begins to separate two questions that until now have often been conflated: whether a company may use a protected work to train a system, and whether it may obtain that work through a pirated copy.

The provisional answer from American courts is controversial. It holds that training a model on copyright-protected works may be considered "legitimate and transformative use." The piracy used to build the training library, however, may give rise to liability running into the hundreds of millions.

A victory and a defeat in the same ruling

The Anthropic case produced a split outcome. Federal judge William Alsup determined that training the Claude models on legally acquired books could be protected under the fair use doctrine. He found that the system did not function simply as a repository of texts, but rather extracted statistical relationships to produce new capabilities.

However, the judge also concluded that the company could not invoke that defense to justify the downloading and retention of millions of books obtained from pirate sites.

The subsequent approval of the settlement closed off the immediate threat of a trial covering hundreds of thousands of works, but did not establish a binding standard for the rest of the country. Nor did it determine that all training on protected material is legal. What it did was offer an initial guide: the provenance of data may prove just as important as the subsequent use made of it.

Large models require enormous volumes of quality text. Books are especially valuable because they contain extended narratives, complex arguments, specialized information and more carefully crafted linguistic structures than much of the content freely available on the internet.

But assembling them through licenses can be slow, costly and administratively complex. So-called "shadow libraries," such as LibGen, which contain millions of downloadable works, offered an immediate alternative for years. The Anthropic settlement reveals the potential price of that decision.

The end of the "build first, negotiate later" era

The initial development of generative artificial intelligence followed a logic well known in Silicon Valley: launch products, gain scale and resolve regulatory problems later.

That method worked because companies could claim that intellectual property laws had been drafted before large language models existed. They could also point out that a system does not necessarily retain a readable copy of each work, but rather learns mathematical patterns from them.

In 2025, two federal judges in California found that the training of Anthropic's and Meta's models could be transformative. But the decisions did not grant blanket immunity. In Meta's case, Judge Vince Chhabria noted that the authors had not presented sufficient evidence of market harm, though he warned that other plaintiffs could build a stronger case. That nuance opened the door to a second generation of lawsuits.

The first cases focused on proving that companies had copied works. The new litigation attempts to prove something more difficult and potentially more dangerous for the industry: that artificial intelligence products directly compete with the authors, media outlets, and publishers whose works they used.

The question that may define the next cases

U.S. fair use doctrine analyzes several factors, including the purpose of the use, the nature of the work, the amount copied, and the effect on the potential market—criteria established in Section 107 of the U.S. Copyright Act.

For years, artificial intelligence companies focused on the first factor. They argued that training is transformative because a novel, an article, or a photograph is not displayed to the user in the same way it was incorporated into the system.

Rights holders are shifting the discussion toward the last factor: economic harm.

The lawsuit filed by major publishers against Google argues that Gemini can generate in minutes texts that compete with the books used to train it. The plaintiffs claim that the company copied millions of works originally submitted for limited services, such as Google Books or Google Scholar, and then repurposed them to develop its commercial models.

But proving it is not so simple. A copy downloaded from an illegal library leaves identifiable traces, yes, but market harm, on the other hand, requires demonstrating how a technology alters sales, subscriptions, licenses, jobs, or incentives to create new works.

From books to news

Media lawsuits could become one of the most important tests. Unlike many books published years ago, journalistic content is part of an immediate market. News outlets charge subscriptions, sell advertising, and license archives. A response generated by a search engine or an assistant can satisfy a user's query without them ever visiting the original page.

Journalism organizations also began turning to the courts. According to a survey cited in the Tow Center for Digital Journalism's document, outlets such as The New York Times, The Intercept, the Center for Investigative Reporting, and Ziff Davis filed lawsuits against artificial intelligence companies for using their content to train models without authorization. The cases argue that this practice results in losses of traffic, advertising, and subscriptions.

The New York Times case is particularly relevant because it challenges both the information used to train the models and the responses they produce.

A New York court has already allowed authors to proceed with claims related to outputs generated by ChatGPT. Judge Sidney Stein held that the plaintiffs could potentially demonstrate that some responses are sufficiently similar to their works as to constitute infringement.

The emergence of a licensing economy

Nearly 20 media groups have signed licensing agreements with OpenAI. Perplexity has established agreements with news companies and created a $42.5 million fund linked to its subscription system. Microsoft is preparing a marketplace in which publishers could receive payments when their content is used by artificial intelligence products.

These negotiations suggest that the question is no longer simply whether companies must pay, but how value will be calculated. A model could pay per work incorporated, per query that uses certain information, per source appearing in a response, or through a flat fee. Each method benefits different players.

Large publishers and national media outlets have valuable catalogs and bargaining power. Independent authors, photographers, illustrators and small publications have less power to demand compensation. For this reason, the outcome of the major cases could affect even those who never set foot in a courtroom.

Google, Meta, Microsoft, Amazon and OpenAI can absorb legal costs, sign agreements with major rights holders and acquire entire databases. A startup would struggle to pay for millions of books, articles, images and songs before training its first model.

Many questions remain open — and in the United States experts say that some will inevitably end up before the Supreme Court when the time comes — but the agreement with Anthropic marks the beginning of the first major reckoning for American artificial intelligence: the moment when the companies that built systems using much of the internet's cultural output must begin to explain where they obtained their materials, what they did with them and how much they are willing to pay for the knowledge that helped create their products.