In a lawsuit that The New York Times filed against OpenAI, the Trump administration has contributed a 20-page brief in defense of the ChatGPT maker’s unlicensed use of copyrighted material to train its LLMs.
“The United States has a strong interest in continuing to develop a robust and competitive artificial intelligence industry that sets the standard for the practice and procedure of AI use globally… As such, it is critical for the United States to ‘retain global leadership in artificial intelligence,’” the brief reads, referencing an executive order that President Donald Trump signed last year.
The LLMs powering chatbots like ChatGPT, Claude, and Gemini are trained on incomprehensibly massive databases of published works, including copyrighted books, articles, and other media that AI companies feed into these databases without permission. Many publishers, including The New York Times in this case, have sought to argue that it is illegal for AI companies like OpenAI to train AI models on their copyrighted material.
This question — can you use copyrighted material to train an AI? — isn’t black and white, hence the extensive legal debate around the subject. These conversations often center on fair use, a carve out of copyright law that makes exceptions for certain scenarios when it can be ruled legal to use someone else’s copyrighted work without permission. In this case, the fair use debate addresses whether AI companies’ use of copyrighted work is “transformative” enough for a judge to rule it legal.
“Constraining LLM development under a misunderstanding of fair use doctrine would thwart such creative and scientific progress while hindering American prosperity and economic mobility,” the brief says.
So far, cases about AI training and copyright infringement have largely been favorable to AI companies. Last year, Judge William Alsup ordered Anthropic to pay a $1.5 billion copyright settlement to a group of writers whose works were used to train the company’s AI models; but Anthropic wasn’t dinged for its AI training. Rather, the company was fined for using illegal shadow libraries to pirate the books it used for training.
“Like any reader aspiring to be a writer, Anthropic’s LLMs trained upon works not to race ahead and replicate or supplant them — but to turn a hard corner and create something different,” Judge Alsup wrote, comparing the LLM’s training to a human reading a book.
This new Trump administration brief is not a ruling, as the case is being tried in the U.S. District Court for the Southern District of New York, and the authors of the brief do not have jurisdiction. However, this intervention by the Trump administration could still carry weight.
Facts Only
* The New York Times filed a lawsuit against OpenAI.
* The Trump administration contributed a 20-page brief in defense of OpenAI’s unlicensed use of copyrighted material for training LLMs.
* The brief references an executive order signed by President Donald Trump last year regarding global AI leadership.
* LLMs powering chatbots like ChatGPT, Claude, and Gemini are trained on databases including copyrighted books and articles without permission.
* Publishers argue that training AI models on copyrighted material is illegal.
* The legal debate centers on whether AI training of copyrighted material qualifies as "transformative" under fair use doctrine.
* A prior ruling ordered Anthropic to pay a $1.5 billion copyright settlement but did not penalize the AI training itself, focusing instead on the use of shadow libraries.
* One judge compared LLM training to a human reading a book, emphasizing creating something different rather than replication.
* The brief intervention is not a ruling, as the case is pending in the U.S. District Court for the Southern District of New York.
Executive Summary
The Trump administration submitted a 20-page brief in defense of OpenAI's use of copyrighted material to train its Large Language Models (LLMs) in a lawsuit filed by The New York Times. The brief references an executive order signed by President Donald Trump last year, asserting the United States' interest in maintaining global leadership in artificial intelligence development. The LLMs powering chatbots like ChatGPT, Claude, and Gemini are trained on massive databases containing copyrighted works without explicit permission from the creators. Legal debate surrounds whether this use constitutes fair use, specifically focusing on whether AI training is "transformative" enough to be legal.
The legal history surrounding these claims shows mixed outcomes; for instance, a prior ruling ordered Anthropic to pay a copyright settlement related to its training data, but did not rule against the training process itself. The administration’s brief argues that constraining LLM development based on a misunderstanding of fair use would impede creative and scientific progress and American economic mobility.
Full Take
The narrative centers on a fundamental tension between the pursuit of technological innovation and established intellectual property rights, framed through the lens of fair use doctrine. The core implication is that society must determine if the process of machine learning constitutes a permissible 'transformative' act analogous to human learning and creation. The intervention by the administration positions the defense of AI development not merely as a technical issue but as an economic necessity tied to national competitiveness, suggesting that stifling this progress harms broader prosperity.
The historical precedents cited—such as the ruling allowing Anthropic training while imposing separate fines for data acquisition methods—suggest a pattern where courts distinguish between the utilization of the raw material and the method of access or replication. This suggests that future legal frameworks may need to establish specific boundaries for derivative works in the context of massive-scale algorithmic ingestion. The argument that constraining AI development hinders progress suggests an assumption that innovation is inherently beneficial, which invites scrutiny regarding who bears the costs associated with that progress.
What is the true cost if the definition of "transformative" is set by the technology developer rather than established through settled law? What precedent does allowing massive data ingestion without permission set for future intellectual property relationships in an increasingly automated knowledge economy? How should the pursuit of global AI leadership balance against the rights of individual creators whose works form the bedrock of that knowledge?
Sentinel — Human
The text is a well-structured synthesis of a specific legal proceeding, successfully framing a complex copyright debate around AI training rather than presenting raw data.
