The Justice Department takes a clear side in a defining AI copyright fight
The U.S. Department of Justice has entered one of the most consequential legal battles in artificial intelligence with a position that could influence how courts treat model training for years to come. In a filing tied to the consolidated lawsuit involving The New York Times and other rights holders, the department argued that training large language models on copyrighted text can qualify as fair use.
That intervention matters because the case has become a bellwether for the broader conflict between publishers and AI companies. The Times sued OpenAI and Microsoft in late 2023 in federal court in Manhattan, alleging that millions of its articles were used without permission to train systems including GPT-4 and to build products that compete with the paper as an information source. According to the supplied source text, the newspaper has sought damages in the billions and demanded the destruction of models trained on its work.
The Justice Department’s argument does not say copyright law is irrelevant to AI. Instead, it makes a narrower but highly important distinction: copying that occurs during training is not the same thing as what a model ultimately produces for users. That separation, in the department’s view, is central to the fair use analysis.
Why the training-versus-output distinction is so important
The supplied source text says the department argued that entire works may be copied during training, but that those works are not made publicly available in the process. It further argued that model outputs “often if not always” lack substantial similarity to the original materials. That point goes to one of the most disputed questions in the AI copyright debate: whether the use of source material during training should be judged as an internal analytical process, or as part of a commercial act that directly substitutes for the original works.
The Justice Department’s view favors the first framing. If the court accepts that logic, AI developers would gain a stronger legal basis for arguing that training on copyrighted corpora is transformative, especially when the resulting systems do not reproduce protected expression in a substantially similar way. If the court rejects it, developers could face much steeper licensing obligations and potentially major operational constraints.
The department also pushed back on what it characterized as an overly broad theory of market harm. As described in the source text, it argued that claims should not simply collapse training and output into a single use. That matters because copyright disputes often turn on whether a challenged activity merely learns from a work or whether it usurps the market for that work. The DOJ position suggests that plaintiffs cannot automatically prove injury just by showing their content was included in training data.
A legal analogy aimed at human creativity
One of the more striking features of the filing, as summarized in the source text, is its analogy to author Joan Didion. The department invoked the familiar practice of writers copying admired prose to learn craft, arguing that the law should recognize a difference between studying existing works and unlawfully republishing them. The filing reportedly suggested that under a more expansive anti-training theory, even that kind of human learning process could be cast as infringing once the learner later creates something new.
The analogy is designed to make a broader policy point. Copyright law is meant to protect original expression, but it also exists within a system that depends on learning, reference, and influence. The Justice Department appears to be arguing that training a model on text, by itself, belongs closer to the act of analysis than to the act of duplication for public distribution. Whether judges find that comparison persuasive is another matter, but it gives AI companies a narrative that reaches beyond technical detail and into first principles about creativity.
What the filing means for publishers, AI firms, and the market
For publishers and other rights holders, the DOJ filing is a setback in a case that many media organizations see as crucial to preserving control over their archives. The New York Times and other plaintiffs have argued that generative AI systems derive value from vast stores of professionally created work without paying for it. They also contend that AI products can compete with the original sources for attention, subscriptions, and advertising.
For AI companies, the department’s position provides more than rhetorical support. It offers a formal statement from the U.S. government that model training can produce creative and social value and that imposing liability for training alone could chill the very innovation copyright law is supposed to encourage. That framing is likely to be cited widely in future cases, negotiations, and policy debates, even beyond this particular lawsuit.
The source text also notes that the case is widely viewed as a landmark test for how courts will handle AI training. That description is not overstated. A ruling that narrows fair use in this context could alter the economics of frontier model development, especially for firms that rely on very large and diverse text datasets. A ruling that affirms broad latitude for training could entrench the current development model while shifting future disputes toward outputs, licensing deals, and product design choices.
The bigger issue is not settled
Even with the Justice Department now backing the fair use theory, the legal picture remains unsettled. The source text itself notes that others disagree. That disagreement reflects the basic fact that AI copyright law is still being built in real time through litigation, not through a mature and stable body of precedent. Courts still have to decide how traditional doctrines such as substantial similarity, transformation, and market substitution apply to systems that do not simply store and replay content, but learn statistical patterns from enormous datasets.
There is also a practical reason this filing matters beyond the courtroom. It could influence how companies, publishers, and lawmakers approach licensing. If training appears legally safer, AI firms may feel less pressure to license broad text archives on a compulsory basis. If publishers see the legal tide moving against them, they may focus more heavily on claims tied to specific outputs, brand confusion, or competitive harm from AI-generated summaries that resemble newsroom products.
What changed on September 2, 2026 is not the final legal answer, but the alignment of a major federal institution with one side of the argument. The Justice Department has now made explicit that, in its view, the act of training a large language model on copyrighted text can fit within fair use. In a dispute that will help define the relationship between AI systems and the written record they learn from, that is a significant development.
This article is based on reporting by The Decoder. Read the original article.
Originally published on the-decoder.com







