Connect with us

NEWS

Washington Puts National Security Into the OpenAI Copyright Fight

The Justice Department told a New York court that AI training is fair use, casting U.S. security and smaller labs as the parties at risk.

Published

on

The Justice Department told Judge Sidney H. Stein on September 1, 2026 that training large language models on copyrighted text is fair use. The United States is not a party in the Manhattan multidistrict suit against OpenAI and Microsoft, yet it filed anyway under 28 U.S.C. § 517, which lets the department speak in a federal case without the court’s leave.

The caption still reads as publishers versus labs. The brief spends its opening pages on a different set of interests: U.S. AI capacity, smaller developers that cannot buy blanket licenses, and foreign rivals that would not be bound by a narrow American fair-use rule.

A 20-Page Brief, Filed Without Leave of Court

The statement of interest filed September 1 sits in In re OpenAI, Inc. Copyright Infringement Litigation, MDL No. 1:25-md-03143, which gathered The New York Times’s December 2023 suit with author and publisher cases before Stein in the Southern District of New York. Associate Attorney General Stanley E. Woodward Jr. and Assistant Attorney General Brett Shumate signed the 20-page paper, joined by Senior Counsel Michael Weisbuch.

The government limited the question it wanted answered. It asked whether copying works at the training stage, “in order to feed data into the model as learning material,” is fair use. Collection of the data and later chatbot answers, it said, can raise separate copyright questions and should not be collapsed into one use.

That split is the whole move. If Stein treats training copies as their own use, a handful of memorized answers cannot be used to enjoin the model that produced them. Footnote 15 adds that any forward-looking remedy would have to stop at complete relief for the plaintiffs, not “massive liability for LLM output uses generally, let alone distinct LLM training uses.”

The Brief Opens on National Security, Not Newsrooms

Section I is not a quiet recitation of § 107. It quotes the January 23, 2025 executive order on American AI leadership and a June 2, 2026 order on AI innovation and security, then a Government Accountability Office warning dated April 19, 2022 that a failure to integrate AI could hinder national security tasks from intelligence analysis to targeting.

Rules of law that make it significantly more difficult to develop a robust AI industry in the United States therefore threaten national security and give a competitive advantage to foreign adversaries who are not so encumbered.

Statement of Interest of the United States, In re OpenAI, Inc. Copyright Infringement Litigation

The March 2026 National Policy Framework, as quoted in the brief, already stated the policy line the department now wants in a court order: the “training of AI models on copyrighted material, in and of itself, does not violate copyright laws.” The same framework, the brief adds, still wants American creators protected from infringing outputs. Training is the use Washington wants locked down. Outputs are the use it is willing to keep litigating.

The competition argument runs the other way from the Times’s attack. The Times’s spokesman, Graham James, said the administration “is siding with a handful of trillion-dollar AI companies at the expense of the countless American creators whose work they stole.” The brief claims a compulsory training license would do the opposite, because “only the largest technology companies might have the capital necessary to pay licensing fees,” with those fees flowing to “legacy media outlets” as “large subsidies.” Independent publishers, it says, already use models themselves, including the Times.

Anupam Chander, a Georgetown law and technology professor, read the paper the next day and wrote that it “strongly argues that LLM reading of copyrighted works for training is fair use” and “rejects Copyright Office’s pre-publication theory of ‘market dilution’ for AI outputs.” The objection on the other side is simpler: if ingesting entire works to build a commercial model is fair use, the exclusive right to copy has been defined out of this industry.

Why Training Copies Are Not ChatGPT Answers

Fair use under 17 U.S.C. § 107 is a four-factor test. The department leans on factor one, purpose and character, and factor four, market effect. It treats factors two and three as supporting fair use because training, by itself, does not put the copied work in front of the public as a competing substitute.

On factor one it borrows the Northern District of California’s 2025 language. Judge William Alsup called Anthropic’s training use “transformative, spectacularly so.” Judge Vince Chhabria called Meta’s training “highly transformative.” The department goes further and calls LLM training “extraordinarily transformative,” because the copy is made to teach statistical relationships among words, not to republish an article for a reader.

The analogy is Google LLC v. Oracle America, Inc., 593 U.S. 1 (2021), where Google copied Java declaring code so programmers could write new programs in a new smartphone environment. The Supreme Court held that this copying enabled new expression. The department says teaching a model to “recognize relationships between data and adapt to new information” is an even cleaner case, because the training copy is not offered to the public at all.

On factor four it takes the Second Circuit’s Google Books test from Authors Guild v. Google, Inc., 804 F.3d 202 (2015): the harm that counts is “significant substitutive competition,” not “some loss of sales.” “Using a copyrighted work to train an LLM, without more, generally does not result in this sort of substitution because it does not ‘reveal’ a significant amount of original ‘authorial expression.’” Outputs that do not copy protected expression, it argues, are in the same genre, and a genre is not copyrightable.

WHAT THE STATEMENT LEAVES OPEN

  • Acquisition: The brief analyzes the training copy as learning material and does not decide whether a pirate download used to build a lasting library is a separate infringing use.
  • Outputs: Reconstructing and disseminating an original work “may not be transformative,” and those uses are to be judged one by one.
  • Licenses already in the market: The United States “takes no position” on whether a full licensing regime would be financially or logistically feasible, and it notes that publishers already sell specialized access to paywalled and real-time content.

That last point is the hinge for rightsholders. Even if Stein adopts the training holding, a lab that stored a pirate corpus, or a chatbot that emits a close copy, is still in the case. The department wants those facts kept off the training verdict.

Anthropic Already Paid $1.5 Billion for Piracy

Two Northern District of California orders from June 2025 are the nearest merits rulings, and the department picks a side. It follows Alsup on transformativeness and attacks Chhabria on market dilution.

HOW THREE FAIR-USE ANALYSES SPLIT

Decision Training copies How copies were obtained Market-harm theory
Bartz v. Anthropic, June 23, 2025 Fair use; “transformative, spectacularly so” Purchased books digitized for a library: fair use. Pirate central library: not fair use Training did not deliver copies or knockoffs to the public
Kadrey v. Meta, June 25, 2025 Fair use on that record Downloads treated as part of one training process Dilution “likely” to win with proof; plaintiffs had none
DOJ statement, September 1, 2026 “Extraordinarily transformative” fair use Not analyzed Rejects dilution; only substitutive revelation of expression counts

Alsup’s piracy holding is the part the department does not touch, and it is the part that has already produced money. He found that Anthropic downloaded more than 7 million pirated books, kept them in a general-purpose library, and could not wrap that library in the training fair-use ruling. “Piracy of otherwise available copies is inherently, irredeemably infringing even if the pirated copies are immediately used for the transformative use and immediately discarded,” he wrote.

Anthropic settled rather than try that claim to a jury. On July 20, 2026, Judge Araceli Martínez-Olguín granted final approval of the $1.5 billion settlement, about $3,000 per work across more than 482,000 books, with Anthropic general counsel Aparna Sridhar saying more than 91 percent of covered authors and publishers had claimed a share. The per-work figure is four times the $750 statutory minimum for ordinary infringement. Class members released past input claims through August 25, 2025. Output claims and future conduct were not released, and Anthropic must destroy the LibGen and PiLiMi files.

Sridhar said the company settled after the ruling that training on books is fair use, “which remains the law today.” That is the map the department is drawing for Stein: training can be fair use and a lab can still write a nine-figure check for the library it built on the way in.

No Deference for the Register’s Training Report

The brief’s sharpest institutional target is not the Times. It is the government’s own copyright adviser. In a footnote on Kadrey, it says Register of Copyrights Shira Perlmutter, “who is currently challenging her removal, appeared to endorse a similar theory” of market dilution in the Office’s May 2025 generative AI training report. “Her understanding does not warrant deference,” it says, citing Loper Bright Enterprises v. Raimondo, 603 U.S. 369 (2024). It calls the Register’s reasoning “threadbare” for skipping the use-by-use case law on what market effect counts.

Perlmutter released that pre-publication Part 3 on May 9, 2025. The next day she received an email saying her position was “terminated effective immediately.” She sued, arguing that only the Librarian of Congress may remove the Register. A divided D.C. Circuit panel on September 10, 2025 held the firing likely unlawful because the Office sits in the legislative branch. On June 30, 2026, the Supreme Court left that restoration order in place and said the denial was not a ruling on the merits. The September 1 brief still describes her as challenging the removal.

Part 3 did not say all training is infringement. After more than 10,000 comments on an August 2023 notice of inquiry, the Office said training acts often implicate exclusive rights and “likely qualify as fair use in some circumstances but not in others,” with facts including the source of the works, the purpose, and any guardrails. It said licensing markets were emerging in some sectors and should be given time before Congress steps in. That is a case-by-case frame. The department wants a categorical one for the training step.

THE ROAD INTO THE MANHATTAN DOCKET

  1. December 2023: The New York Times sues OpenAI and Microsoft in the Southern District of New York.
  2. April 11, 2025: The OpenAI copyright cases are centralized as MDL 25-md-3143 before Judge Stein.
  3. May 9, 2025: The Copyright Office releases Part 3 of its AI study, on generative training.
  4. May 10, 2025: Perlmutter is told she is removed as Register.
  5. June 23, 2025: Alsup holds Anthropic’s training use is fair use and its pirate library is not.
  6. June 25, 2025: Chhabria holds Meta’s training is fair use on a thin market record and warns dilution could flip later cases.
  7. July 20, 2026: Martínez-Olguín grants final approval of Anthropic’s $1.5 billion piracy settlement.
  8. September 1, 2026: The United States files the statement of interest.
  9. September 4, 2026: OpenAI, Microsoft, and the publishers file opening summary judgment briefs.

The Office’s broader AI docket is still public on its Copyright Office artificial intelligence reports page, including Part 1 on digital replicas (July 31, 2024) and Part 2 on the copyrightability of AI-assisted works (January 29, 2025). The department is telling Stein he does not have to follow Part 3’s dilution discussion when he weighs factor four.

Contracts, the DMCA, and a Trip to Congress

If Stein accepts the training holding, the remaining pressure points are the ones the brief either flags or ignores. Publishers can still sell licenses for live, paywalled, or otherwise specialized feeds, a market the department expressly preserves. They can still sue on outputs that are substantially similar to protected expression. They can still try acquisition claims of the Bartz library type. And they can still plead Digital Millennium Copyright Act claims tied to copyright-management information stripped from files, which do not rise or fall with a fair-use finding on the training copy.

Congress is the other forum the brief points to. The National Policy Framework, as described in the filing, asked lawmakers to consider licensing or collective-rights systems without deciding in the statute “when or whether such licensing is required,” plus a federal regime for unauthorized digital replicas of a person’s identifiable attributes. Part 1 of the Office’s study had already recommended a federal replica law. That is not a verdict on ChatGPT’s training set. It is a map of where a training-stage loss in court still leaves work to do.

The Joan Didion aside in the brief is the department’s answer to dilution. Didion typed out Hemingway stories as a teenager “to learn how the sentences worked.” The United States says a model doing the same thing, without handing the public the story, is in that tradition, not in the tradition of a competing edition. Rightsholders will say a commercial system that can emit unlimited substitute text is not a teenager at a typewriter. Stein has to pick which analogy fits the record now in front of him.

Judge Stein Now Has the Summary Judgment Record

The statement of interest is not an order. Stein still decides fair use on the evidence the parties filed. Opening summary judgment and related papers in the news and books tracks went in on September 4, 2026, three days after the government spoke. Microsoft asked him to hold that training a general-purpose model is fair use as a matter of law. The publishers asked him to find liability for copying at several stages and to reject the fair-use defense.

A training-only ruling would track Bartz and the department and still leave output, piracy, and DMCA counts alive, which is already how the Anthropic case ended. A ruling that folds outputs and training together, in the Kadrey dilution style, is the outcome the brief is written to block. Either way, the United States has now told a trial judge that the copy made to teach the model, standing alone, should walk.

Frequently Asked Questions

What Is a Statement of Interest Under 28 U.S.C. § 517?

Section 517 lets Department of Justice lawyers “attend to the interests of the United States in a suit pending in a court of the United States.” The brief’s own footnote says the statute has no time limit and does not require the court’s leave, citing district-court orders that have read it that way, including Gil v. Winn Dixie Stores. The filing is not binding on Judge Stein and does not make the United States a party.

Does the Justice Department Say ChatGPT Answers Are Fair Use?

No. The March 2026 National Policy Framework quoted in the brief says American creators “should be protected from AI-generated outputs that infringe their protected content, without undermining lawful innovation and free expression.” The department argues only that those output questions cannot dictate the legal treatment of the separate training copies.

What Did the Bartz Court Hold Besides Training Fair Use?

Alsup also held that turning lawfully purchased print books into digital files for Anthropic’s own central library was fair use, because the company was replacing copies it had bought with searchable substitutes it did not redistribute. That format-shift holding is separate from both the training ruling and the pirate-library ruling, and the department does not ask Stein to adopt it.

Who Appoints the Register of Copyrights, and Why Does That Matter Here?

The Librarian of Congress appoints the Register under 17 U.S.C. § 701(a). Perlmutter has held the post since October 2020. The department’s refusal to give her Part 3 report Loper Bright deference is a claim that a legislative-branch advisory study does not control a district court’s fair-use analysis, even while her removal case is still pending.

Disclaimer: This article is news reporting and analysis of a court filing and related public orders. It is informational only and is not legal advice, is not an opinion on the merits of any pending claim, and is not a prediction of how Judge Stein or any appellate court will rule. Readers with rights, licenses, or exposure in AI-training disputes should consult a qualified copyright attorney before changing contracts, preservation practices, or litigation strategy. Figures, docket events, and office-holder statuses reflect the public papers described here and may change as the MDL and the Register’s removal case move.

Harry is the editor of THE iBULLETIN, an independent publication he owns and runs. He has been in journalism for ten years, first reporting and later editing, and much of what the site covers now begins in its inbox. Reader mail is read in full, every message of it. A tip is treated as a lead to be verified, not a story to be printed, and a challenge to a published fact is checked against the original filing, statement or transcript within the day, with the article corrected under a public policy if the reader is right. Questions that several readers ask become articles. That exchange feeds coverage of news, business and technology, of science and sports, and of entertainment, lifestyle, travel, auto and gaming, written for readers spread across many countries rather than one. Harry works from primary sources and checks each number himself before publication, and he would rather run a shorter story than an unconfirmed one. The address for all of it, tips, corrections and questions alike, is support@theibulletin.com.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending