Almost every publisher AI strategy written since 2023 rests on an assumption nobody states out loud: that the courts will eventually make the models pay, and that a licensing cheque is therefore a matter of timing rather than a matter of whether. That assumption took its most serious hit yet last week. The United States government, appearing as a non-party in the consolidated litigation against OpenAI and Microsoft, filed a nineteen-page brief arguing that training on copyrighted written works is not infringement at all. The practical consequence for a publisher is not that licensing disappears. It is that licensing stops being something the law might compel and becomes something a buyer chooses, which means the only leverage left is whatever a publisher physically controls: the live retrieval that happens on its own servers, in real time, long after training finished.
What the Department of Justice filed, and when
On 1 September 2026 the United States filed a Statement of Interest under 28 U.S.C. section 517 in In re OpenAI Copyright Infringement Litigation, the multidistrict proceeding carrying case number 25-md-3143 before Judge Sidney H. Stein in the Southern District of New York, with Magistrate Judge Ona T. Wang handling discovery. A Statement of Interest is advisory. The government is not a party and Judge Stein is not bound by it. It is nonetheless an unusual intervention by the executive branch in a private copyright dispute between named commercial parties, and it landed three days before both sides filed cross-motions for summary judgment on 4 September.
The plaintiffs are not one publisher. The News plaintiffs include The New York Times Company, which asserts 6,030,928 individual articles; the Daily News group, comprising the New York Daily News, Chicago Tribune, Orlando Sentinel, Sun-Sentinel, Denver Post, Pioneer Press, Orange County Register and Mercury News; Ziff Davis, whose properties include IGN, CNET, ZDNET, PCMag, Everyday Health and Mashable; the Center for Investigative Reporting, publisher of Mother Jones and Reveal; and The Intercept. A separate Class plaintiff group covers book authors and the Authors Guild. OpenAI and Microsoft are the defendants. On the same day briefing opened, the Seattle Times filed its own suit against both companies.
The argument: training is transformative, dilution is not cognisable harm
The government's brief concentrates on the first and fourth fair use factors. On the first, it argues that training is "exceedingly transformative" (p. 10) and later "extraordinarily transformative" (p. 12), because a model does not consume an article the way a reader does. It converts text into numerical representations and learns statistical relationships across vocabulary, syntax and knowledge. The brief leans on Bartz v. Anthropic PBC, where Judge Alsup called AI training "transformative - spectacularly so", on Kadrey v. Meta Platforms, and on Google v. Oracle and the Second Circuit's Authors Guild v. Google, for the proposition that copying can be fair when it enables a new technological function even where whole works are reproduced at an intermediate step.
On the fourth factor, the government attacks the market dilution theory directly. Dilution is the argument that a flood of machine-generated material occupying the same genre destroys the market for human work even where nothing is copied verbatim. The DOJ calls it "deeply flawed" (p. 15) on two grounds: that under Warhol each use must be analysed separately, so outputs cannot be folded into an assessment of training; and that output which is not substantially similar to a protected work is not a copyright substitute merely because it competes in the same category. The brief singles out the Kadrey court's fourth-factor analysis for combining training and output into one continuous use, and offers the illustration that Joan Didion retyped Hemingway's stories to study his sentences without thereby owing him a share of her later career.
What the filing deliberately did not decide
This is the part most trade coverage skipped, and it matters commercially. The government confined itself to the use of works during training. It expressly did not address the acquisition and storage of training data, and it did not address outputs that reproduce protected expression. It also did not contend that the government had authorised any of the challenged conduct.
So the carve-outs are real. How a dataset was obtained remains live. Whether specific outputs infringe remains live. And nothing in the brief touches retrieval: the separate act of an assistant fetching a live page from a publisher's server to answer a question today. A ruling that training is fair use would say nothing about whether a model may crawl your site tomorrow, what it owes you for doing so, or what you are entitled to serve it when it arrives.
The line that should worry small publishers most
The brief's most consequential passage for the industry is not about doctrine. On page 4 the government argues that an erroneous fair use ruling "would hamper competition in the market for LLMs, because only the largest technology companies might have the capital necessary to pay licensing fees", and that it is not in the public interest for those companies "to have an oligopoly on LLM training due to licensing entry barriers that function primarily as large subsidies for old mainstream media companies".
Read that as a publisher rather than as a lawyer. The government has adopted the argument that mandatory licensing helps large publishers and hurts small ones, and has deployed the interests of small newsrooms against the plaintiffs, a bloc that includes The Intercept and the Center for Investigative Reporting. Nieman Lab summarised the position as the administration saying a New York Times win would threaten national security and hurt small newsrooms. Whatever the merits, this is now the framing a publisher will meet across the table in any licensing conversation, and it is a framing under which the natural counter-argument, that small publishers deserve to be paid too, has been pre-emptively turned around.
What OpenAI's summary judgment motion says about your traffic
The defendants' 4 September motion is worth reading for its evidence about how retrieval actually works, independent of who wins. It is built on counting. Against 10.8 million asserted articles and 20 million ChatGPT conversations produced in discovery, OpenAI's experts report 24 instances of verbatim reproduction, the longest running 29 and 43 words, at rates the brief puts between 0.00011 and 0.000002 percent.
The Browse analysis is the operationally important one. OpenAI describes the sequence as: a user question becomes a search query passed to Microsoft Bing; Bing returns links and snippets; ChatGPT selects which pages look relevant; and only then does it request full content from the underlying site. On that account more than 96 percent of the claimed Browse infringement involved Bing snippets alone, with no request ever reaching the publisher's server. If the court accepts it, most of what publishers experience as an AI assistant reading their site is an argument about search snippets. It also means the retrieval that does hit your origin is a smaller, more selective and more commercially interesting slice of traffic than raw crawler logs suggest.
Two further defence arguments bear on operational decisions publishers have already made. On implied licence, OpenAI notes it disclosed the ChatGPT-User agent in March 2023 and stated it respected robots.txt, while the plaintiff publishers did not block it until April 2024, in a period when 48 percent of major news sites had already blocked some AI crawler by the end of 2023. Copies taken in that window, the defendants argue under Field v. Google, were impliedly licensed. On the DMCA claims, an expert review of 68 alleged regurgitations found 63 of them, or 92.7 percent, would have had an obvious source to the user. On market harm the brief reaches for the plaintiffs' own results: the Times exceeded 12 million subscribers during 2025 and targets 15 million by 2027, and its second-quarter 2026 digital advertising revenue rose 20.7 percent to $114.0m.
What the publishers are arguing back
The plaintiffs assert copying at acquisition, training, grounding and output, with a fifth category redacted. Acquisition covers the WebText, WebText2 and Common Crawl datasets plus the New York Times Annotated Corpus, 1.8 million articles from 1987 to 2007 obtained from the Linguistic Data Consortium under a non-commercial licence. Custom GPTs named in the filings include News Summarizer Ace, NYTimesGPT, Bypass Paywall, Remove Paywall and Article Reader.
Their sharpest point is procedural. Fair use is an affirmative defence, so the burden of showing no harm to existing or potential markets sits with the party asserting it, not with the publisher. Their dilution evidence is the part with the widest implications for advertising: Prism News, roughly 200 AI-generated publications presenting themselves as local newsrooms and staffed by four people in total; Integral Ad Science's forecast that as much as 90 percent of web content could be machine-generated by 2026, with individual sites capable of 1,200 articles a day; and the cost comparison, at approximately $6,800 to generate a million 500-word articles against close to $2bn a year for the Times to produce about half a million works. Underneath all of it sit the traffic ratios: OpenAI's crawl-to-referral ratio reached 1,500 to 1 by June 2025, Google's moved from 2 to 1 in 2015 to 18 to 1, and separate research has measured AI Overviews cutting publisher clicks by 39.8 percent.
Both defendants have requested oral argument. No date has been set, and no ruling should be expected quickly.
What leverage survives, and what to do about it
Two precedents already frame the range of outcomes. Meta won summary judgment on fair use in June 2025. Anthropic settled the authors' class action for $1.5bn in September 2025, the largest copyright settlement on record, which established a price without establishing a rule. Neither produced a recurring publisher revenue line, and a defence win here would produce less than that.
The honest planning assumption is therefore this. Legal leverage over training is now genuinely uncertain and may be worth close to nothing. Leverage over acquisition and outputs survives but is slow, expensive and available in practice only to publishers who can fund years of litigation. Leverage over live retrieval survives completely, because it was never a copyright question in the first place. When an assistant requests a page from your origin to answer a question now, you control what is served, on infrastructure you already pay for, under terms you set, with no counterparty agreement required. That is the asset a fair use ruling does not touch.
The practical sequence for the next quarter is unglamorous. Measure the retrieval traffic hitting your origin and separate it from crawler volume, because the two are different products and OpenAI's own Browse evidence suggests the split is severe. Stop modelling licensing as a base case and model it as upside. Where a licensing conversation is live, note that the market dilution argument has just been called deeply flawed by the US government and price accordingly. And treat the retrieval layer as the part of the business you can act on this quarter without waiting for Judge Stein. blankspace operates at that layer, injecting contextual brand facts into live agent retrievals at the CDN edge, which is one answer among several to the same question: what can a publisher monetise that no court is currently deciding.
Frequently asked questions
Did the DOJ rule that AI training is legal?
No. The Department of Justice is not a party and cannot rule on anything. It filed a Statement of Interest under 28 U.S.C. section 517, which is a formal expression of the government's view submitted for the court's consideration. Judge Sidney Stein decides the fair use question in In re OpenAI Copyright Infringement Litigation, and he is free to disagree with the government entirely.
Does this mean publishers will never be paid for AI training?
Not necessarily, but it removes a large part of the reason a model developer would feel obliged to pay. Licensing deals have always been commercial rather than compelled, and the strongest argument for signing one has been litigation risk. If training is held to be fair use, that risk falls sharply for written works, and licensing becomes a negotiation about data quality, freshness and indemnity rather than about liability.
What did the government not address in its filing?
Three things, all deliberately left open. It did not address how training data was acquired or stored. It did not address outputs that reproduce protected expression. And it did not address live retrieval, meaning the separate act of an AI assistant fetching a page from a publisher's server to answer a question today. Those questions remain fully live regardless of how the training issue is decided.
What is the market dilution theory and why does it matter to advertising?
Market dilution is the argument that mass machine-generated content destroys the market for human work even when no protected expression is copied, because it floods the same category. It matters commercially because it is the legal theory that most closely matches what publishers actually experience: not plagiarism, but substitution. The government calls it deeply flawed. If courts agree, publishers lose the doctrine that best describes the harm to their advertising business.
What should a publisher do while the case is pending?
Separate the two revenue questions. Rights revenue from training is now contingent on litigation nobody can time, so it should be treated as upside rather than budget. Retrieval revenue, from what AI agents fetch from your servers in real time, is available now, requires no agreement with an AI company, and is unaffected by the fair use ruling either way. Measuring the retrieval traffic you already receive is the cheapest first step and the one that informs every later decision.
