News publishers presented core legal arguments to a New York federal court this month, asking a judge to decide whether OpenAI must pay to train ChatGPT on copyrighted journalism.
Key points
- Publishers and OpenAI submitted core arguments to a federal court in New York regarding AI training on copyrighted news articles.
- Ziff Davis CEO Vivek Shah said an estimated $20 billion annual licensing fee is small compared to AI infrastructure spending.
- The Department of Justice filed a statement supporting fair use claims to protect domestic artificial intelligence development.
- More than 550 regional and national publications have joined consolidated legal actions challenging OpenAI and Microsoft.

The consolidated case combines lawsuits from The New York Times, Ziff Davis, and the New York Daily News. The outcome will establish whether scraping online articles to build artificial intelligence models falls under fair use or constitutes mass copyright infringement under United States law.
The Financial Argument for Content Licensing
Publishers argue that commercial AI companies rely on original reporting to make their conversational tools authoritative. Media executives assert that unauthorized ingestion of their reporting reduces direct web traffic and replaces source reporting with machine-generated summaries.
Writing in Fortune, Ziff Davis CEO Vivek Shah addressed the financial balance between the technology and publishing sectors. Shah calculated that paying annual licensing royalties across the entire news industry would cost roughly $20 billion. In contrast, OpenAI projects spending $750 billion on computing infrastructure through 2030, supported in part by a $122 billion funding round closed in early 2026 at an $852 billion valuation.
A sensible fee structure creates a flywheel of quality inputs and quality outputs, benefiting the AI consumer and the public good.
Shah described the estimated licensing cost as little more than a rounding error for frontier AI companies. Publishers maintain that statutory or mandatory licensing fees are necessary to sustain newsrooms, warning that selective commercial partnerships with select outlets fail to support the wider press corps.
OpenAI Defends Web Scraping as Fair Use
OpenAI asked the court to reject the publishers’ claims, maintaining that pretraining language models on publicly accessible web pages qualifies as fair use. The company argued that the process transforms factual text into entirely new functional systems without creating market substitutes for the underlying articles.
OpenAI also asserted that it held an implied license to scrape public sites. In legal filings, the firm noted that media organizations did not initially implement technical restrictions, such as the standard robots.txt exclusion protocol, to prevent automated web crawlers from reading their archives.
OpenAI likened its data collection practices to the early development of search engines like Google, which indexed billions of pages without upfront licensing agreements. The media industry has since shifted course, with hundreds of publishers now deploying robots.txt rules and technical barriers to block AI data collection.
Government Intervention and Coalition Growth
The dispute has drawn direct interest from federal regulators concerned with national competitiveness. The U.S. Department of Justice submitted a Statement of Interest to the court cautioning against restrictions on model development.
Constraining LLM development under a misunderstanding of fair use doctrine would thwart…creative and scientific progress while hindering American prosperity and economic mobility.
Despite federal support for broad data ingestion, legal pressure on AI developers continues to expand. Platkin LLP now represents a coalition of more than 550 regional, local, and specialty publications suing OpenAI and Microsoft over content usage.
Over 500 news publications are standing together because what happens here will help determine whether local and independent journalism can continue to thrive in communities across the country.
With evidentiary discovery drawing to a close, the federal judge presiding over the consolidated action is expected to address the core legal questions surrounding fair use and copyright liability in generative AI development.





