Technology / Intellectual Property

USA Today publishers seek over $250 million in OpenAI copyright lawsuit

Fourteen publishing entities behind USA Today and 18 regional newspapers allege that OpenAI copied hundreds of thousands of articles for AI training and produced outputs that can substitute for their reporting. The claims have not been decided.

INNOVOX News DeskOct 9, 2026 · 6 min read
A blue USA Today newspaper vending machine standing beside other newspaper boxes in Fort Wayne, Indiana
Twzink · CC0 1.0 via Wikimedia Commons

The story

USA Today Co. and 13 affiliated publishing entities sued OpenAI on October 8, alleging that the artificial-intelligence company copied protected journalism without permission to train its models and then used that material to generate competing answers. The complaint was filed in the U.S. District Court for the Southern District of New York as USA Today Co., Inc. et al. v. OpenAI Foundation et al., case 1:26-cv-08892. The court docket classifies it as a copyright action under 17 U.S.C. Section 501.

The plaintiffs represent 19 publications: USA Today and 18 regional newspapers, including The Arizona Republic, the Detroit Free Press and The Des Moines Register. Reuters reported that the publishers accuse OpenAI of taking hundreds of thousands of articles and other materials for ChatGPT training without authorization. The complaint seeks more than $250 million in damages as well as a court order restricting the alleged conduct. Those remedies are requests, not an award, and no judge has ruled that OpenAI infringed the publishers' rights.

The lawsuit advances three related theories. It alleges direct copyright infringement by OpenAI entities, vicarious infringement based on control over products that allegedly reproduce or derive value from the works, and unlawful removal or alteration of copyright-management information under Section 1202 of the Digital Millennium Copyright Act. The last claim focuses on information such as author, publication and copyright notices. It can raise different questions from whether using a work for training is fair use.

To connect its broad legal claims to specific data, the complaint says the publishers found more than 160,000 entries associated with their domains in the WebText dataset and more than 122 million tokens from their material in the C4 dataset. Those numbers are allegations drawn from the publishers' analysis, not court findings, and public datasets do not by themselves establish which items appeared in every OpenAI training run. Their significance will depend on evidence about provenance, model versions, access and how particular materials were processed.

The publishers also try to show market harm at the output stage. The complaint includes examples that it says were generated with GPT-5.6 and that summarize or reproduce reporting without sending a reader to the original publication. The theory is that a sufficiently complete AI answer can substitute for visiting a news site, reducing advertising impressions, subscriptions and the value of licensing. OpenAI had not publicly responded to this new complaint when Reuters and The Verge reported the filing.

OpenAI has previously taken a different position in other news-copyright disputes. In a 2024 public statement about litigation brought by The New York Times, the company said training on publicly available internet material is fair use, described regurgitation as a rare failure and highlighted opt-out tools and licensing partnerships. That statement is not an answer to the USA Today complaint, but it identifies the defenses and product safeguards likely to frame the wider debate.

The unresolved legal boundary is important. Copyright law protects original expression, not facts, and fair use examines purpose, the nature of the work, the amount used and market effect. AI developers argue that training extracts statistical relationships for a new, transformative purpose. Publishers respond that copying at model scale and producing article substitutes threatens the market that funds the underlying journalism. A ruling may therefore turn on technical evidence about both training and outputs rather than on a single rule covering every model and dataset.

The case joins a growing group of U.S. lawsuits by authors, newspapers, visual artists and other rights holders against generative-AI companies. Consolidation can improve efficiency when cases share technical and legal questions, but the plaintiffs and works still differ. USA Today's filing adds a particularly broad network of national and regional reporting, potentially giving the court evidence about how the same model behavior affects publications with different audiences and business models.

INNOVOX analysis: the lawsuit matters because it links dataset-level claims, output examples and an asserted licensing market in one record. A publisher does not prove infringement merely by showing that its pages appeared in a web corpus, while a model developer does not resolve the dispute merely by calling training transformative. The decisive questions are likely to include which works were copied, what each system retained, whether outputs are substantially similar, how users obtain them and whether those uses displace plausible markets for access or licensing.

The next signals will come from procedure and evidence. OpenAI's response could seek dismissal, transfer or coordination with related cases and will show how it applies its fair-use position to this complaint. Any discovery order covering training records, data deletion, output testing or copyright notices would be technically consequential. A settlement or licensing agreement would also matter, but it would set commercial terms rather than a judicial precedent. Until then, the complaint is a detailed allegation and a high-value demand, not a decision on the legality of AI training.

INNOVOX analysis

The filing turns the AI-training dispute into a test involving a national title and a broad regional-news network. Its most consequential move is to join training-data allegations with a market-substitution theory about generated summaries. If the case advances, evidence about data provenance, model behavior, attribution and licensing economics could matter as much as abstract arguments over whether training is transformative.

What to watch

Watch for OpenAI's formal response, any request to move or consolidate the case, and early rulings on fair use and the Digital Millennium Copyright Act claim. Discovery over training datasets and model outputs, evidence of traffic or subscription harm, and any licensing talks will show whether the dispute moves toward a precedent or a commercial settlement.