Seattle Times and Newsday are the latest publications to sue OpenAI and Microsoft
By Jakub Antkiewicz
•2026-09-06T11:57:50Z
Publishers Intensify Legal Fight Over AI Training Data
The Seattle Times and Newsday have joined the expanding legal battle against OpenAI and its primary investor, Microsoft, filing a lawsuit that alleges copyright infringement through the use of their journalism to train large language models. This action escalates the conflict between media outlets and AI developers, highlighting a fundamental dispute over the value and ownership of content in the era of generative AI. The suit argues that AI systems are not creating new content but are instead consuming and repurposing existing works, threatening the financial viability of the original publishers.
The complaint echoes arguments from a similar 2023 lawsuit by The New York Times, describing generative AI as “a snake eating its own tail” that could ultimately “destroy the very organizations” that produce the high-quality information it relies on. The situation is particularly complex because The Seattle Times had previously received funding from both Microsoft and OpenAI for journalism projects. In response, a Microsoft spokesperson expressed surprise at the lawsuit but stated the company is open to exploring a resolution.
- The lawsuit describes AI products like ChatGPT and CoPilot as “rapacious consumers” of human-authored content.
- It claims these tools deliver “copies and derivative imitations” of the original works they ingest.
- The core argument is that this practice constitutes massive copyright infringement for commercial gain.
This growing wave of litigation signals a significant operational and financial risk for the AI industry. The outcomes of these cases could set a critical precedent for data acquisition, potentially forcing AI labs to pivot from scraping public web data to negotiating expensive licensing deals with content creators. A legal victory for the publishers would compel a systemic re-evaluation of how foundation models are built, possibly increasing costs and creating a competitive advantage for firms with access to large, proprietary, or fully licensed datasets.
The persistent legal action from publishers is escalating the operational risk for AI developers, suggesting that relying on scraped web data is an increasingly untenable long-term strategy. The future of foundation model training will likely depend on building defensible, licensed data pipelines.