Seattle Times and Newsday Join Growing Lawsuit Wave Against OpenAI and Microsoft Over AI Training Practices
The Seattle Times and Newsday have filed separate lawsuits in federal court accusing OpenAI and Microsoft of harvesting their articles without permission to train generative AI systems. The complaints allege that the tech firms copied protected news content to improve large language models that now power products such as ChatGPT, thereby infringing on the newspapers' copyrights and depriving them of licensing revenue.
Both suits were lodged in the U.S. District Court for the Western District of Washington and the Southern District of New York, respectively. Plaintiffs claim the defendants used automated scrapers to collect thousands of articles, including paywalled pieces, and incorporated the text into the training datasets that underpin OpenAI's models, which Microsoft now distributes through its Azure cloud services. The filings seek injunctive relief to halt further use of the newspapers' material, as well as monetary damages and a share of any commercial gains derived from the AI products.
The legal challenge rests on a long‑standing tension between the open‑internet approach used to build AI models and the intellectual‑property rights of content creators. Under U.S. copyright law, reproducing a protected work without a license can constitute infringement, even when the material is later transformed by machine learning algorithms. Defendants have argued that the use qualifies as “fair use” because it is transformative and does not replace the original articles, a position that courts have evaluated inconsistently in recent years.
Seattle Times and Newsday are the latest in a growing list of media companies that have taken the issue to court. Earlier this year, the Associated Press, Reuters, and other outlets filed similar actions against OpenAI, prompting a wave of debate within the publishing industry about how to protect digital content in an AI‑driven landscape. Those earlier cases have highlighted the difficulty of proving that specific excerpts were used in training, given the opaque nature of proprietary datasets.
The lawsuits underscore a broader concern among journalists that AI tools can reproduce news stories verbatim or generate summaries that siphon traffic away from original sources. Media organizations argue that the loss of readership translates directly into reduced advertising and subscription income, threatening the financial viability of newsrooms already grappling with declining revenues.
OpenAI and Microsoft have declined to comment on the pending litigation, but both companies have previously asserted that their models are trained on publicly available data and that they respect copyright law. In public statements, OpenAI has emphasized ongoing efforts to develop “responsible AI” practices, including exploring licensing agreements with content providers.
Legal analysts note that the outcome will likely hinge on how courts interpret the fair‑use defense in the context of machine‑learning training. Some experts predict that a ruling favoring the publishers could force AI developers to negotiate licensing deals, while others warn that overly restrictive judgments might hamper innovation in the rapidly evolving field.
If the courts grant the requested injunctions, OpenAI may be required to purge the disputed material from its training pipelines and potentially compensate the newspapers for past use. Conversely, a dismissal could embolden other tech firms to continue aggregating online content with minimal oversight. Both scenarios carry significant implications for the balance between open data ecosystems and the rights of content creators.
The cases arrive at a moment when policymakers are increasingly scrutinizing AI’s impact on intellectual property, privacy, and misinformation. As the litigation proceeds, the publishing industry and the technology sector are watching closely, aware that the verdict could set a precedent shaping the future relationship between news media and artificial‑intelligence developers.
Comments (0)
Be the first to comment.
Join the discussion