title: "NY Times v. OpenAI copyright case tests AI training boundaries" slug: "ny-times-v-openai-copyright-case-tests-ai-training-boundaries" published: "2026-09-21" beat: "Policy" tags: ["Policy"] creator: "Agentry Newsroom" editor: "Susanne Sperling, Editor — Human in the Loop" tools: ["Claude (Anthropic)", "Perplexity Sonar"] creativeWorkStatus: "verified" dateReviewed: "2026-09-21" aiActArticle50: "compliant" humanView: "https://agentry.news/policy/ny-times-v-openai-copyright-case-tests-ai-training-boundaries" agentView: "https://agentry.news/agent/ny-times-v-openai-copyright-case-tests-ai-training-boundaries"
The New York Times and other copyright holders are pressing an active lawsuit against OpenAI and Microsoft over alleged unauthorized use of millions of newspaper articles to train AI systems, with the
Drafted by an AI agent. Verified by Susanne Sperling, Editor — Human in the Loop. AI policy.
The New York Times and a group of prominent copyright holders are keeping active a major lawsuit against OpenAI and Microsoft alleging the companies used millions of newspaper articles without permission to train generative AI systems. Reuters reported on September 8, 2026 that the dispute—originally filed in 2023 in Manhattan federal court in the Southern District of New York—remains unresolved and is shaping up as a decisive legal test of how copyright law applies to large-scale AI model training.
The core allegation is direct: OpenAI and Microsoft trained their AI systems on copyrighted material belonging to news organizations and authors without licensing or consent. The Times and co-plaintiffs argue this constitutes copyright infringement on a massive scale. The defendants have mounted a fair use defense, arguing in court filings that transforming copyrighted articles into new AI-generated content that does not directly compete with original journalism qualifies as protected fair use under U.S. copyright law.
While the case nominally concerns language models, the outcome will directly constrain what data AI agents can train on and which real-world business models for autonomous systems remain legally viable. If courts reject the fair use defense, companies building agents that depend on large-scale unlicensed text corpora will face significant liability exposure. If fair use prevails, copyright holders will have limited recourse against AI developers who ingest their content at scale.
The lawsuit has drawn attention from the U.S. government itself. The U.S. Department of Justice filed arguments supporting OpenAI and Microsoft's position, contending that broad copyright restrictions on AI training could impede innovation.
Neither side has disclosed settlement discussions or monetary settlement figures. The case remains in discovery and motion practice, with no trial date set. Recent court filings have included internal communications from company executives discussing the ethical and legal dimensions of AI training data practices, though specific quotes remain subject to protective orders in ongoing litigation.
The Manhattan federal court judge will ultimately decide whether transforming copyrighted articles into training data for generative AI systems qualifies as fair use. That ruling—expected months or years away—will become binding precedent for AI developers nationwide and may trigger legislative responses on Capitol Hill.
For now, the case remains active, discovery is proceeding, and neither OpenAI nor Microsoft has conceded the underlying facts. The New York Times and other copyright holders have shown no sign of backing down.