Core Event: Microsoft Submits Legal Filing to Counter NYT Copyright Claims
Microsoft has filed legal documents in its copyright fight with The New York Times and book authors, asserting that its Copilot chatbot rarely reproduces news articles or books verbatim. The filing is part of a key stage in the case: Microsoft is asking the judge to issue a summary judgment, which could end the case early if granted.
Key facts:
- Microsoft provided 8.2 million Copilot chat logs to an expert hired by news publishers, describing them as logs “specifically chosen” because they hit keywords tied to the news plaintiffs’ websites
- Microsoft says 59,545 of those logs contained at least 16 words in common with news content used to ground the AI model, representing fewer than 1% of the dataset
- In the authors’ case, an expert found only 24 responses with at least 30 matching words across 8.2 million Copilot conversations; only 10 of 212 books evaluated had any matches
- Microsoft says an expert for the Center for Investigative Reporting found 51 instances of “substantial overlap” with CIR work in the dataset
Data Contrast: Filtered Logs vs. Actual Reproduction Rates
Microsoft emphasizes that the 8.2 million logs were not a random sample. They were selected because they hit keywords connected to the news plaintiffs’ websites, and therefore, in Microsoft’s framing, were among the conversations most likely to contain the plaintiffs’ works. Microsoft is using that point to argue that even in a higher-risk dataset, actual reproduction of protected text was rare.
The numbers show the contrast:
- 59,545 logs contained at least 16 matching words with news content, or roughly 0.7% of the 8.2 million logs
- Only 24 responses in the books-related analysis contained at least 30 matching words
- CIR-related material produced 51 instances described as “substantial overlap”
That sets up the central dispute: publishers and authors argue Microsoft and OpenAI built commercial products on their work that can substitute for the originals, while Microsoft argues that occasional textual overlap does not undermine the transformative purpose of large language model training.
Legal Positions and Fair Use Arguments
Microsoft argues that using copyrighted material in AI training datasets should qualify as “fair use” under U.S. copyright law. Its core position is that systems like Copilot may rely on copyrighted material during training, but the resulting tools are used for purposes significantly different from the original works. In Microsoft’s view, occasional reproduction of text does not defeat the transformative nature of LLM training.
The New York Times disagrees. Its lead counsel, Ian Crosby, said in a statement that documents and testimony uncovered during discovery “lead to only one conclusion”: that Microsoft and OpenAI stole from The New York Times to make commercial products that substitute for its journalism, threaten its business, and undermine its industry. He added that the Times looks forward to Microsoft and OpenAI being held accountable.
The news publishers’ and book authors’ claims have been consolidated under one judge to streamline the process, despite objections from the publishers and authors. Microsoft submitted the filing as it seeks summary judgment; if the judge sides with the publishers and authors, the case will continue in court. The original report also notes that the Trump administration filed a statement of interest this week in the New York Times case, supporting OpenAI.
Reader Recommendations
Who should pay attention, and who should be cautious:
- General Copilot users: Microsoft’s disclosed data suggests a low chance of encountering long verbatim reproductions, but AI outputs should still be checked
- Content creators and media professionals: The outcome could influence how courts define the boundaries of AI training data use
- Enterprise users: Human review remains prudent for outward-facing Copilot-generated content, especially long-form text or material involving quotations and copyrighted sources
Best practice: Until the litigation is resolved, organizations using Copilot-generated content should avoid publishing long passages without review. Individual users can continue using the tool, but should not assume AI outputs are automatically copyright-safe or factually reliable.
Bottom Line
The core dispute is not simply whether AI learned journalistic or literary style. It is about how the law should treat the relationship between training data and final model outputs. Microsoft’s filtered log data highlights a key question for AI copyright cases: when a large-scale system produces only sporadic textual matches, is that enough to support claims of systemic infringement? The answer could become an important test of fair use in the AI era.



