# Microsoft: Copilot rarely reproduces NYT content in lawsuit

**Published:** 2026-09-07T07:52:55.771Z  
**Topic:** Microsoft  
**Sentiment:** neutral  
**Publisher:** TrendWatcher — https://www.trendwatcher.in/article/cac6acd1-1362-4ce2-95d1-8143e978d0ee

Microsoft claims under 1% of 8.2 million Copilot chat logs contained 16+ words from news articles, challenging copyright infringement lawsuits from publishers.

Microsoft has asserted in new legal filings that its Copilot AI chatbot rarely reproduces substantial portions of copyrighted content, with fewer than 1% of 8.2 million analyzed chat logs containing at least 16 words from news articles [1]. This claim is central to Microsoft's defense against copyright infringement lawsuits brought by publishers, including The New York Times, and authors, who allege that AI models were built on their works and now compete directly by regurgitating content [1, 3].

| At a glance | |
|---|---|
| Company | Microsoft |
| Product | Copilot AI chatbot |
| Key Figure | Under 1% of chat logs reproduced 16+ words from news |
| Context | Legal defense against copyright claims |

## Copyright dispute over AI training data

Microsoft submitted 8.2 million Copilot chat logs during the discovery phase of the lawsuit, stating these logs were specifically chosen for hitting keywords related to news plaintiffs' websites, making them most likely to contain relevant content [1, 3]. The company's analysis of this dataset found that 59,545 conversations, or less than 1%, contained at least 16 words in common with news content used to ground the AI model [1, 3]. An expert for the Center for Investigative Reporting (CIR) identified 51 instances of "substantial overlap" with CIR work within this dataset [1].

In a separate but consolidated lawsuit by authors, an expert found only 24 responses out of the 8.2 million Copilot conversations that contained at least 30 matching words [1]. Furthermore, only 10 of 212 evaluated books had any matches [1]. Microsoft argues these figures support its position that using copyrighted content for AI training datasets should be considered fair use, contending that while systems like Copilot rely on such material, they serve "significantly different purposes than the original" [1]. The company concludes that occasional text reproduction "hardly undermines the transformative purpose of LLM training" [1].

The New York Times, a lead plaintiff, has rejected Microsoft's conclusions. Its lead counsel, Ian Crosby, stated that discovery documents and testimony indicate Microsoft and OpenAI "stole" from the Times to create commercial products that substitute for its journalism, threaten its business, and undermine the industry [1, 3].

## Legal implications and fair use argument

Microsoft filed its arguments as part of a motion for summary judgment, seeking to end the case at an early stage [1, 3]. The lawsuits, which consolidate claims from major news publishers and the Authors Guild against Microsoft and OpenAI, allege that the companies built products on their works and now compete directly by regurgitating copyrighted content [1, 3]. The Trump administration also filed a statement of interest supporting OpenAI in The New York Times case [1].

If the judge grants summary judgment, the case would conclude; otherwise, the legal battle will proceed to full litigation [1, 3]. The core of Microsoft's defense rests on the fair use doctrine, asserting that the transformative nature of AI model training outweighs instances of text reproduction [1].

## What to watch

*   **Summary Judgment Decision:** The judge's ruling on Microsoft's motion for summary judgment will determine if the case proceeds to trial or is dismissed [1].
*   **Further Discovery:** If the case continues, additional discovery could reveal more details about the extent of copyrighted material use and reproduction [1].
*   **Industry Response:** The outcome of this litigation could set precedents for how AI companies use copyrighted content for training and how content creators are compensated [1].

The dispute highlights a fundamental tension between the development of AI technologies that rely on vast datasets and the rights of content creators whose work forms the basis of these datasets. The court's decision will significantly influence the legal framework for AI development and intellectual property in the digital age.

## Sources
1. The Verge — [Microsoft says virtually nobody was grabbing NYT articles through its chatbot](https://www.theverge.com/policy/990267/microsoft-openai-new-york-times-authors-lawsuit)
2. Nytimes — [Contests - The New York Times](https://www.nytimes.com/spotlight/learning-contests)
3. Rocketnews — [Microsoft says virtually nobody was grabbing NYT articles...](https://rocketnews.com/2026/09/microsoft-says-virtually-nobody-was-grabbing-nyt-articles-through-its-chatbot/)
4. Brutalist — [The day's headlines delivered to you without bullshit.](https://brutalist.report/?before=2026-09-05)

---
Cite as: TrendWatcher, "Microsoft: Copilot rarely reproduces NYT content in lawsuit", https://www.trendwatcher.in/article/cac6acd1-1362-4ce2-95d1-8143e978d0ee
