Editorial illustration for Microsoft Claims Chatbot Logs Skewed to Overstate NYT Article Use
Microsoft: Less Than 1% of Copilot Logs Used NYT
Microsoft Claims Chatbot Logs Skewed to Overstate NYT Article Use
Microsoft has a number for the judge in its ongoing copyright fight with The New York Times, and it's a small one. In court filings submitted as part of discovery, the company says fewer than 1 percent of 8.2 million Copilot chat logs turned up even 16 matching words from news content the chatbot was trained on. That's the figure Microsoft is leaning on to argue that Copilot doesn't function as a substitute for reading the original articles, books, or investigative reports it was trained on.
The logs themselves weren't random. Microsoft says it handed over conversations specifically flagged for keywords tied to plaintiff websites, meaning the sample was already stacked toward finding overlap. Even so, the company says the results came back thin: 59,545 logs out of millions showed any meaningful word overlap with news content, and separate analyses tied to the Center for Investigative Reporting and a group of book authors turned up similarly small numbers.
Microsoft is now using that scarcity as evidence in its defense against The New York Times, CIR, and the Authors Guild, all of which have accused the company of building its AI on copyrighted work without permission or payment.
Microsoft’s Copilot rarely reproduces even full sentences from news articles and books, let alone substantive chunks that could substitute for the original, the company says in new legal filings as it fights copyright claims from publishers including The New York Times and book authors.
Why this matters
The case turns on a numbers fight, and that should worry anyone building products on licensed or scraped content. Microsoft's own framing admits the sample was cherry-picked for keywords tied to NYT domains, then used that skewed pool to argue infringement is rare. That's a tell: if the company needs a curated dataset to make its case look good, the underlying exposure is probably larger than 59,545 logs out of 8 million suggests.
For developers and founders shipping chatbots on top of licensed news or book content, the lesson isn't "regurgitation is rare, relax." It's that courts are going to demand exact methodology, sample selection, and word-count thresholds before anyone gets to claim their tool is clean. Sixteen words is an oddly specific bar, and whoever sets that bar in this case will shape fair-use arguments for every RAG product and search-grounded assistant that touches paywalled journalism. Watch how Judge Sidney Stein's court treats Microsoft's sampling methodology, because that ruling, not the 1 percent figure, is what will actually set precedent.
Common Questions Answered
What percentage of Copilot chat logs contained matching words from New York Times articles according to Microsoft's court filing?
Microsoft claims that fewer than 1 percent of 8.2 million Copilot chat logs contained even 16 matching words from news content the chatbot was trained on. The company is using this statistic to argue that Copilot does not function as a substitute for reading original articles, books, or investigative reports it was trained on.
How does Microsoft argue that Copilot differs from direct article substitution in its copyright defense?
Microsoft contends in its legal filings that Copilot rarely reproduces even full sentences from news articles and books, let alone substantive chunks that could serve as replacements for the original content. The company uses this argument to demonstrate that users cannot simply rely on Copilot output instead of accessing the actual published materials.
What concern does the article raise about Microsoft's methodology for presenting the 1 percent statistic?
The article suggests that Microsoft's sample was cherry-picked for keywords tied to New York Times domains and then used that skewed pool to argue that infringement is rare. This selective framing indicates that if Microsoft needs a curated dataset to make its case appear favorable, the actual underlying copyright exposure is likely larger than the 59,545 logs out of 8 million initially presented.
Why should developers and founders building chatbots be concerned about this copyright case?
The case demonstrates that the outcome depends heavily on how companies frame their data and statistics regarding content usage, which creates uncertainty for anyone building products on licensed or scraped content. The dispute highlights that copyright liability for AI chatbots trained on published materials remains a significant legal and financial risk for developers in the industry.
Further Reading
- Microsoft says virtually nobody was grabbing NYT articles through its chatbot - The Verge
- OpenAI may have made a fatal misstep in copyright fight with news outlets - Ars Technica
- News outlets ask judge to sanction OpenAI in copyright fight - AP News
- OpenAI loses fight to keep ChatGPT logs secret in copyright case - Reuters
- OpenAI appeals data preservation order in NYT copyright case - Reuters