Skip to main content
Microsoft internal documents reveal AI scraping details in NY Times lawsuit, showing legal battle over data use.

Editorial illustration for Microsoft internal docs reveal AI scraping details in NY Times lawsuit

Microsoft AI Scraping Exposed in NY Times Lawsuit

4 min read

A federal court in New York unsealed a batch of internal Microsoft and OpenAI documents on Thursday, filed as part of a motion for summary judgment by news plaintiffs led by The New York Times. The filings mark a shift in a copyright case that has run for months under heavy redaction, with both AI companies working to keep specifics about their data practices sealed. That effort broke down this week.

The documents, according to the news organizations, show how Microsoft and OpenAI privately assessed the risks of scraping journalism to train products like ChatGPT and Copilot, even as they built and released those tools. Named executives appear throughout the filings, including Microsoft's Brent Hecht and OpenAI's Nick Turley, whose internal messages are cited directly by the plaintiffs. The News organizations argue the records undercut the fair use defense both companies have leaned on in court.

What's striking is the language used inside these companies to describe their own AI training practices, language now at the center of the plaintiffs' argument that Microsoft and OpenAI understood the stakes long before publishers ever filed suit.

Perhaps most explosively, Microsoft Director of Applied Science Brent Hecht repeatedly warned in documents that scraping news for AI training was “an astonishing theft of unprecedented proportions,” calling it perhaps the “largest theft of labor in human history,” news orgs said.

Why this matters

For anyone building products on top of licensed or scraped data, this unsealing is a reminder that the legal exposure isn't hypothetical. An internal Microsoft executive reportedly describing the practice as the "largest theft of labor in human history" is not the kind of language that survives discovery without consequence. If that characterization came from inside the company that co-built Copilot, it suggests Microsoft and OpenAI understood the copyright risk well before ChatGPT ever shipped, not after the lawsuits started piling up.

For founders training models on web-scraped corpora, the case is now a live test of how courts will treat "we knew but proceeded anyway" evidence. For researchers, it's worth watching what else surfaces as these seals get pulled back, since summary judgment filings tend to include the documents each side thinks are most damaging to the other. The New York Times case has been the bellwether for a reason, and this is the first real look at what Microsoft and OpenAI didn't want the public to see.

Common Questions Answered

What did Microsoft Director Brent Hecht say about AI scraping in the unsealed documents?

Microsoft Director of Applied Science Brent Hecht repeatedly warned in internal documents that scraping news for AI training was "an astonishing theft of unprecedented proportions," calling it perhaps the "largest theft of labor in human history." These characterizations were revealed in documents unsealed during the New York Times lawsuit against Microsoft and OpenAI.

Why were the Microsoft and OpenAI internal documents previously kept sealed in the NY Times lawsuit?

Both Microsoft and OpenAI worked to keep the specifics about their data practices sealed throughout the copyright case that had run for months under heavy redaction. The companies' effort to maintain confidentiality over these documents broke down this week when a federal court in New York unsealed the batch of filings.

What does the unsealing of these documents reveal about Microsoft and OpenAI's understanding of copyright risks?

The internal characterization of AI scraping as potentially the "largest theft of labor in human history" suggests that Microsoft and OpenAI understood the copyright risks associated with their data practices well before the lawsuit was filed. The language used by executives in these documents indicates the companies were aware of the legal exposure involved in their AI training methods.

How does this court filing impact companies building products on licensed or scraped data?

The unsealing of these documents serves as a reminder that legal exposure from data scraping practices is not hypothetical but represents real legal consequences. Companies building products on top of licensed or scraped data now have evidence that even major tech executives have internally acknowledged the serious legal and ethical concerns with such practices.

LIVE05:05UN partners with Google to prepare global data for AI analysis