Skip to main content
Courtroom scene with a gavel, OpenAI and Microsoft logos on a screen, and newspaper clippings stacked beside a robot.

Editorial illustration for OpenAI and Microsoft Sued Over Alleged Newspaper Text Scraping Without License

OpenAI, Microsoft Sued Over Unlicensed News Content Scraping

OpenAI, Microsoft sued for scraping newspapers and using text without license

Updated: 3 min read

This lawsuit isn't about copyright. It's a $10 billion bet on whether you can build an industry by ignoring property rights. Eight newspaper publishers, including heavyweights like the New York Daily News and the Chicago Tribune, have filed suit against OpenAI and Microsoft.

The charge is theft, pure and simple. They allege their articles were scraped, stripped of copyright markers, and fed into AI models without payment or permission. Microsoft, they argue, wasn't just hosting the servers; it was a co-designer and direct beneficiary.

The text now surfaces in products like Copilot, placing their reporting directly in the AI's output.

Nine US regional newspapers have slapped OpenAI and Microsoft with a massive copyright lawsuit, seeking damages that could top $10 billion. At the same time, a federal court has ordered OpenAI to hand over internal communications about book datasets it allegedly sourced from a pirate library.

But journalism is just one battlefield. Look at the separate case over OpenAI's "Books1" and "Books2" datasets. There, evidence points to Library Genesis—"LibGen"—a sprawling pirate site.

If that's true, foundational AI models were built on a mountain of stolen books. That context makes the newspaper suit's apocalyptic demand for the destruction of infringing GPT models and datasets more than a bargaining tactic. It's a logical conclusion.

The plaintiffs argue the entire enterprise is poisoned at the root, built on a habit of taking first and negotiating later. Now the bill is due, and it includes a clause to dismantle the factory. A judge will finally rule on Silicon Valley's long-standing assumption: that the internet is a free buffet.

Common Questions Answered

How much in damages are OpenAI and Microsoft facing in this lawsuit?

The lawsuit is seeking damages exceeding $10 billion, with potential penalties of up to $150,000 per copyrighted work for willful infringement. This substantial financial claim underscores the serious nature of the alleged intellectual property violations by the tech companies.

What specific allegations does the lawsuit make about OpenAI and Microsoft's content scraping practices?

The lawsuit alleges that OpenAI and Microsoft systematically harvested news content without proper authorization, stripping copyright notices and using text without licensing. Microsoft is not just viewed as an infrastructure provider, but as a co-designer of AI models and a direct beneficiary of the alleged content theft.

Why are news organizations concerned about AI companies' data training practices?

News organizations are increasingly worried about how their intellectual property is being used without compensation in AI training datasets. The lawsuit highlights the growing tension between tech giants and media companies over the unauthorized use of copyrighted content for artificial intelligence development.

LIVE17:24Silicon Valley Split on Regulating Chinese AI Models