Skip to main content
DOJ building facade, Washington D.C. AI training, copyright, fair use, legal case.

Editorial illustration for DOJ Backs Fair Use for AI Training in Copyright Case

DOJ Backs AI Fair Use in NYT Copyright Case

4 min read

The Justice Department filed a brief this week in the Southern District of New York, staking out a position that could shape every AI copyright lawsuit filed since 2023. The case in question started when The New York Times sued OpenAI and Microsoft in late 2023, claiming the companies trained GPT-4 and other models on millions of its articles without permission. The Times wants billions in damages and asked the court to force destruction of any model trained on its work. Other rights holders have since joined the fight, turning a single lawsuit into a consolidated case that lawyers and judges around the country are watching closely.

Now the DOJ has weighed in, and it's not on the side of the publishers. The department argues that feeding copyrighted text into an AI model during training doesn't amount to infringement, drawing a line between the act of training and whatever the model later produces. That argument, if it holds, would give AI companies a legal foundation that extends well past this one case.

The US Department of Justice has sided with AI companies in the consolidated lawsuit involving the New York Times and other rights holders, arguing that training AI models on copyrighted material qualifies as fair use.

Why this matters

This filing gives developers and founders a preview of how the executive branch wants judges to draw the line between training data and outputs, not a settled rule. The DOJ's brief isn't binding on the court hearing the New York Times case, and Judge Sidney Stein could easily reject its reasoning. But it signals that the current administration sees fair use as the legal scaffolding it wants under the AI industry, which matters for anyone building products on scraped or licensed text right now.

Worth watching: the DOJ's own example, cited almost as an aside, that Times writers reportedly use LLMs to draft copy. That detail cuts against the paper's infringement claim more than any abstract legal theory does, and plaintiffs' lawyers will have to address it directly. For research teams weighing whether to license training data or bet on fair use protections, this is a data point, not a resolution.

The next real signal comes from Stein's ruling, not from DOJ filings, and that's where founders should be watching for the durable precedent rather than treating this brief as one.

Common Questions Answered

What position did the DOJ take in the New York Times vs. OpenAI copyright lawsuit?

The Justice Department filed a brief arguing that training AI models on copyrighted material qualifies as fair use, effectively siding with OpenAI and Microsoft in the consolidated lawsuit. This position could significantly influence how courts interpret copyright law in relation to AI model training across future cases.

What specific claims did The New York Times make against OpenAI and Microsoft?

The New York Times sued OpenAI and Microsoft in late 2023, claiming the companies trained GPT-4 and other models on millions of its articles without permission. The Times is seeking billions in damages and asking the court to force destruction of any model trained on its copyrighted work.

How binding is the DOJ's brief on the court's final decision in this case?

The DOJ's brief is not legally binding on Judge Sidney Stein, who is hearing the New York Times case, and the judge could easily reject its reasoning. However, it signals the current administration's position that fair use should be the legal framework supporting the AI industry.

Why does the DOJ's filing matter for AI developers and founders?

The DOJ's brief provides developers and founders insight into how the executive branch wants judges to distinguish between training data and AI outputs, establishing a preview of potential legal scaffolding for the AI industry. This matters significantly for anyone building products using scraped or licensed data, as it could influence future copyright litigation outcomes.

LIVE21:38Switchyard: A Rust Library That Routes LLM Traffic Between OpenAI and Anthropic