Skip to main content
Artists protest AI data use, holding signs about emotional harm from unauthorized training.

Editorial illustration for Artists sue over AI training, citing emotional harm from data use

Artists Sue Over AI Training Data Without Consent

4 min read

Kirk Wallace Johnson found his own books in a database he never agreed to join. The Feather Thief and The Fishermen and the Dragon, two nonfiction works that took him five to six years apiece to research and write, showed up in a searchable dataset The Atlantic published listing works used to train AI systems. His books had been pirated and fed into a chatbot's training data without his knowledge or consent.

Johnson isn't alone. Dozens of authors, musicians, illustrators, and other artists have discovered their work in similar datasets, and many are now suing. Susman Godfrey, the firm already representing authors in a case against Anthropic, has become a rallying point for creators who see the legal fight as bigger than any single settlement. Johnson himself reached out to the firm after reading about its case, calling the lawsuit a stand-in for every artist who's ever tried to make something of their own.

The lawsuits target AI companies mostly on copyright grounds, though some raise other claims, like violations of terms of service. Timelines vary widely. Some cases have dragged on for years.

Others wrapped up fast. None of it has been simple.

Johnson is just one of the dozens of authors, musicians, illustrators, and artists of all stripes taking the fight against AI to the courts. The lawsuits they’ve filed have primarily targeted the companies on copyright grounds, though some have sought other avenues, like terms of service violations.

Why this matters

Johnson's lawsuit is a preview of the discovery fights coming for every company that trained on scraped text without asking. The Atlantic's searchable dataset did something most litigation can't: it let individual authors find their own books inside the machine, by name, in seconds. That's a different kind of evidence than aggregate claims about "millions of works." For builders and founders leaning on models trained on similar corpora, the legal exposure isn't hypothetical anymore, it's searchable and citable in a complaint.

The emotional-harm framing Johnson uses ("violated, shocked, alarmed") also signals where plaintiffs' lawyers are headed: not just copyright infringement, but personal injury language that could shift how courts and juries weigh damages. If more datasets get this kind of public, queryable treatment, expect more Johnsons, and more lawsuits with names attached instead of abstractions. For anyone fine-tuning on web-scraped text now, the question isn't whether your training data included pirated work.

It's whether someone's already built the tool to prove it.

Common Questions Answered

How did Kirk Wallace Johnson discover his books were used to train AI systems?

Johnson found his own books, including The Feather Thief and The Fishermen and the Dragon, listed in a searchable dataset that The Atlantic published showing works used to train AI systems. His nonfiction works, which took five to six years each to research and write, had been pirated and included in chatbot training data without his knowledge or consent.

What legal grounds are artists using in their AI training lawsuits?

Artists have primarily targeted AI companies on copyright grounds, though some have pursued other legal avenues such as terms of service violations. The lawsuits involve dozens of authors, musicians, illustrators, and other creative professionals seeking compensation for unauthorized use of their work.

Why is The Atlantic's searchable dataset significant for AI training litigation?

The Atlantic's searchable dataset allows individual artists to find their own specific works inside AI training data by name within seconds, providing concrete evidence of infringement. This type of targeted evidence is different from aggregate claims about millions of works and strengthens the legal cases against companies that trained models on scraped text without permission.

What is the broader legal exposure for AI companies that trained on scraped text?

Companies that trained AI models on scraped text without authorization face significant discovery fights and legal liability similar to Johnson's lawsuit. The ability to identify specific copyrighted works in training datasets means the legal exposure for these builders and founders is not hypothetical but represents a concrete and growing risk.

LIVE16:35Google Expands SynthID Watermark to Label AI Content