Editorial illustration for AI Pioneer Warns Synthetic Data Can't Capture Human Minds
Synthetic Data Won't Build Human-Like AI, Sutton Warns
AI Pioneer Warns Synthetic Data Can't Capture Human Minds
Richard Sutton doesn't think synthetic data will save the current approach to building large language models. The Turing Award winner, who co-wrote the standard textbook on reinforcement learning and mentored AlphaGo researcher David Silver, has spent decades arguing that AI progress comes from methods that scale with compute rather than from stuffing in human knowledge by hand. That argument, laid out in his 2019 essay "The Bitter Lesson," made him one of the field's most cited voices on what actually works over time.
Now Sutton is pointing that same argument at the labs racing to scale up LLMs. In a recent conversation announcing his new venture, Oak Lab, founded with former student Khurram Javeed, he laid out where he thinks the field went wrong. Large language models proved the Bitter Lesson right by scaling with compute and swallowing the internet whole.
But the internet, Sutton says, is finite, and the world it describes is not. That mismatch is where his critique of synthetic data begins.
Asked whether synthetic data could break through this bottleneck, Sutton doesn't mince words. "No, that's that's just a big mistake." The reasoning comes from the "Big World Hypothesis" that Javeed formulated and the group in Alberta has been working on for years.
Why this matters
Sutton's argument lands squarely on labs like OpenAI and DeepMind, which have leaned harder on synthetic data as they run short of fresh web text. If he's right that no dataset can stand in for another person's mind, that's a ceiling on how far current scaling strategies can go, not just a technical footnote. It also exposes a labor problem the industry doesn't talk about enough: synthetic pipelines still need human experts to judge what counts as good data, which means the bottleneck everyone hoped to engineer away just moves upstream.
For researchers, this is a reason to stop treating synthetic data generation as a solved pipeline and start asking who's actually curating it and how that scales. For founders building on top of frontier models, it's a signal that the next wall in LLM progress might not be compute or parameters but something closer to an old problem: getting enough good judgment into the loop. Sutton built the field's foundations.
When he calls a core industry bet a mistake, that's worth sitting with, not dismissing.
Common Questions Answered
Why does Richard Sutton believe synthetic data cannot solve the bottleneck in large language model development?
Sutton argues that synthetic data fundamentally cannot capture the complexity of human minds, based on the "Big World Hypothesis" that his research group in Alberta has been developing. He contends that no dataset can adequately stand in for another person's mind, which means synthetic data represents a ceiling on how far current scaling strategies can progress rather than a viable solution to data scarcity.
What is the "Bitter Lesson" and how does it relate to Sutton's current stance on synthetic data?
The "Bitter Lesson" is Sutton's 2019 essay arguing that AI progress comes from methods that scale with compute rather than from manually stuffing in human knowledge by hand. This foundational argument supports his current position that synthetic data, which attempts to artificially generate training material, represents a misguided approach that contradicts the proven scaling principles he outlined years ago.
What labor problem does the synthetic data pipeline create that the industry overlooks?
While synthetic data pipelines are promoted as a solution to data scarcity, they still require human experts to judge and evaluate what counts as good data. This hidden labor requirement undermines the efficiency gains that synthetic data is supposed to provide and represents a significant cost that the industry doesn't adequately discuss.
How does Sutton's warning about synthetic data specifically impact companies like OpenAI and DeepMind?
OpenAI and DeepMind have increasingly relied on synthetic data as they exhaust supplies of fresh web text for training large language models. If Sutton is correct that synthetic data cannot replicate human knowledge, these companies may have hit a fundamental ceiling on their current scaling strategies, forcing them to reconsider their approach to model development rather than simply generating more artificial training data.
Further Reading
- Using Synthetic Data for AI Training Is 'a Big Mistake' - Business Insider
- A Turing Award winner says the industry's fix for running out of training data is a mistake - The Next Web
- Rich Sutton and Khurram Javed: Why AI Models Stop Improving - Apple Podcasts / Sequoia Training Data
- Best Practices and Lessons Learned on Synthetic Data for AI - arXiv
- Leveraging Synthetic Data from Large Language Models to Improve Model Performance - University of Pennsylvania