Editorial illustration for Leading mathematicians call AI strong at calculation, weak at creative thought
AI Excels at Math but Fails at Creative Thinking
Leading mathematicians call AI strong at calculation, weak at creative thought
Timothy Gowers has a Fields Medal. Peter Sarnak holds the Eugene Higgins Chair at Princeton. Both have spent time testing large language models against real mathematical work, and both come back with the same split verdict: these systems are fast, capable calculators that stall the moment a problem demands something genuinely new.
The distinction matters because it cuts against the usual AI progress narrative, where bigger models simply get better at everything. Gowers and Sarnak aren't disputing that LLMs can execute known methods with speed and range. Their concern sits with what happens before the execution, the step where a mathematician has to guess which of thousands of possible approaches is worth pursuing at all.
That's not a calculation problem. It's a judgment problem, and judgment built on intuition doesn't show up cleanly in training data.
A DeepMind researcher, Tom Zahavy, has been chasing a related question from inside the industry itself, giving the failure mode a name borrowed from philosophy of science. Between the mathematicians' outside view and Zahavy's inside one, a clearer picture starts to form of exactly where today's models run out of road.
DeepMind researcher Tom Zahavy reached a similar conclusion. In his paper "LLMs Can't Jump," he pins the bottleneck on "manipulative abduction," the ability to invent new foundational assumptions with no linguistic precedent.
Why this matters
For anyone building on top of these models, Gowers and Sarnak's read is worth sitting with. Search-and-combine is genuinely useful: it can grind through known techniques faster than any grad student and surface results that would take a human weeks to verify. That's real value for research teams doing literature-heavy or computation-heavy work.
But the abstraction problem they describe, the failure to invent new frameworks rather than recombine old ones, points to a ceiling that more compute alone probably won't fix. If two mathematicians who watch these systems closely are converging on the same critique, founders pitching "AI mathematician" products should be specific about which part of the job gets automated. Calculation and search, yes.
The Grothendieck-style leap to a new abstraction, not yet. For researchers, the useful move is to treat current models as fast collaborators inside a search space someone else defined, not as sources of the next big conjecture. That distinction will matter a lot once funding decisions start hinging on it.
Common Questions Answered
Why do Timothy Gowers and Peter Sarnak believe large language models have limitations in mathematical work?
Gowers and Sarnak conclude that while LLMs excel at fast calculation and combining existing techniques, they struggle when problems require genuinely creative or novel thinking. Their testing revealed that these systems stall the moment a mathematical problem demands something fundamentally new rather than recombination of known approaches.
What does Tom Zahavy identify as the core bottleneck preventing LLMs from creative mathematical thinking?
In his paper "LLMs Can't Jump," Zahavy pins the bottleneck on "manipulative abduction," which is the ability to invent new foundational assumptions with no linguistic precedent. This limitation explains why LLMs cannot generate truly original mathematical frameworks and concepts.
How can research teams practically benefit from LLMs despite their creative limitations?
LLMs provide genuine value for literature-heavy or computation-heavy research work by grinding through known techniques faster than graduate students and surfacing results that would take humans weeks to verify. The search-and-combine capability, while not creative, represents real utility for accelerating verification and synthesis of existing mathematical knowledge.
What does the split verdict from leading mathematicians suggest about the AI progress narrative?
The mathematicians' findings contradict the typical AI progress narrative that assumes bigger models simply get better at everything across all domains. Instead, their research reveals a fundamental ceiling in LLM capabilities related to creative abstraction, suggesting that scaling alone cannot solve the problem of generating genuinely new mathematical frameworks.
Further Reading
- Fields Medalist Timothy Gowers says ChatGPT solved PhD-level math problems in under two hours - VnExpress International
- A recent experience with ChatGPT 5.5 Pro - Timothy Gowers’s Weblog
- OpenAI's math breakthrough played to AI's strengths - Understanding AI
- An AI solution to an 80-year-old problem has shocked mathematicians - The Conversation
- Can LLM generate interesting mathematical research problems? - arXiv