Editorial illustration for OpenAI's Astra Skips Math Optimization Despite Earlier Focus
OpenAI's Astra Tops Math Benchmark Despite Skip
OpenAI's Astra Skips Math Optimization Despite Earlier Focus
OpenAI's newest model, GPT-6 Astra, just posted the top score on ErdosBench, a set of 226 open math problems built by ulam.ai around the tradition of Erdős-style questions. Astra solved 106 of them, closed out 43 completely, and disproved 27 others, landing a benchmark score of 3.23. Benchmark creator Przemek Chojecki pegged the jump at roughly 5 to 10 percent across various math-research skills compared to the field's previous leader, Sol, which managed 78 solved problems at maximum reasoning settings. Astra also wrote up its results with less exaggeration, in some cases underselling what it had actually proven.
That result lands oddly given what OpenAI has said about its own priorities. The company built its first Astra announcement around math achievements, yet chief scientist Jakub Pachocki now says math research wasn't where the team focused its effort this time. Mathematicians who've spent the past stretch of time nervously watching model releases for signs their field was about to get automated away can exhale, at least for now. The gap between what OpenAI chose to build and what Astra turned out to be capable of raises its own questions about how the company is thinking about progress toward AGI.
OpenAI's GPT-6 Astra tops the ErdosBench for open math problems, even though chief scientist Jakub Pachocki says the company didn't prioritize math. That says something about where "AGI" actually stands.
Why this matters
Pachocki's admission is the real story here, not ErdosBench's score. If Astra tops open math problems as a side effect of general capability, without OpenAI tuning for it, that's a bigger signal about model progress than any benchmark headline. For researchers and founders building on top of these models, it means the old playbook of chasing narrow benchmark wins is giving way to something harder to plan around: capability that shows up sideways, unannounced, wherever it happens to land.
That's good news if you're a mathematician hoping for tools you didn't ask for, and unsettling if you're trying to forecast what a lab will ship next. We'd push back on reading this as proof AGI is near. It's proof that OpenAI itself doesn't fully control, or maybe doesn't fully understand, where its own models get strong.
Watch what Pachocki's team decides to optimize for next. That choice will tell us more about their roadmap than any leaderboard will.
Common Questions Answered
What score did GPT-6 Astra achieve on ErdosBench and how does it compare to previous models?
GPT-6 Astra achieved a benchmark score of 3.23 on ErdosBench by solving 106 of the 226 open math problems, closing out 43 completely, and disproving 27 others. This represents approximately a 5 to 10 percent improvement across various math-research skills compared to Sol, the previous leader, which managed 78 solved problems at maximum reasoning.
Why is OpenAI's lack of math optimization focus significant according to Jakub Pachocki?
Chief scientist Jakub Pachocki stated that OpenAI didn't prioritize math optimization for GPT-6 Astra, yet the model still topped ErdosBench. This suggests that strong mathematical capabilities emerged as a side effect of general capability improvements rather than targeted tuning, indicating broader progress toward AGI than narrow benchmark optimization would suggest.
What is ErdosBench and who created it?
ErdosBench is a benchmark set consisting of 226 open math problems built by ulam.ai around the tradition of Erdős-style questions. The benchmark was created by Przemek Chojecki to evaluate AI models' performance on complex mathematical research problems.
How does the emergence of unplanned capabilities in GPT-6 Astra change the approach for researchers and founders?
The unplanned emergence of strong math capabilities in GPT-6 Astra signals that the traditional playbook of chasing narrow benchmark wins is becoming obsolete. Researchers and founders building on these models now face a harder challenge to plan around: capabilities that show up unexpectedly across various domains without deliberate optimization.
Further Reading
- OpenAI's GPT-6 Astra on ARC-AGI-3 | Hacker News - Hacker News
- OpenAI's GPT-6 Astra scores 97.6% on FrontierMath after internal version started at just 17% - CryptoBriefing
- OpenAI's Astra solves 10 long-open math problems and publishes the proofs - SiliconANGLE - SiliconANGLE
- Announcing FrontierMath Erdős - Epoch AI - Epoch AI
- FrontierMath Erdős - Epoch AI