Skip to main content
GPT-6 Astra AI solving complex Erdős problems with a limited budget, showcasing advanced computational power.

Editorial illustration for GPT-6 Astra solves two open Erdős problems with USD 300 budget

GPT-6 Astra Solves Erdős Problems on $300 Budget

4 min read

OpenAI released GPT-6 Astra this week, and the model has already split the two firms that track frontier AI performance for a living. Epoch AI, which aggregates more than 50 benchmarks into a single score, ranks Astra first among 267 models tested. Artificial Analysis, running its own mix of knowledge, coding, and comprehension tests, scores it dead even with its predecessor Sol and behind Anthropic's Claude Fable 5.1.

That gap alone would be worth a story. What's pulling attention away from the scoreboard fight is ARC-AGI-3, the reasoning benchmark built by François Chollet's ARC Prize team. For the first time, a model on that test is completing tasks with less effort than the average human tester needs.

Chollet, who has spent years pushing back on inflated AGI timelines, says the jump caught him off guard, fast enough that he's now revising his own forecast for when machines catch up to people on the metric he built specifically to resist gaming. What he said about the pace of that shift is below.

The biggest surprise comes from ARC-AGI-3, where Astra works more efficiently than the average human for the first time. ARC Prize chief François Chollet calls the progress "2x faster" than he expected and is moving up his AGI forecast.

Why this matters

The $300-versus-$220,000 split is the story here, not the two solved Erdős problems. Epoch AI's own numbers show that once you strip out the marketing-friendly cheap run and look at what it actually took to squeeze three more proofs out of Astra, the "efficiency" narrative gets shakier. That gap matters for anyone budgeting compute against benchmark claims: a model that's cheap in a curated demo and absurdly expensive in unconstrained use isn't the same product.

Artificial Analysis and Epoch AI can't even agree on whether Astra beats its predecessor, which should make founders wary of leaning on any single leaderboard when pricing out deployment. Chollet's ARC-AGI-3 read, that Astra beats average humans on efficiency and that this arrived twice as fast as he expected, is the number worth tracking, because forecast revisions from someone who built the benchmark carry more weight than vendor-selected wins. Watch for independent, non-cherry-picked runs on FrontierMath Erdős before treating "$300 budget" as anything more than a headline.

Common Questions Answered

How does GPT-6 Astra perform on the ARC-AGI-3 benchmark compared to human efficiency?

GPT-6 Astra works more efficiently than the average human for the first time on ARC-AGI-3, marking a significant milestone in AI development. ARC Prize chief François Chollet described this progress as "2x faster" than expected and has moved up his AGI forecast as a result of this breakthrough.

Why do Epoch AI and Artificial Analysis disagree on GPT-6 Astra's ranking?

Epoch AI ranks Astra first among 267 models tested using its aggregated benchmark of more than 50 tests, while Artificial Analysis scores it equal to its predecessor Sol and behind Anthropic's Claude Fable 5.1 using its own mix of knowledge, coding, and comprehension tests. The disagreement highlights how different benchmark methodologies can produce significantly different performance assessments.

What is the significance of the $300-versus-$220,000 cost difference mentioned in the article?

The $300 cost represents a curated demo run, while $220,000 reflects the actual expense of unconstrained use to solve the two Erdős problems. This gap is critical for compute budgeting because a model that appears cheap in marketing-friendly demonstrations may be prohibitively expensive in real-world applications, making them fundamentally different products despite similar benchmark claims.

What are the two open Erdős problems that GPT-6 Astra solved?

While the article headline references that GPT-6 Astra solved two open Erdős problems, the specific problems are not detailed in the provided excerpt. The article focuses more on the efficiency metrics and benchmark performance rather than the mathematical problems themselves.

LIVE15:44OpenAI Agents Accessed German Wiki to Share Sandbox Exploits, Logs Show