Editorial illustration for GPT-6 Astra gains 4 points in revised Artificial Analysis index, still trails Claude Fable
GPT-6 Astra gains 4 points in revised Artificial...
Artificial Analysis pushed out version 4.2 of its Intelligence Index this week, four points better than before for GPT-6 Astra, the OpenAI model whose earlier score had become something of a running argument in AI benchmarking circles. Epoch AI had ranked Astra first among 267 models tested, with 169 points spread across more than 50 benchmarks. ARC-AGI-3 showed jumps ranging from large to very large depending on which harness ran the test. Artificial Analysis, by contrast, had scored Astra roughly even with its predecessor, a result that drew pushback from people who felt the firm's methodology was missing something obvious about the model's actual capabilities.
The revised index still puts Astra behind Anthropic's Claude Fable 5.1, which holds the top spot, with Meta close behind in third. But the update touches more than one model's score. Artificial Analysis added two benchmarks, dropped one that models had effectively solved, and shifted how much weight private test data carries in the final ranking. The company says it normally avoids changes like this during major launches, but felt it had no choice this time.
Artificial Analysis has released version 4.2 of its Intelligence Index, likely in response to criticism that its benchmarks failed to capture GPT-6 Astra's actual progress.
Why this matters
Benchmark churn like this should make developers and founders nervous, not reassured. A four-point bump for Astra after a full index overhaul looks less like a correction and more like a scoreboard adjusted to match expectations set by OpenAI's own numbers and Epoch AI's 169-point ranking. When a firm revises its methodology right after a model underperforms relative to competing evaluations, the timing itself becomes part of the story, regardless of what the new numbers say.
Claude Fable 5.1 still tops the list, Astra sits second, Meta third, and Astra's edge in token efficiency is genuinely worth watching if you're optimizing for cost per task. But the bigger lesson for anyone picking models for production is that leaderboard position is only as stable as the rubric behind it. Before you swap architectures based on a single index jump, check whether the benchmark changed or the model did.
Watch whether Artificial Analysis publishes the specific weighting changes behind version 4.2, and whether Epoch AI's ranking moves at all in response.
Further Reading
- GPT-6 Astra vs Claude Fable 5.1 - Release Intelligence ... - Artificial Analysis
- GPT-6 Astra: Release Intelligence, Performance & Price - Artificial Analysis
- Claude Fable 5.1 tops the Artificial Analysis Intelligence Index - Artificial Analysis
- GPT-6 Astra's Independent Benchmarks Land Flat on General ... - FourWeekMBA
- Why GPT-6 Astra bei Artificial Analysis schlechter als GPT ... - all-ai.de