Skip to main content
Grok 4.6 AI logo on a screen, symbolizing SpaceXAI's advanced model matching OpenAI's top performance and reduced price.

Editorial illustration for SpaceXAI's Grok 4.6 Matches Top OpenAI Model, Cuts Price

Grok 4.6 Matches GPT-5.6 Sol, Costs Less

SpaceXAI's Grok 4.6 Matches Top OpenAI Model, Cuts Price

4 min read

xAI released Grok 4.6 this week, and the numbers put it right in the middle of the current top tier of language models. The Artificial Analysis Intelligence Index gives it 61 points, the same score as OpenAI's GPT-5.6 Sol. Only two models beat it: Anthropic's Claude Opus 5 at 63 and Claude Fable 5 at 62. That's a five-point gain over Grok 4.5, its predecessor, and enough to close most of the gap with Anthropic's best work.

The bigger story is what xAI is charging for it. Grok 4.6 holds steady at $2 per million input tokens and $6 per million output tokens, pricing that undercuts Claude Opus 5's $5/$25 and GPT-5.6 Sol's $5/$30 by more than 60 percent. For a model scoring within two points of the market leader, that's a real shift in what buyers get for their money.

xAI is also pointing to agentic performance, where Grok 4.6 handles multi-step tasks with fewer steps than its closest competitor. The model is live now through xAI's API, Cursor, Grok Build, and partners including OpenRouter, Vercel, and Cloudflare.

Grok 4.6 catches up to frontier models while costing far less. According to the Artificial Analysis Intelligence Index, SpaceXAI's new model scores 61 points, tying OpenAI's GPT-5.6 Sol.

Why this matters

For developers building agentic products, Grok 4.6 changes the calculus on which model sits behind the workflow. A five-point jump on the Intelligence Index is real progress, and ranking second on GDPval-AA v2 suggests SpaceXAI has actually put engineering effort into the boring, expensive part of AI, getting models to complete multi-step tasks without falling over. Tying GPT-5.6 Sol while undercutting it on price is the kind of detail that matters more to a founder watching API bills than a leaderboard position ever will.

We'd still push back on treating benchmark scores as the whole story: Claude Opus 5 and Claude Fable 5 still score higher, and a two-point Intelligence Index gap can mean very different things depending on your actual use case. But for teams running high-volume agentic pipelines where cost per task compounds fast, a model that matches frontier performance at a lower price is worth testing directly against your own workloads, not just against Artificial Analysis's leaderboard. Watch how OpenAI and Anthropic respond on pricing next.

Common Questions Answered

How does Grok 4.6's performance compare to other top language models according to the Artificial Analysis Intelligence Index?

Grok 4.6 scores 61 points on the Artificial Analysis Intelligence Index, tying OpenAI's GPT-5.6 Sol for third place among top models. Only Anthropic's Claude Opus 5 (63 points) and Claude Fable 5 (62 points) rank higher, making Grok 4.6 competitive with the frontier models in the current market.

What performance improvement did Grok 4.6 achieve over its predecessor Grok 4.5?

Grok 4.6 gained five points on the Intelligence Index compared to Grok 4.5, representing significant progress in closing the gap with Anthropic's best models. This five-point jump demonstrates meaningful engineering improvements in the model's capabilities.

Why is Grok 4.6's pricing strategy significant for developers building agentic products?

Grok 4.6 undercuts OpenAI's GPT-5.6 Sol on price while matching its performance score, fundamentally changing the cost-benefit analysis for developers choosing which model to use in their workflows. This pricing advantage combined with strong performance on multi-step task completion makes it an attractive alternative for founders building agentic applications.

What does Grok 4.6's ranking on GDPval-AA v2 indicate about SpaceXAI's engineering focus?

Grok 4.6's second-place ranking on GDPval-AA v2 suggests that SpaceXAI has invested significant engineering effort into the expensive and complex work of enabling models to complete multi-step tasks reliably. This indicates the company is focused on practical, production-ready capabilities rather than just raw benchmark scores.

LIVE21:45Twitch streamers can now opt out of Amazon AI training