Skip to main content
Grok 4.7 AI model pricing at $2/million tokens, underperforming rivals in benchmark tests, shown on a screen.

Editorial illustration for xAI's Grok 4.7 priced at USD 2 per million tokens, lags rivals in benchmarks

xAI's Grok 4.7 priced at USD 2 per million tokens, lags...

3 min read

xAI put a price tag on Grok 4.7 that looks more like a Chinese lab's rate card than a Western frontier model's. Two dollars per million input tokens, six dollars per million output tokens, that's the deal Elon Musk's company is offering for what it calls its most capable model yet for coding and knowledge work. The pitch is a bigger base model, longer reinforcement learning, and better self-checking of its own answers before it hands back a result.

The catch shows up once you check the scoreboard. Artificial Analysis runs an independent Intelligence Index, version 4.3.2, built from ten benchmarks, and Grok 4.7 doesn't come close to the top. Rivals from Anthropic and OpenAI are already shipping models under names like Claude Fable 5.1 and GPT-6, and both are priced well above xAI's new release.

Grok 4.7 is now live through the Grok API, in Cursor, and inside Grok Build, so developers can start testing the cost-versus-capability tradeoff for themselves. The numbers on agentic coding tasks, where xAI has pushed hardest, tell a rougher story than the price sheet suggests.

On Terminal-Bench 4.0, Grok 4.7 hits just 26 percent, versus 60 percent for GPT-6 Astra and 55 percent for Claude Fable 5.1. Even the cheaper DeepSeek V4.1 Flash edges past it at 27 percent.

Why this matters

The pricing tells its own story. When a Musk-backed lab with access to a 200,000-GPU cluster ships a model priced like a budget Chinese offering, that's not aggressive market strategy, that's a company pricing to the score it got. A 46 on Artificial Analysis's index isn't competitive with what Claude and GPT-6 are reportedly putting up, and no amount of "longer reinforcement learning" language in a press release changes what the benchmark shows.

For developers and founders building on Grok, the calculus is straightforward now: cheap tokens matter less if you're burning more of them to get comparable output, or worse, shipping code that needs more verification passes because the model's self-checking claims don't hold up under real workloads. For researchers, this is a useful data point on where scaling base models and RL training time actually plateaus versus where the frontier labs are pulling ahead through other means.

We'd watch whether xAI adjusts pricing further or pushes a faster follow-up. A gap this visible on a public index rarely survives a single quarter without a response.

LIVE20:18xAI's Grok 4.7 priced at USD 2 per million tokens, lags rivals in benchmarks