Editorial illustration for Qwen3.8 Max Catches Claude Opus in Test, But Uses 64 Steps and 15x More Tokens
Qwen3.8 Max Matches Claude Opus, But Costs 15x More Tokens
Alibaba's Qwen3.8 Max landed a 10-point jump on the Artificial Analysis Intelligence Index, climbing from 46 to 56 and pulling even with Claude Opus 4.8 in the process. That's the headline number. Artificial Analysis also puts the new model ahead of GLM-5.2, which sits at 51, and on GDPval-AA, a benchmark built around work-related tasks, Qwen3.8 Max gained 468 Elo points to hit 1,739, enough to pass Kimi K3's 1,685.
But scores alone don't tell you what it took to get there. Alibaba cut list prices on the new model, with input tokens dropping from $2.50 to $2.00 per million and output falling from $7.50 to $6.00. Cache hits got cheaper too, down to $0.25 from $0.50. On paper, that reads as a win twice over: higher intelligence, lower rates.
The actual cost per task tells a different story, and so does the number of steps the model needs to reach its answer. Kimi K3 still beats Qwen3.8 Max on the Intelligence Index by a single point, and it does so while costing 25 percent less per task. What's driving that gap is worth a closer look.
Alibaba's Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index, a 10-point jump over Qwen3.7 Max (46). According to Artificial Analysis, that puts it on par with Claude Opus 4.8 and ahead of GLM-5.2 (51), but behind Kimi K3 (57), which also runs 25 percent cheaper.
Why this matters
The Intelligence Index score is the headline Alibaba wants, and it got it: 56, tied with Claude Opus 4.8. But the 64-step, 15x-token cost to get there tells the real story. Kimi K3 hits 57 for 25 percent less, without the extra steps.
For anyone shipping products on these APIs, that gap between benchmark score and actual bill is the number worth tracking, not the leaderboard rank. GDPval-AA looks more encouraging for Qwen3.8 Max, a 468-point jump to 1,739 that beats Kimi K3 on work-related tasks specifically, which suggests the model's strength is task completion rather than efficient reasoning. That's a useful distinction for founders picking models: raw intelligence scores and cost-adjusted performance are now diverging enough that vendors will lean on whichever number flatters them.
Artificial Analysis publishing both cuts through that. Our advice: check the step count and token multiplier before trusting any "on par with Opus" claim, because thoroughness and value are clearly not the same metric anymore.
Common Questions Answered
How does Qwen3.8 Max's performance compare to Claude Opus 4.8 on the Artificial Analysis Intelligence Index?
Qwen3.8 Max achieved a score of 56 on the Artificial Analysis Intelligence Index, matching Claude Opus 4.8's score and representing a 10-point improvement over the previous Qwen3.7 Max model which scored 46. This tied performance places Qwen3.8 Max ahead of GLM-5.2 (51) but behind Kimi K3 (57) on the same benchmark.
What is the computational cost difference between Qwen3.8 Max and its benchmark competitors?
Qwen3.8 Max requires 64 reasoning steps and uses 15 times more tokens than competing models to achieve its benchmark score, making it significantly more computationally expensive despite matching Claude Opus 4.8's Intelligence Index score. In contrast, Kimi K3 achieves a higher score of 57 while running 25 percent cheaper and without requiring the extra inference steps.
How did Qwen3.8 Max perform on the GDPval-AA work-related tasks benchmark?
On the GDPval-AA benchmark focused on work-related tasks, Qwen3.8 Max achieved a score of 1,739 Elo points, representing a 468-point improvement over its predecessor. This performance was sufficient to surpass Kimi K3's score of 1,685 on the same work-related benchmark.
Why is the token cost more important than the Intelligence Index score for API users?
For companies deploying these models through APIs, the actual computational cost measured in tokens and inference steps directly impacts their operational expenses, making it more relevant than leaderboard rankings. While Qwen3.8 Max ties with Claude Opus 4.8 on the Intelligence Index at 56, the 15x token overhead means users will pay significantly more per query despite achieving equivalent benchmark scores.
Further Reading
- Alibaba says Qwen3.8-Max coded autonomously for 16 days - InfoWorld
- Qwen3.8-Max arrives with a bold claim: it outperforms GPT-5.6 Sol and Fable 5 on agentic computer use - VentureBeat
- Qwen 3.8 Benchmarks: What Alibaba's Table Shows, and What It Doesn’t - Apidog
- How Alibaba’s New Qwen3.8-Max Stacks Up Against US AI Giants - Mitrade
- Qwen 3.8 vs Claude Opus 4.8: Raw Scale vs Reasoning - OrcaRouter