Skip to main content
Chart showing GLM-5.3's 246-point jump on a benchmark, ranking second to Claude Opus. AI model performance comparison.

Editorial illustration for GLM-5.3 Jumps 246 Points on Key Benchmark, Ranks Second to Claude Opus

GLM-5.3 Surges 246 Points, Ranks Second to Claude

3 min read

Z.ai's GLM-5.3 just closed the gap with the best closed model on the market. On the GDPval-AA v2 benchmark, which measures agentic task performance, the model's Elo score climbed from 1,524 to 1,770. That's a 246-point jump, and it leaves GLM-5.3 trailing only Claude Opus 5, which sits at 1,855. Overall, on the Artificial Analysis Intelligence Index, GLM-5.3 scores 60 points, tying Kimi K3 for the top spot among open models and beating its own predecessor, GLM-5.2, by seven points.

Price is where the model makes its real case. Artificial Analysis puts GLM-5.3 at $0.68 per task, undercutting Kimi K3's $0.84 by 19 percent, even though it costs 1.5 times more than GLM-5.2's $0.44. The model is already live through Z.ai's API, so developers can test the agentic gains now.

The open weights are another matter. Z.ai says it's holding them back for roughly two weeks, and the reason ties directly to what the benchmarks show about the model's capabilities.

GLM-5.3, the AI model from Chinese startup Z.ai, scores 60 points on the Artificial Analysis Intelligence Index. That ties it with Kimi K3 for the top spot among open models, and it's seven points ahead of the previous GLM-5.2.

Why this matters A 246-point Elo jump on GDPval-AA v2 is the kind of number that would normally headline itself, but the fine print matters more here. Z.ai is closing the agentic-task gap with Claude Opus 5 while still pricing GLM-5.3 well below it, at $0.68 versus whatever premium Anthropic charges, which is the real pitch to developers weighing frontier capability against burn rate. The catch: that price tag is 1.5 times what GLM-5.2 cost, so "cheap" is relative, and the delayed release means none of this is testable in production yet.

For founders building agent pipelines on open models, the tie with Kimi K3 at the top of the open-model index is worth watching, but benchmark Elo and real-world reliability on messy, multi-step tasks are different animals. We'd treat this as a strong signal that Chinese labs are optimizing specifically for agentic workflows rather than general chat quality, and that's the metric to track once GLM-5.3 actually ships and outside teams can run it against their own task suites.

Common Questions Answered

How much did GLM-5.3's Elo score improve on the GDPval-AA v2 benchmark?

GLM-5.3's Elo score on the GDPval-AA v2 benchmark jumped 246 points, climbing from 1,524 to 1,770. This significant improvement measures the model's performance on agentic task execution, demonstrating substantial progress in Z.ai's latest release.

How does GLM-5.3 rank compared to Claude Opus 5 on agentic task performance?

GLM-5.3 ranks second to Claude Opus 5 on the GDPval-AA v2 benchmark, with an Elo score of 1,770 versus Claude Opus 5's 1,855. This represents a significant narrowing of the gap between Z.ai's model and the top-performing closed model on the market.

What is GLM-5.3's score on the Artificial Analysis Intelligence Index and how does it compare to other open models?

GLM-5.3 scores 60 points on the Artificial Analysis Intelligence Index, tying it with Kimi K3 for the top spot among open models. This score represents a seven-point improvement over GLM-5.3's predecessor, GLM-5.2.

How does GLM-5.3's pricing compare to Claude Opus 5 and its predecessor?

GLM-5.3 is priced at $0.68, significantly undercutting Claude Opus 5's premium pricing while offering comparable frontier capabilities. However, GLM-5.3's price is 1.5 times higher than GLM-5.2's cost, making the price increase a notable trade-off for developers.

LIVE16:12GLM-5.3 Jumps 246 Points on Key Benchmark, Ranks Second to Claude Opus