Skip to main content
Gemini 2.5 Pro (K2.5) beats GPT-5.2 and Claude Opus 4.5 in AI benchmarks, showing cost-effective agentic and video performanc

Editorial illustration for K2.5 Beats GPT-5.2 and Opus 4.5 on Agentic and Video Benchmarks, Cuts Costs

K2.5 Beats GPT-5.2 and Opus 4.5 on Agentic and Video...

Updated: 3 min read

Bigger models are no longer better models. A company called Moonshot just proved it with K2.5, an open-source release that beats OpenAI's GPT-5.2 and Anthropic's Opus 4.5 on the two things that were supposed to be walled gardens for giants: agentic tasks and video reasoning. It's cheaper, too.

The details: K2.5 tops GPT-5.2 and Opus 4.5 in on key benchmarks for agentic tasks and video reasoning, though it trails slightly on pure coding evals. K2.5 shows massive cost savings over top rivals, is natively multimodal, and comes in as the top open model on Artificial Analysis' leaderboard. The model also features Agent Swarm, allowing K2.5 to manage up to 100 AI sub-agents running tasks at once across up to 1,500 steps and tools. Moonshot also open-sourced Kimi Code, an agentic coding agent that works in terminals and IDEs like VSCode and Cursor.

The coding lag is a footnote. What matters is the cost and the context. Agent Swarm means you can run a hundred specialized sub-agents in parallel, a chaotic but potentially powerful way to solve complex problems.

And by open-sourcing Kimi Code, Moonshot is planting its flag directly in the developer's workspace. This isn't about winning a single benchmark. It's about making the expensive, proprietary approach to advanced AI look suddenly fragile.

The economics of building with this stuff just tipped.

Common Questions Answered

How does K2.5 outperform GPT-5.2 and Opus 4.5 on agentic and video benchmarks?

K2.5 demonstrates superior performance on agentic tasks and video reasoning, two capabilities that were previously considered exclusive strengths of larger proprietary models from OpenAI and Anthropic. The open-source model achieves this while maintaining lower operational costs, challenging the assumption that bigger models are inherently better for these specialized tasks.

What is Agent Swarm and how does it enable parallel processing in K2.5?

Agent Swarm is a capability that allows K2.5 to run up to a hundred specialized sub-agents in parallel simultaneously. This approach enables a chaotic but potentially powerful method for solving complex problems by distributing tasks across multiple specialized agents working concurrently.

Why is Moonshot's open-sourcing of Kimi Code significant for developers?

By open-sourcing Kimi Code, Moonshot is directly positioning its technology in the developer's workspace, making advanced AI capabilities more accessible and affordable. This move challenges the expensive, proprietary approach to advanced AI that has been dominated by larger companies, potentially disrupting the economics of AI development.

What cost advantages does K2.5 offer compared to GPT-5.2 and Opus 4.5?

K2.5 is significantly cheaper to operate than both GPT-5.2 and Opus 4.5 while delivering superior performance on agentic and video reasoning benchmarks. This cost efficiency, combined with its open-source nature, makes advanced AI capabilities more economically viable for developers and organizations.

LIVE20:07Nimble's New Web Search Agents Cut AI Token Costs by Half