Skip to main content
Cognition's SWE-2 coding model, a powerful AI, shown on a screen with code, demonstrating its efficiency.

Editorial illustration for Cognition's SWE-2 Coding Model Matches Rival at 64% Lower Cost

SWE-2 Matches Rivals at 64% Lower Cost

4 min read

Cognition put a number on its newest coding model Tuesday: 50.0% on FrontierCode 1.1 Main, one point behind Fable 5.1, at 64% of the cost. The company, best known for its Devin coding agent, calls the model SWE-2, and it's built by post-training Moonshot AI's Kimi K3, a 2.8-trillion-parameter open model, with reinforcement learning.

That's a jump in scale from Cognition's last release. SWE-1.7 was post-trained from Kimi K2.7, a much smaller base. SWE-2 runs on a foundation almost three times the size, and Cognition says its RL recipe still found 5 to 6 points of headroom on top of K3 across several benchmarks, despite the bigger starting point.

There's a catch for anyone hoping to run it themselves. SWE-2 has no open weights and no standalone API. It only works inside Devin, currently on Desktop and CLI, with Web and Fusion support coming later. It's also Cognition's first model to offer selectable reasoning-effort levels, all trained together in a single RL run rather than as separate models.

Cognition, the company behind the Devin coding agent, has released SWE-2, its most capable coding model to date. SWE-2 is post-trained with reinforcement learning from Kimi K3, Moonshot AI’s 2.8T-parameter open model. Cognition reports a score of 50.0% on FrontierCode 1.1 Main, within 1 point of Fable 5.1 at 64% lower cost.

Why this matters

For teams choosing coding models, price-to-performance just moved again, and it moved fast. A 64% cost cut for near-parity results on FrontierCode is the kind of number that gets put in front of a CFO, not just a CTO. But we'd push back on treating this as a settled win.

SWE-2 isn't something you can pull down and run on your own servers, so "deployable" here means deployable through Cognition's own stack, on their terms, at their pricing. That's a meaningfully different proposition than an open weights release, and worth remembering before anyone rewrites their infrastructure plans around it.

The more interesting engineering story is the single RL run producing three selectable effort levels, each with its own cost penalty. That's a real technical move, not a marketing wrapper, and it hints at where post-training work on top of open base models like Kimi K3 is headed. For founders building on borrowed foundation models, the lesson is that the post-training layer, not just the base model, is becoming the competitive battlefield. Watch whether Cognition publishes reproducible benchmarks beyond its own reporting.

Common Questions Answered

How does SWE-2's performance compare to Fable 5.1 on FrontierCode 1.1 Main?

SWE-2 achieves a score of 50.0% on FrontierCode 1.1 Main, which is only one point behind Fable 5.1's performance. Despite this near-parity in results, SWE-2 accomplishes this at 64% lower cost, making it a significantly more cost-effective option for teams evaluating coding models.

What base model does Cognition use for SWE-2 and how does it differ from SWE-1.7?

SWE-2 is built by post-training Moonshot AI's Kimi K3, a 2.8-trillion-parameter open model, using reinforcement learning. This represents a substantial upgrade from SWE-1.7, which was post-trained from the much smaller Kimi K2.7, making the new model's foundation almost three times larger in scale.

What are the limitations of deploying SWE-2 according to the article?

SWE-2 is not available as a downloadable model that you can run on your own servers. Instead, deployment is limited to Cognition's own stack on their terms and at their pricing, which represents a meaningfully different proposition than having full control over a self-hosted solution.

Why is SWE-2's cost advantage particularly significant for enterprise adoption?

The 64% cost reduction while maintaining near-parity performance with Fable 5.1 is a compelling metric that appeals to both technical and financial decision-makers. This price-to-performance improvement is substantial enough to warrant CFO-level attention, not just CTO consideration, making it an attractive option for cost-conscious organizations.

LIVE02:13Cognition's SWE-2 Coding Model Matches Rival at 64% Lower Cost