Cognition, the company behind the Devin AI software engineer, launched SWE-2 on Thursday, describing it as its most advanced coding model yet. In a blog post published September 10, the company said the new model pushes what it calls the Pareto frontier of capability and cost, scoring 50.0% on the FrontierCode 1.1 Main benchmark while costing 64% less than Fable 5.1, the Anthropic model that sits just one point ahead of it at 50.9%.

The launch lands barely a day after Cognition confirmed a $2 billion funding round at a $48 billion valuation, and it signals how aggressively the agentic coding startup is investing in its own model development rather than relying solely on third-party frontier models. For continuous coverage of the AI coding race and everything else moving the industry, bookmark AI Buzz Wire, your daily source for breaking AI news.

Where SWE-2 Lands on the Benchmarks

According to Cognition's published results, SWE-2 posts 50.0% on FrontierCode 1.1 Main, ahead of Grok 4.6 at 48.0% and GPT-5.6 Sol at 47.5%, and within striking distance of GPT-6 Astra, which leads the field at 53.3%. On the DeepSWE 1.1 benchmark, SWE-2 reaches 73.0%, second only to Astra's 74.1% among the models Cognition lists.

The strongest showing comes on Terminal-Bench 2.1, where SWE-2 scores 92.8%, the highest figure in Cognition's comparison table and ahead of Fable 5.1 at 91.4% and GPT-6 Astra at 89.9%. The picture changes on the harder Terminal-Bench 4, where SWE-2 manages 27.3% against 57.9% for GPT-6 Astra and 55.8% for Fable 5.1, a reminder that the model's efficiency-first design leaves headroom at the very top end of task difficulty.

Cognition sums up the positioning bluntly: on FrontierCode 1.1 Main and DeepSWE 1.1, SWE-2 beats its previous model SWE-1.7 and Grok 4.6 on both score and cost, matches GPT-5.6 Sol and Fable 5/5.1 at a fraction of their price, and comes within a few points of GPT-6 Astra at a quarter of the cost.

Post-Trained From Kimi K3 at Multi-Trillion Scale

SWE-2 is not built from scratch. Cognition post-trained the model from Kimi K3, a 2.8-trillion-parameter base model from Moonshot AI that had already undergone extensive reinforcement learning for agentic coding. On top of that foundation, Cognition says its RL process added five to six points on many benchmarks and shifted K3's entire cost-performance frontier.

The company says SWE-2 is the first time it has scaled reinforcement learning to the multi-trillion-parameter regime, building on the training infrastructure and recipe developed for SWE-1.7. The key addition is an RL algorithm that trains all reasoning-effort levels in a single run, rather than treating low, medium and high effort as separate training targets.

Cost Penalties Shape the Whole Frontier

A central idea in SWE-2's training is a linear cost penalty applied per effort level within a single RL run, with each penalty tuned to the local slope of the base model's Pareto frontier. Cognition says this approach, derived from first principles, advances the model's entire cost-performance frontier while preserving its shape and reflecting real user costs as directly as possible in training.

The company also detailed a length-weighted reward baseline it has used since SWE-1.6, which it credits with significantly stabilizing training, alongside improvements to RL rollout serving that include an online draft model to raise decoding throughput. With NVFP4 and FP8 kernels and quantization-aware training, Cognition says it reduced overall memory usage despite the base model having almost three times the parameters of its predecessor.

The training data pipeline expanded as well: Cognition tripled the number of RL environments, added instruction-following overlays, and built a flywheel powered by previous SWE-2 checkpoints that iteratively hardens the verifiers used to score the model.

Efficiency as Much as Intelligence

Cognition frames SWE-2's gains as closely tied to efficiency. Stronger engineering judgment, the company says, lets the agent write more complete solutions with fewer detours and redundant reads. On FrontierCode 1.1 Main, the medium-effort configuration of SWE-2 scores higher than SWE-1.7 while taking 58% fewer turns and costing 81% less.

That framing matters for Cognition's business. Devin is sold as an autonomous software engineer, and inference cost is a direct input to pricing and margins. A model that holds frontier-adjacent quality while cutting turns and tokens per task changes the economics of every Devin session the company serves.

Availability

SWE-2 is available starting today in Devin Desktop and Devin CLI, with a rollout on Devin Web and Fusion underway, according to the company. Cognition says users should see the model appear across surfaces over the coming days.

What It Means for the Coding Model Market

The release tightens an already crowded pack behind GPT-6 Astra. Cognition now owns a model that trades blows with offerings from Anthropic, OpenAI and xAI at markedly lower cost, and it controls the full stack from model to agent product. Rivals will face pressure to respond on price as much as on capability, particularly for the high-volume, medium-difficulty tasks where SWE-2's cost profile is strongest.

For buyers of AI coding tools, the practical takeaway is that frontier-level coding help no longer requires frontier-level pricing. For the rest of the industry, SWE-2 is evidence that post-training and infrastructure efficiency, not just base model scale, have become a decisive competitive lever.

---

Stay Ahead of AI

Get the latest AI news, analysis, and breakthroughs — all in one place.

Read more AI news →