Merit AC
2026-09-10

Cognition ships SWE-2: near-frontier coding quality at roughly a quarter of GPT-6 Astra's cost

Post-trained from Moonshot's Kimi K3, the new Devin model scores within one point of Claude Fable 5.1 on Cognition's own benchmark while running about 64% cheaper -- a direct cost-per-task data point, not just a capability claim.

Cognition released SWE-2 on September 10, 2026, the next model behind its Devin coding agent, post-trained from Moonshot AI's 2.8-trillion-parameter Kimi K3. On Cognition's own FrontierCode 1.1 Main benchmark, SWE-2 scores 50.0% -- within one point of Anthropic's Claude Fable 5.1 -- while running roughly 64% cheaper at that score, and at about one-quarter the cost of GPT-6 Astra by Cognition's own figures. The headline technical feature is configurable reasoning-effort levels trained in a single reinforcement-learning run. It's rolling out immediately in Devin Desktop and CLI, with Web and Fusion to follow.

A cost claim stated as a benchmark, not a slogan

Cognition's own announcement is direct about availability: "SWE-2 is available starting today in Devin Desktop and CLI." The more consequential number is the cost-per-quality-point comparison -- Cognition is explicitly positioning this as a Pareto-frontier move, matching near-frontier output quality at a fraction of the price, rather than chasing the top benchmark score outright.

Why this is the exact tradeoff this site tracks

Cost-per-task at a given quality level is closer to what actually determines whether AI coding spend produces value than raw benchmark leaderboard position -- a vendor publishing its own cost comparison this explicitly gives engineering and finance leaders a concrete number to check against their own measured output, not just a marketing claim to take on faith.

Sources

← All news