
Cognition SWE-2: Frontier-Adjacent Coding at 64% Off, and the Benchmark Trap Underneath
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-10$0.15 / $0.60 per 1M tokens
- openaiNEWOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleNEWGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenNEWQwen: Qwen3.8 Max (0902)2026-09-0240Intelligence72Coding
- anthropicNEWAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.24 / $0.73 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0340Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2451Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2134Intelligence69Coding
Cognition SWE-2 shipped this week with a number built to travel: 50.0% on FrontierCode 1.1 Main, within a point of Claude Fable 5.1 at 50.9%, for 64% less money. The model is a post-train of Moonshot AI's Kimi K3 — the 2.8-trillion-parameter open-weight base — and Cognition says its reinforcement learning added 5–6 points across several benchmarks. Both figures are the vendor's own, and neither has been independently reproduced. The second one, though, is the claim that should change how you read the first, because "plus five or six" turns out to describe nearly everything SWE-2 does, including the places where it loses badly.
That is not a knock. It is the most useful thing about this release, and it is the part almost none of the launch coverage says out loud.
What Cognition actually shipped
SWE-2 is Devin's coding model, not a general-purpose API model with a price list. It shipped available starting today in Devin Desktop and CLI, in Cognition's own words, with rollout underway on Devin Web and Fusion. Third-party coverage dates the announcement to September 10, 2026; the byline on Cognition's own blog post reads 09.11.26. Either way it is days old. The weights are proprietary and there is no published per-token rate card. If you want to use SWE-2, you use it inside Devin.
The base is Kimi K3, a 2.8-trillion-parameter Mixture-of-Experts model with roughly 104B parameters active per token — about three times the size of SWE-1.7's base. Cognition notes K3 had already been heavily RL-trained for agentic coding before it got there, and that its own RL "still finds substantial headroom." That mirrors the pattern from SWE-1.7, which was post-trained on a Kimi base as well.
Here is the detail that matters more than any single benchmark score: FrontierCode is Cognition's benchmark. The company writes the tasks — with 20-plus open-source maintainers, more than 40 hours per task, each maintainer defining what "mergeable" means in their own repo — runs the leaderboard, and grades the submissions. It is a serious piece of evaluation design, and it is also a house instrument. Cognition reports each model's best score across reasoning-effort settings, and by its own methodology appendix the open-weight comparison models, including Kimi K3, were run inside Cognition's Devin CLI. None of that makes the numbers wrong. All of it belongs in the sentence where you repeat them.

Reading the benchmark table honestly
Four benchmarks, all vendor-reported. Here is what each one actually says, one line per row.
• FrontierCode 1.1 Main — SWE-2 50.0%, against Kimi K3 44.2%, Grok 4.6 48.0%, GPT-5.6 Sol 47.5%, Claude Fable 5.1 50.9%, GPT-6 Astra 53.3%, SWE-1.7 42.0%. Within a point of the frontier, ahead of everything else in the table.
• DeepSWE 1.1 — SWE-2 73.0%, ahead of Claude Fable 5.1 at 67.4% and Grok 4.6 at 67.5%, behind GPT-6 Astra at 74.1%. This is the row where SWE-2 most credibly beats a frontier model outright.
• Terminal-Bench 2.1 — SWE-2 92.8%, the top score in Cognition's table, ahead of Claude Fable 5.1 at 91.4% and GPT-6 Astra at 89.9%. This is the number in most of the launch headlines.
• Terminal-Bench 4 — SWE-2 27.3%, against Claude Fable 5.1 at 55.8% and GPT-6 Astra at 57.9%. A gap of roughly 28 points, and the row that gets quoted least.
The obvious objection is that everyone drops on Terminal-Bench 4 — it is a harder, longer-horizon successor, so of course scores fall. The interesting part is by how much, and the fall is not proportional. Going from Terminal-Bench 2.1 to 4, Claude Fable 5.1 loses about 35.6 points and GPT-6 Astra about 32.0. SWE-2 loses about 65.5, and its base Kimi K3 about 66.8. So on long-horizon agentic work, SWE-2 does not merely sit below the frontier models — it degrades roughly twice as fast as they do when the task runs long.
Whether Terminal-Bench 4 is the "live" version is worth stating plainly: it is the current edition, and 2.1 is the superseded one. The row where SWE-2 leads is the retired benchmark. The row where it trails is the current one. Both facts are true, and the launch coverage tended to pick one.
"Kimi K3 plus five" — the heuristic that predicts the scores nobody published
Line the four benchmarks up against SWE-2's own base model and something almost mechanical falls out. Against Kimi K3, SWE-2 gains 5.8 points on FrontierCode 1.1 Main, 4.5 on DeepSWE 1.1, 4.5 on Terminal-Bench 2.1, and 5.8 on Terminal-Bench 4.
Four benchmarks, four gaps, all between 4.5 and 5.8. The post-training moved the level, not the shape. Wherever K3 is strong, SWE-2 is strong plus about five. Wherever K3 is weak — Terminal-Bench 4 being the clear case — SWE-2 is weak plus about five.
That gives you a working forecast for any task Cognition did not benchmark, which is most of them. Look up how Kimi K3 performs on your workload, add five points, and you have a better estimate of SWE-2 than any of the headline numbers will give you. It also explains the release's actual thesis: Cognition is not claiming to have beaten the frontier. It is claiming to have moved the whole cost-performance curve outward while staying a fixed distance behind the best model on whatever axis you measure.
The 64% is real, but you cannot shop it
"64% cheaper than Claude Fable 5.1" is a true statement about a specific measurement — and the measurement is mean US dollars spent per rollout on FrontierCode, not a token price. There is no per-token rate for SWE-2 because there is no SWE-2 API. The cost claim describes what a task costs inside Devin, which bundles the model, the harness, the scaffolding and Cognition's serving stack into one invoice.
That distinction has teeth. You cannot compare SWE-2's cost against a token price from another vendor, because they are not the same unit. And you cannot take SWE-2 elsewhere: proprietary weights, one vendor, one agent product. The "up to 70% lower cost" figure in some coverage is the top of the range Cognition quotes across benchmarks; the FrontierCode 1.1 Main figure is 64%.
The efficiency numbers are, if anything, the more practical story, and those are vendor-reported too. SWE-2 at medium effort beats SWE-1.7's FrontierCode score while using 58% fewer turns and costing 81% less on average. Mean steps per run drop from 127 on SWE-1.7 to 53 at medium effort (80 at high, 98 at max), and the median first real code edit moves from step 48 to step 18. That is a direct answer to the complaint that SWE-1.7 over-explored simple tasks, and it is the kind of improvement that shows up on your bill long before a benchmark gap of one point does.

The training trick is the actual news
Strip away the leaderboard positioning and there is one genuinely novel piece of engineering here, which is why the release is worth reading even if you never open Devin.
Cognition trained all three reasoning-effort levels — medium, high and max — in a single RL run. The mechanism is a linear cost penalty per effort level, each one tuned to match the slope of the base model's own cost-performance Pareto curve at that point. Cognition's argument for why the penalty has to be linear is the interesting bit: only a linear penalty keeps the training objective a function of average cost and solve rate alone, rather than something that depends on the shape of the distribution.
The usual approach trains effort levels separately, or bolts a cheaper checkpoint on afterwards. Training them together inside one run is what lets the company claim it advanced the whole frontier rather than one point on it. Supporting changes: a length-weighted reward baseline, tripling the number of RL environments, instruction-following overlays, and a flywheel that uses earlier SWE-2 checkpoints to harden the verifiers against reward hacking. On the serving side, NVFP4 and FP8 kernels with quantization-aware training, a DSpark speculative-decoding draft model retrained for 15% longer accept lengths, and a prefill delayer worth 10–20% on throughput per GPU at the cost of worse time-to-first-token.
If you run post-training, that recipe is the transferable part of this release. If you do not, the takeaway is simpler: the reason SWE-2 is cheap is not that it is a smaller model. It is that the cost of running it was written into the reward function.
If you want the Kimi K3 capability outside Devin
SWE-2 is not on OrcaRouter, and it will not be — proprietary weights, no API, Devin-only distribution. Its base model is a different situation. Kimi K3 is routed on OrcaRouter as kimi/kimi-k3, at $3.00 per 1M input tokens and $15.00 per 1M output tokens with no markup over the provider rate, a 1M-token context window, native tool calling, image input, and reasoning depth controlled through a top-level reasoning_effort parameter instead of sampling settings.
That is the honest version of this release's practical takeaway. SWE-2 shows you what a well-executed RL pass on top of K3 looks like at the top of its range — a genuine 4.5-to-5.8-point lift, plus large efficiency gains, at a real cost. What it does not give you is a model you can call. If you want the underlying capability pointed at your own code, your own harness or your own pipeline, K3 is the part of this stack you can actually reach, and it is reachable through one OpenAI-compatible endpoint.
It is also the safer half of the bet while the numbers are unaudited. SWE-2's benchmarks are Cognition's own and nobody outside the company has reproduced them. Kimi K3's are the ones the comparison actually rests on — and if you route through OrcaRouter you can put a fallback chain behind it, so a model you are trialling does not become a production dependency the day it disappoints you.

Who should act, and what would settle it
If you already pay for Devin, SWE-2 is straightforwardly good news: it is rolling out to your product, it is cheaper per task than what it replaces, and the efficiency numbers suggest it will waste fewer turns on easy work. Try it on the simple end of your queue first — the effort-level design means medium is where the cost win lives.
If you are choosing a coding model from outside, wait for independent evaluation. There is none yet: Artificial Analysis has no SWE-2 entry, and the FrontierCode leaderboard is Cognition's own instrument. Three things would settle it. A third-party run of SWE-2 on Terminal-Bench 4, which is the row that decides whether this is a frontier-adjacent model or a fast one. Any independent reproduction of the FrontierCode gap. And a per-token price, which would require Cognition to sell the model outside Devin — the single change that would make the 64% comparable to anything else you are quoted.
Until then the fair summary is the one the base-model gap already gives you. SWE-2 is Kimi K3 plus five points, at a cost Cognition designed into the reward function. That is a real achievement in training and a genuinely good deal inside Devin. It is not yet a frontier model, and the one benchmark where it appears to beat everything is the one that stopped counting.
Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
