
Cognition SWE-2: Who Makes It, What Harness It Runs In, and What the Numbers Actually Say
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 517 tok/s
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 196 tok/s
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1327 tok/s
- deepseekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- tencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 111 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 221 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
Cognition SWE-2 is a coding model made by Cognition, the company behind Devin, and it is served inside Devin rather than over a model API: Cognition's own launch post says SWE-2 is available in Devin Desktop and Devin CLI first, with rollout to Devin Web and Fusion. It is post-trained from Kimi K3, Moonshot AI's 2.8-trillion-parameter model. That is the whole answer to the question most people are actually asking when they arrive here, and it is worth stating before anything else, because a measurable share of the searches that lead to this page add a fourth word — Devin — and attribute the model to Devin rather than to Cognition. Devin is the agent Cognition sells. SWE-2 is the model Cognition trained to run inside it.
This page is a reference page rather than a launch story. SWE-2 shipped on September 10, 2026, which is more than a week behind us; the date is here to fix the model in time, not to report an event. What follows is what is documented, who documented it, and what nobody has checked yet.
Who makes SWE-2, and why the attribution keeps drifting
The confusion is not unreasonable. Cognition does not sell SWE-2 the way a lab sells an API model. There is no model card with a context window, no published token price, no identifier you can pass in a request body. There is a product — Devin — and the model arrives as part of it. Cognition's own prose reinforces the collapse: the launch post is written from Devin's point of view, the benchmark table is unexplained to anyone who has not already read the company's earlier FrontierCode work, and the availability line names Devin four times and Cognition once — in the byline.
So, plainly: Cognition makes SWE-2. Cognition makes Devin. The two are not the same thing, but you cannot currently have one without the other. This is also not a one-off naming accident. Cognition has shipped a run of software-engineering models under the SWE prefix — SWE-1.5, SWE-1.6, SWE-1.7 and now SWE-2, with an SWE-1.6 Preview checkpoint along the way — alongside narrower tools like SWE-grep and SWE-check. The prefix is the company's shorthand for software engineering, and it predates every model in this paragraph by roughly a year.
The name: a model, not a benchmark
This is the second reason searchers attach SWE-2 to the wrong owner, and it deserves its own paragraph rather than a footnote.
SWE-bench is a benchmark. It is an independent evaluation maintained by the SWE-bench team, hosted at swebench.com, and it ships in several editions: the original 2,294-instance set drawn from real GitHub issues, plus the Verified, Lite, Multilingual and Multimodal subsets. It is used to score coding agents generally, Devin included — Cognition published a Devin-on-SWE-bench technical report back in March 2024 that is still on its blog.
SWE-2 is a model. It is not an edition of SWE-bench, and it is not short for one. There is no SWE-bench 2: the published family runs SWE-bench, Verified, Lite, Multilingual and Multimodal, and the version numbering that exists belongs to the benchmark's subsets, not to a numbered successor. Cognition's benchmark of record for these models is called FrontierCode, which is a different instrument entirely and, as the next section explains, one Cognition owns.
If you came here expecting a second generation of the SWE-bench evaluation, that is the whole of the misunderstanding.
Release date and how SWE-2 is distributed
Cognition's announcement post carries a September 10, 2026 byline, and the page's own structured data gives the publication timestamp as 2026-09-10. Both agree. SWE-2 was available in Devin Desktop and Devin CLI on that date, with Devin Web and Fusion following.
That is the distribution surface, and it is the entire distribution surface. There are no open weights — nothing on Hugging Face from Cognition for this model — and no vendor-hosted endpoint that takes a SWE-2 identifier. If your harness is your own, SWE-2 is not a component you can install into it today. If your harness is Devin, SWE-2 is already in front of you.
For anyone who needs the base model's envelope rather than SWE-2's, Kimi K3 is a separate and better-documented thing, and it is routed on OrcaRouter as kimi/kimi-k3 — one API for 200+ models at the provider's list price with no markup. That matters here for a specific reason: the underlying capability Cognition started from is reachable independently, even though the post-trained model is not.
The envelope: what Cognition documents, and what it does not
A reference page should be honest about the shape of its own evidence, and this is where SWE-2's is thinnest.
• Reasoning effort — three documented levels: medium, high and max. Cognition writes that the levels behave differently by design: medium steps into action sooner and carries simple and intermediate work, while high and max plan more, explore more of the codebase, and spend more on verification.
• Distribution — Devin Desktop and Devin CLI at launch, Devin Web and Fusion rolling out. No API, no weights, no self-host path.
• Base model — Kimi K3, described by Cognition as a 2.8T-parameter model that had already undergone extensive RL for agentic coding. The Kimi K3 paper gives that base a 1M-token context window and native vision.
• Context window, max output, modalities, post-trained parameter count — not published for SWE-2. Cognition's announcement says nothing about any of them, and no other Cognition page fills the gap.
The distinction in that last row is not pedantry. Kimi K3's million-token window is a fact about Kimi K3. Cognition does not say that SWE-2 inherits it, and post-training on a different recipe is exactly the kind of change that can alter what a model accepts. If you need a context-window number for planning, the number you can verify belongs to the base model, and you should treat SWE-2's as unknown rather than assume the two agree.

The rate card, in the only currency Cognition quotes
There is no per-token price for SWE-2, because there is no SWE-2 endpoint. Cognition's cost claim is denominated differently: it says SWE-2 lands within one point of Claude Fable 5.1 on FrontierCode 1.1 Main — 50.0% against 50.9% — while being 64% cheaper. Read the unit carefully. The 64% is mean US dollars spent per rollout on FrontierCode, a task-level figure that bundles the model, the harness, the scaffolding and Cognition's serving stack into one invoice. It is not a token rate, and it cannot be compared against one.
The rate card that does exist is Devin's. On Devin's pricing page, read on September 28, 2026: Free at $0, Pro at $20 per month, Max at $200 per month, Teams at $80 per month plus $40 per full developer seat. The SWE-2 line sits in Pro and carries a date — "Free use of SWE-2 — Free in Devin Desktop and CLI through October 10, 2026." That is a promotional window with roughly two weeks left on it, and it is the one SWE-2 figure on that page that is genuinely time-sensitive: the model's cost inside Devin after October 10 is not stated there.
Cognition's efficiency numbers are worth quoting next to the cost claim, and they are vendor-reported like everything else on this page. On FrontierCode 1.1 Main, SWE-2 at medium effort scores above SWE-1.7 while taking 58% fewer turns and costing 81% less on average. Mean steps per run fall from 127 on SWE-1.7 to 53 at medium, 80 at high and 98 at max, and the median first real code edit moves from step 48 to step 18.
The benchmarks, with the harness attached
Software-engineering scores are unusually sensitive to the scaffold they were produced in, so a number without its harness is not a number. Cognition states its harness policy in the announcement's methodology appendix: results are reported from public sources where one exists, and otherwise run on Cognition's internal evaluation framework using the harness each model was primarily developed for — Claude Code for Anthropic models, Codex for OpenAI models, Grok Build for xAI models, and Devin CLI for open-weight models. For every model, Cognition reports the best score across reasoning-effort settings.
Two consequences follow, and both belong in the same breath as the numbers. First, several rows are not like-for-like: a Fable 5.1 figure produced in Claude Code and a Kimi K3 figure produced in Devin CLI are measurements of two different agent systems, not two models in a neutral harness. Second, "best score across effort settings" means each cell is a model's best case, not a matched setting. A comparison that holds effort constant would look different, and Cognition did not publish one.
All four rows below are vendor-reported and unaudited.
• FrontierCode 1.1 Main — SWE-2 50.0%, Kimi K3 44.2%, Grok 4.6 48.0%, Claude Fable 5.1 50.9%, GPT-5.6 Sol 47.5%, GPT-6 Astra 53.3%, SWE-1.7 42.0%. This is the headline row, and FrontierCode is Cognition's own benchmark — Cognition writes the tasks with 20-plus open-source maintainers, runs the leaderboard, and grades the submissions. None of that makes the figures wrong; all of it belongs in the sentence where you repeat them.
• DeepSWE 1.1 — SWE-2 73.0%, Kimi K3 68.5%, Grok 4.6 67.5%, Claude Fable 5.1 67.4%, GPT-5.6 Sol 72.7%, GPT-6 Astra 74.1%, SWE-1.7 37.7%. The row where SWE-2 most credibly beats a frontier model outright.
• Terminal-Bench 2.1 — SWE-2 92.8%, Kimi K3 88.3%, Grok 4.6 88.4%, Claude Fable 5.1 91.4%, GPT-5.6 Sol 88.8%, GPT-6 Astra 89.9%, SWE-1.7 81.5%. SWE-2's best-looking number, and the current edition of Terminal-Bench is not 2.1.
• Terminal-Bench 4 — SWE-2 27.3%, Kimi K3 21.5%, Grok 4.6 20.3%, Claude Fable 5.1 55.8%, GPT-5.6 Sol 37.3%, GPT-6 Astra 57.9%, SWE-1.7 7.6%. The row quoted least, and the one where the model sits roughly 28 points behind the frontier.
The comparison also needs one more piece of context, because it explains where the frontier row's unusual shape comes from. Cognition's own footnote to its charts notes that Fable 5.1's max-effort setting on FrontierCode 1.1 Main scores 50.3% at $12.83 per task — below its own medium setting at 50.9% and $3.28. More effort bought a lower score at roughly four times the cost. That is the frontier model's curve bending back on itself, and it is part of why 50.0% against 50.9% is closer than the two numbers alone suggest.
The trustworthiness numbers
Cognition's separate trustworthiness section is the part of the release that has nothing to do with leaderboards, and it reports that SWE-2 passed 98.0% of its propaganda-and-censorship prompts overall — 99.8% in English, 95.2% in Simplified Chinese and 99.1% in Traditional Chinese — and that its context-dependent vulnerability evaluation found no framing condition producing a statistically significant change for any model tested. Those are Cognition's judgments, produced with a GPT-5.6 Luna judge on a 145-question set, and they are vendor-reported in the same sense the benchmarks are.
The independent read: there isn't one yet
As of September 28, 2026, SWE-2 has no entry on Artificial Analysis. Searching the model index for the name returns nothing, and a direct model-page request 404s. The only scoreboard carrying SWE-2 is FrontierCode, which belongs to Cognition. That leaves the release with vendor-reported numbers and nothing else, which is not a criticism of the numbers — it is a statement about how much of them a reader can check.
Two specific things would settle it, and neither exists today. A third-party run of SWE-2 on Terminal-Bench 4 would decide whether this is a frontier-adjacent model that happens to be cheap or a fast one that happens to lead a retired edition. And any independent reproduction of the FrontierCode gap would tell you whether the one-point margin against Fable 5.1 is a property of the model or a property of the instrument.
Unlike Cognition's own models, Artificial Analysis does carry Kimi K3 — and Kimi K3's figures are third-party, which makes the base model the better-evidenced half of this stack. That asymmetry is the practical shape of SWE-2 today: the post-train is the interesting part and the least verifiable part; the base is the documented part and the one you can actually call.

Using SWE-2, and what to do while it is unaudited
If you already pay for Devin, the decision is easy and does not depend on any of the caveats above. SWE-2 is in your product, Cognition says it costs less per task than the model it replaces, and the efficiency figures — 58% fewer turns, 81% lower average cost, a median first edit at step 18 instead of step 48 — describe a model that wastes less of your time on ordinary work. Start it at medium effort on the simple end of your queue, which is where Cognition's own effort-level design puts the cost win.
If you are choosing a coding model from outside Devin, the honest position is that SWE-2 is not a candidate you can evaluate, because there is nothing to call. There is no identifier, no price per token, and no endpoint. What you can evaluate is the base model.

Kimi K3 is routed on OrcaRouter, at $3.00 per 1M input tokens and $15.00 per 1M output tokens with no markup over the provider rate, a 1M-token context window, native tool calling, image input and a top-level reasoning-effort control. That is the reachable half of this release: the capability Cognition post-trained from, available through one OpenAI-compatible endpoint rather than one vendor's agent product. And because it is unaudited territory either way, you can put a fallback chain behind it, so a model you are trialling does not become a production dependency the day it disappoints you.
What SWE-2 tells you, if you read it as evidence rather than as a product page, is narrower and more useful than the headline: a well-executed RL pass on top of Kimi K3 moved the level by roughly five points across four benchmarks without changing the shape, and took 58% fewer turns to do it. That is a real result inside Devin. It is not yet a result anyone outside Cognition has reproduced, and the one benchmark where SWE-2 appears to beat everything is the one the field has stopped scoring on.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
