THURSDAY 10 SEPTEMBER 2026 latent·wire 92 PIECES ON FILE
← ModelsModels

Cognition launches SWE-2 coding model, claiming frontier-level scores at 64% lower cost

Cognition released SWE-2, a coding model it calls its most advanced yet, claiming 50.0% on FrontierCode 1.1 Main and describing the result as within one point of Fable 5.1 at 64% lower cost. The company published the launch post on September 10, 2026 and said SWE-2 is available starting today in Devin Desktop and CLI, with a rollout on Devin Web and Fusion. The Hacker News submission for the launch drew 111 points and 62 comments.

Cognition ships Devin, its autonomous software engineer, and SWE-2 is the company's second model pitched at frontier labs. The model post-trains from Kimi K3, a 2.8 trillion-parameter base that had already undergone extensive reinforcement learning for agentic coding. Cognition says its own RL stage still finds substantial headroom on top of K3, adding 5 to 6 points on many benchmarks and shifting the base model's entire cost-performance frontier.

The scale of that RL run is the central technical claim. Cognition says it trained SWE-2 with reinforcement learning in the multi-trillion-parameter regime for the first time, building on the infrastructure and recipe developed for SWE-1.7. On cost, the post places SWE-2 within a few points of GPT-6 Astra at a quarter of the price, and at a fraction of what Sol and Fable 5 or 5.1 cost for comparable scores.

Memory and efficiency work backed the run. Cognition used NVFP4 and FP8 kernels with quantization-aware training, which it says reduces overall memory use and produces lower train-inference mismatch than SWE-1.7 at similar throughput, despite a base model with almost three times the parameters. The company also tripled the number of RL environments it trains against, added instruction-following overlays, and built a flywheel in which earlier SWE-2 checkpoints iteratively harden the verifiers used to grade the model.

One behavior change addresses user complaints about the previous model. SWE-1.7 gained a reputation for working carefully through thorough exploration of a codebase before making edits, and Cognition says that while this boosted performance it produced feedback that the model over-explored and overthought simple tasks. Cognition reports the largest efficiency gains in SWE-2 come from focused exploration, with higher intelligence letting the model judge which parts of a codebase matter for a given task.

Every performance figure in the launch rests on Cognition's own reporting. The post reports 50.0% on FrontierCode 1.1 Main and does not cite third-party or independently run evaluations of SWE-2. The Hacker News thread had accumulated 62 comments alongside the 111 points without an independent replication of those scores.

What remains unclear is how SWE-2 behaves outside Cognition's own harness and what it charges per token, neither of which the launch post discloses. Cognition said it is rolling SWE-2 out across Devin Desktop, CLI, Web and Fusion, so third-party results should surface as access widens beyond the company's early users.

Why it matters

Cognition is claiming a coding model that matches frontier performance at a fraction of the cost, a claim that would change what teams pay for agentic software work if independent testing confirms it.