Back to News
User submissioncognition.com

Cognition Introduces SWE-2 Coding Model With Lower-Cost Frontier Performance

Summary

Cognition has introduced SWE-2, its most advanced coding model, reporting 50.0% on FrontierCode 1.1 Main, within one percentage point of Fable 5.1 while costing 64% less. The company says SWE-2 is the first model for which it scaled reinforcement learning to the multi-trillion-parameter regime, building on the SWE-1.72 training infrastructure and using one run to train multiple reasoning-effort levels. It post-trains Kimi K3, a 2.8-trillion-parameter model that had already received extensive reinforcement learning for agentic coding, and Cognition reports gains of 5–6 points on many benchmarks. On FrontierCode 1.1 Main and DeepSWE 1.1, SWE-2 reportedly beats SWE-1.7 and Grok 4.6 on both score and cost, matches GPT-5.6 Sol and Fable 5/5.1 at a fraction of their price, and comes within a few points of GPT-6 Astra at a quarter of the cost. Across the listed benchmarks, SWE-2 scores 73.0% on DeepSWE 1.1, 92.8% on Terminal-Bench 2.1, and 27.3% on Terminal-Bench 4. Cognition’s training method applies effort-level-specific linear cost penalties calibrated to the local slope of the base model’s cost-performance frontier, while a length-weighted reward baseline is intended to reduce gradient variance and keep inference and training policies closer. The rollout system combines scheduling changes, online draft-model training for speculative decoding, NVFP4 and FP8 kernels, and quantization-aware training. The company also tripled its reinforcement-learning environments, added instruction-following overlays, and iteratively hardened verifiers against false positives, false negatives, and reward hacking. In an internal 100-task FrontierCode evaluation, SWE-2 medium used a mean of 53 steps versus 127 for SWE-1.7, with 58% fewer turns and 81% lower average cost; it also made its first real edit after a median of 18 steps versus 48. SWE-2 is available in Devin Desktop and CLI, with rollout planned for Devin Web and Fusion. Cognition reports a 98.0% overall pass rate on its propaganda and censorship evaluation, while no tested customer or language framing produced a statistically significant change in vulnerability on its coding safety evaluation.