Cognition shipped SWE-2 yesterday, and the framing is "pushing the Pareto frontier": within a point of Fable 5.1 on FrontierCode 1.1 Main while being 64% cheaper. That's a real number on an easy benchmark. But the scorecard has a second column that tells a different story, and it's the one nobody is looking at.
Take Terminal-Bench, the pair TB2.1 and TB4. TB2.1 is the cheap slice, the everyday stuff. TB4 is the hard, adversarial tail. The gap between them is what the model gives up when you dial up the difficulty.
SWE-2: 92.8% on TB2.1, then 27.3% on TB4. That's a 65.5 point collapse.
GPT-6 Astra: 89.9% to 57.9%. 32 points.
Fable 5.1: 91.4% to 55.8%. 35.6 points.
GPT-5.6 Sol: 88.8% to 37.3%. 51.5 points.
So SWE-2 has the best easy-slice score in the group and tied with Grok for the worst hard-tail score out of the real contenders. The "within one point of Fable" line is true on the slice where everyone is basically saturated. On the slice where capability actually separates, Fable beats it by 28 points.
This is the pattern, and it's a training artifact before it's a capability one. When you train with a linear cost penalty per effort level, tuned to the local slope of the Pareto frontier, you get a model that's excellent at the cheap end because that's what the optimizer is rewarded for. The tail gets given up because the reward slope there is thinner. Cognition literally wrote this in the post: the RL penalty is tuned to the base model's frontier shape. Cheap wins, tail sacrificed.
None of this means SWE-2 is bad. If your workload is the easy slice, it's probably the best price-performance you can point at today. But if you're using benchmark tables to pick a model for genuinely hard, adversarial bugs, the headline number is telling you about the flat part of the curve. The decay column is where the actual separation lives.
My rule now: always read the scorecard in pairs, an easy slice and a hard tail, before believing a Pareto claim. A model can "push the frontier" on the slice it optimized and look catastrophic on the slice you actually needed.
Top comments (0)