· AI

Cognition's Cheap Coding Model Wins the Benchmark It Chose. It Loses the One That Matters Most.

Cost-aware training just got real, but the discount headline is measuring the wrong task.

Cognition released SWE-2 on September 10, a coding model built specifically to prove that reinforcement learning can be trained to care about price, not just correctness. On FrontierCode 1.1 Main, a coding benchmark, SWE-2 scores 50.0 percent against Anthropic's Fable 5.1 at 50.9 percent, a gap Cognition calls essentially a tie, while charging what it says is 64 percent less per task, according to Cognition's own announcement. Against OpenAI's GPT-6 Astra, the company claims comparable results at roughly a quarter of the cost. Read only that far and SWE-2 looks like the cleanest proof yet that cheap and good are no longer opposites.

Read the rest of Cognition's own table and the story gets more complicated. On Terminal-Bench 4, a harder benchmark built around long, multi-step agentic tasks, SWE-2 scores 27.3 percent. Fable 5.1 scores 55.8 percent. GPT-6 Astra scores 57.9 percent. The cheap model does not trail a little here. It finishes fewer than half as many of these tasks as its full-price competitors, on the exact category of work where a wrong answer costs the most to catch.

My position: the engineering behind SWE-2 is real progress, and the standard reflex when a benchmark like this shows up (ignore it, quote the flattering number) is exactly the trap a buyer needs to avoid.

Steelman the discount first, because it earns one. Cognition post-trained SWE-2 from Moonshot AI's open-weight Kimi K3 using a single reinforcement learning run that penalizes token cost directly inside the reward function, rather than bolting a cheaper model on afterward, according to Cognition's technical description. The method visibly worked. SWE-2 beats its Kimi K3 base by 5.8 points on FrontierCode and by an identical 5.8 points on Terminal-Bench 4, and it beats Cognition's prior model, SWE-1.7, using 58 percent fewer agent steps and 81 percent less cost per run on average. For the large share of coding work that is routine, refactors, test writing, small bug fixes, a model within one point of frontier performance at a fraction of the price is close to free money. Cognition also deserves credit here for publishing the Terminal-Bench 4 number at all instead of quietly leaving its worst result off the table.

Here is where the discount headline breaks down. FrontierCode 1.1 Main and Terminal-Bench 2.1, the two benchmarks where SWE-2 looks strongest, both test relatively short, well-scoped coding tasks. Terminal-Bench 4 is built for the opposite: long-horizon work where a model has to plan, act, and recover from its own mistakes without a person checking every step. That is precisely the kind of task a business is most tempted to hand off entirely, because it is the kind eating the most engineer time. A model that finishes barely a quarter of those tasks is not failing a fifth more often than one that finishes over half. It is failing almost twice as often, and a failure discovered at the end of a long, unsupervised run costs more to unwind than one caught five minutes in. Cognition has not published per-token pricing for SWE-2, so I cannot turn that failure rate into a dollar figure the way I can the ticket price, and I am not going to invent one. What is not speculation is the direction of the risk: a 64 percent discount measured on the tasks a vendor chose tells you nothing about the tasks you actually need done. An independent read of the release makes the same point from outside the company, noting that every number in Cognition's table, the flattering ones included, is Cognition's own run, with no independent replication published yet.

None of this makes SWE-2 a bad model or Cognition's release dishonest. Publishing your worst number is more candor than most vendors offer, and the underlying method, training a model to trade capability for cost on purpose, is exactly the kind of engineering this blog has been arguing will do more for AI affordability than any policy fix. It does mean the free lunch a 64 percent discount implies is only free for the slice of work that resembles the benchmark it was measured on. Assuming that slice is your whole workload is how a cost-saving swap turns into a rework bill nobody budgeted for.

The practical move is not to distrust SWE-2 or the discount. It is to stop treating any vendor's headline benchmark as a stand-in for your own task list. Pull ten tickets from your hardest backlog, the ones closer to a long agentic rebuild than a unit test, and run the cheaper model against them before routing a dollar of production work there. That is the same audit that belongs in any model or vendor decision, and it is worth having someone outside the vendor's own benchmark table run it. Reach out if you want a second opinion on that audit before you commit to a cheaper model across the board.

Sources

References used in this article. Links also appear alongside the relevant claims.

Let's talk

Tell us what's on your mind.

You don't need a polished brief to reach out. A two-line email about what's bugging you is plenty; we'll tell you straight if we're the right fit, and what we'd tackle first.

We'll scope the work around your workflow, goals, and timeline before quoting anything, so you know what's included before committing.

LocationBoca Raton, Florida
CoverageSouth Florida + remote nationwide
Status Now accepting clients