· AI

Claude Sonnet 5.5 Kept Sonnet 5's Price. Independent Testing Found a Bigger Bill Anyway.

The sticker price never moved. What you actually pay depends on a dial most buyers never touch.

Anthropic released Claude Sonnet 5.5 on September 28 at the exact same price as its predecessor: $2 per million input tokens, $10 per million output tokens, unchanged since Sonnet 5 shipped in June (Anthropic). The company's pitch is that the flat price is a discount in disguise. Sonnet 5.5 needs far fewer tokens to finish the same work, Anthropic says, and "costs up to 30% less per task than its predecessor" in its own testing.

I do not think that claim is false. I also do not think it is the number a buyer should budget against, because independent testing found roughly the opposite result at the settings Anthropic's own API defaults to.

Steelman first, because Anthropic's evidence here is specific, not marketing fluff. Slack reported about 14 percent fewer output tokens on its internal evaluations, reaching equal or better results in fewer steps. Box measured the model running 2.4 times faster while using 12 percent fewer total tokens on its own workloads. Base44 said its average app build dropped from 7.7 iterations on Opus 5 to 3.6 on Sonnet 5.5. Zendesk processed support tickets 20 percent faster. That is four unrelated companies reporting the same direction of change on their own production traffic, not a cherry-picked demo. On ordinary work, the efficiency gain looks real, and Anthropic deserves credit for shipping it without raising the price, which is what real price competition is supposed to look like.

Here is where it gets complicated. Artificial Analysis, an independent benchmarking firm with no stake in Anthropic's sales numbers, tested Sonnet 5.5 across its available reasoning-effort settings, including "high," the tier Anthropic's API falls back to by default absent a developer override. Push the model one parameter further, to its top "max" setting, and the firm measured about 193,000 output tokens per task on its Intelligence Index, the most it has recorded for any model it has tested, roughly seven times what GPT-6 Astra uses at its own maximum setting and about 60 percent more than Anthropic's own Opus 5.5 (Artificial Analysis). Cost per task at that setting came out near $7.60, about 50 percent higher than Sonnet 5, not lower. The extra reasoning bought almost nothing in score. Artificial Analysis places the model off its own intelligence-per-dollar frontier at that configuration, a worse trade than Opus 5.5 or several cheaper rivals, despite the lower list price sitting right there on the same page.

Both findings are true at once, and neither company is lying. Anthropic's 30 percent number describes typical production work, the kind Slack and Box actually run through it every day. Artificial Analysis's benchmark tasks push the model toward its ceiling, where Sonnet 5.5 apparently chooses to keep reasoning for marginal benchmark points, the same instinct you would not want showing up on a routine support ticket. The token rate held still. The bill moved a great deal in both directions, and which direction it moved for you depends on a reasoning-effort setting that most developers never touch after picking a model, and that Anthropic ships turned up higher than "medium" by default.

I made a version of this point when Sonnet 5's price froze back in August: the number on a pricing page was never what decided your actual bill. This is the sharper version. Now even a vendor's own "percent cheaper" claim is a single data point measured at one configuration among several, and the configuration it shipped as the default is not obviously the cheap one. A business that lifts Sonnet 5.5 into production on the API defaults could plausibly pay more per completed task than it did on Sonnet 5, price freeze notwithstanding. A business that tunes the effort setting down to match how hard its actual work needs to think could beat Anthropic's 30 percent figure by a wide margin.

The fix is boring, and it works: before moving a production workload to a new model release, test it at the effort setting you will actually run at scale, on your own tasks, not the vendor's demo tasks or whatever the API defaults to. That is half a day of engineering time set against a bill that compounds every month it goes unchecked. It is also, not coincidentally, the kind of vendor call I help clients make at Mojo, because the gap between a model that is "cheaper" and one that is actually cheaper for you almost never shows up on the pricing page. If you are about to make that switch, test it before you trust the announcement: mojoaiservices.com/#contact.

Sources

References used in this article. Links also appear alongside the relevant claims.

Let's talk

Tell us what's on your mind.

You don't need a polished brief to reach out. A two-line email about what's bugging you is plenty; we'll tell you straight if we're the right fit, and what we'd tackle first.

We'll scope the work around your workflow, goals, and timeline before quoting anything, so you know what's included before committing.

LocationBoca Raton, Florida
CoverageSouth Florida + remote nationwide
Status Now accepting clients