September 6, 2026 · AI
Anthropic Cut Cache Prices 75 Percent. Its Own Customers Already Solved the Problem Differently
A 75 percent price cut on cache reads is good engineering. It is not the fix for the bill problem most businesses actually have.
Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 on September 1, and the headline number is a 75 percent cut to cache-read pricing, from $1.00 to $0.25 per million tokens, according to the company's own announcement. Anthropic says that translates to roughly 25 percent lower effective cost for typical workloads and up to about 45 percent for heavy agentic and coding use, where an agent rereads the same system prompt and file context on nearly every turn. Base input and output pricing did not move: still $10 per million input tokens and $50 per million output tokens, identical to Fable 5.
My position: the cache cut is real and worth using if you run agentic workloads on Fable-tier models. But it answers a narrower question than the one most businesses are actually asking, which is why their AI bill is unpredictable at all. And Anthropic's own customers, per independent transaction data, have already answered that question a different way: by not buying the flagship model in the first place.
Steelman the cut on its own terms first, because it deserves one. A caching discount is exactly the kind of cost reduction this blog likes to see: an infrastructure improvement passed to the customer, not a headline discount funded by burning more venture money. VentureBeat's coverage notes Terminal-Bench coding scores rose from 42.0 percent on Fable 5 to 55.8 percent on Fable 5.1, so this is not a pure price move dressed up as a release. And the mechanics genuinely favor agent-style usage: a coding agent that calls tools in a loop resends most of its context on every step, so if cache reads make up the bulk of a session's billed tokens, a 75 percent cut on that line item is not a rounding error. Anyone running long Claude Code sessions or similar agent loops on Fable should benefit on day one, with no code changes required.
Here is where the picture gets more complicated than the press release. The Financial Times reported, using Ramp's transaction data across roughly 70,000 companies, that more than two months after Fable 5 launched it accounted for only about 11 percent of Anthropic spending, with the cheaper Opus 5 and Opus 4.8 taking share instead. Opus 5 runs $5 per million input tokens and $25 per million output, exactly half of Fable's rate card, for performance that is close enough on most real tasks that businesses are voting with their spend. The same Financial Times reporting, cited by VentureBeat, points to a separate strand of the story: enterprise customers describing unpredictable AI bills as a real source of friction, independent of which model they picked. That is consistent with what surfaced publicly in May, when Axios reported an AI consultant's client burned through half a billion dollars in a single month on Claude licenses because nobody had set a per-seat usage limit.
Put those two facts together and the math on Anthropic's own price cut looks smaller than the percentage suggests. A 45 percent reduction applies only to the cache-read portion of a bill, on a model that, per the FT's data, is already a small and shrinking slice of what most organizations spend. If your actual problem is that 10 to 15 percent of your team is generating 60 percent of an unbounded bill, as Rippling found in its own AI spend audit this summer, a cache discount on the model that small group happens to be using does not touch the rest of the org. It is a real savings for a real subset of workloads, not a fix for cost governance.
The practical read for a business owner: check whether you are actually running Fable-tier agentic workloads with heavy repeated context before assuming this cut moves your number. If you are, take it, it is free money. If your problem instead looks like the market's broader move toward Opus-tier and cheaper models for routine work, the fix is task-to-model routing and usage limits, not waiting on the next vendor price cut. That is the same conversation we have with clients under MojoAI advisory: which tasks genuinely need the frontier model, and which ones a model at half the price clears just as well. If nobody at your company has done that audit yet, that is worth a conversation.
Sources
Every factual claim above is drawn from these independently published sources, linked inline where first referenced.
Tell us what's on your mind.
You don't need a polished brief to reach out. A two-line email about what's bugging you is plenty; we'll tell you straight if we're the right fit, and what we'd tackle first.
We'll scope the work around your workflow, goals, and timeline before quoting anything, so you know what's included before committing.