· AI
Open-Weight Models Hit 56 Percent of AI Traffic in August. Anthropic Still Kept 64 Percent of the Money.
The market did not pick a winner between cheap and expensive AI. It split the work between them, on purpose.
Open-weight models ran the majority of AI Gateway traffic for the first time in August, according to Vercel's latest AI Gateway Production Index, published September 17. Fifty six percent of tokens routed through Vercel's gateway went to open-weight models, up from 13 percent in April and under 10 percent last December. Average price per token fell 23.2 percent that same month, the third straight monthly drop. Read the headline number alone and the story writes itself: open weights won, closed frontier labs are losing the price war they started.
That story is at best half right, and the same report proves it. Open-weight models took 56 percent of volume and only 14 percent of spending. Anthropic, meanwhile, has held 61 to 64 percent of gateway dollars every month since December, closing August at 64 percent, even as its own cheaper flagship, Opus 5, launched in late July and reached 22.5 percent of spend by August while the pricier Fable 5 fell from 13.2 percent of spend in July to 4.9 percent. The market did not choose cheap over expensive. It chose cheap for most of the work and expensive for the work that matters, in the same month, from the same buyers.
Steelman the "nothing actually changed" reading first, because Gennaro Cuofano's critique of the same dataset earns it. His July numbers, pulled from the same gateway a month earlier, show open weights at 36 percent of volume and 8.6 percent of spend, against Anthropic's 30 percent of volume at 65 percent of spend, priced 4.4 times the gateway average. If that spend ratio barely budges while volume jumps from 36 to 56 percent in a month, the fair read is that buyers are running cheap experiments and routine batch jobs through open models while never seriously trusting them with anything consequential. Cuofano's sharper point is about lock-in: developer tooling, IDEs, and SDKs are still mostly built around one vendor's conventions, so moving a real workload off a frontier model costs engineering time even when that model runs, per Vercel's own pricing comparison, roughly nine times what Claude Opus 5's own smaller open rival Gemini 3 Flash costs per token. On his account, inertia, not quality, is propping up two thirds of the spend.
I think that reads the data too narrowly. What happened in August looks less like inertia and more like a market pricing risk correctly, faster than any regulator or standards body could manage. A support ticket triage, a first-pass code lint, a document summary: none of that needs a model priced nine times higher per token, and buyers moved more than half of all gateway volume to cheaper alternatives in four months flat, with no mandate telling them to. The direction shows up outside Vercel's data too. OpenRouter usage figures Bloomberg cited in June put combined US frontier-lab share at 70 percent of that platform's tokens in June 2025 and 30 percent a year later, with DeepSeek alone taking 16.3 percent on its own. Different platform, different methodology, same conclusion: the routine bulk of AI work is being priced like the commodity it is, without a single new regulation making that happen. The premium that remains is not proof the frontier labs escaped discipline. It is the price buyers are still willing to pay for the sliver of work where a wrong answer is expensive, and right now Anthropic and OpenAI are who they trust to sell it.
One caution before you copy this pattern: Cuofano also flags that Kimi K3 reportedly burns about twelve times the tokens per request that its predecessor K2.5 did. A model that is cheap per token but verbose per answer is not automatically a cheap model to run. Price the completed task, not the rate card, before you route anything to it.
The practical move if you are the one signing the AI bill: stop treating "which model" as one decision for the whole company. Build the barbell deliberately instead of finding out later that you already have one, unmanaged. Route the routine work, the 90 percent that tolerates an occasional miss, to an open-weight model behind a router or your own hosted endpoint. Measure cost per completed task, not the advertised per-token price. Keep a frontier model on a short leash for the work where being wrong is expensive, and know which category each workload falls into before you spend on either side of it. That split is most of what we walk clients through at Mojo before they commit real budget to any model. If your AI bill still looks like one number instead of two, let's talk.
Sources
References used in this article. Links also appear alongside the relevant claims.
Tell us what's on your mind.
You don't need a polished brief to reach out. A two-line email about what's bugging you is plenty; we'll tell you straight if we're the right fit, and what we'd tackle first.
We'll scope the work around your workflow, goals, and timeline before quoting anything, so you know what's included before committing.