· AI
Open Models Now Carry Most AI Traffic. Most Company Bills Do Not Reflect That Yet.
The market has already voted with its traffic. Most companies are still paying as if the vote never happened.
Open-weight models carried the majority of AI traffic for the first time in August 2026. Vercel's AI Gateway Production Index, built from real production traffic flowing through its own gateway rather than a survey, found open-weight models reached 56 percent of gateway tokens that month, up from 7 percent in December 2025. Those same models accounted for only 14 percent of the dollars spent. Anthropic alone took 64 percent of spend. That gap between where the tokens go and where the money goes is the whole story, and it tells you something concrete about your own AI bill whether you have looked or not.
This is not one platform's quirk. An independent read of OpenRouter's usage data, compiled by venture investor Tomasz Tunguz, put open-weight models at 69.1 percent of a named open-versus-closed snapshot by early June 2026, up from a minority position in 2025. Two different gateways, tracking different customer bases, are showing the same direction of travel.
The honest case for staying on a frontier model
I want to steelman the case for staying put, because the capability argument is real and I would be lying to pretend otherwise. Epoch AI's research tracks the gap between the best open-weight model and the best closed frontier model on its own capability index, and as of a May 2026 update, that gap was not closing. It had widened slightly, from roughly three months behind in October 2025 to roughly four months, or 8 points on Epoch's index. If your work is the kind where that gap actually bites, difficult multi-step reasoning, ambiguous judgment calls, the 10 percent of tasks a wrong answer actually costs you money on, paying for the frontier model is not brand loyalty. It is a correctly priced insurance premium, and I would tell a client who needs that ceiling to keep paying for it.
The mistake is assuming your whole workload needs that ceiling. Most of it does not, and the market's own traffic proves the point better than I can argue it.
What the price gap actually buys you
Anthropic's Claude Opus 5, which launched July 24, 2026 explicitly as a cheaper sibling to its own flagship, prices at $5 per million input tokens and $25 per million output. DeepSeek's V4 Pro runs $1.74 per million input tokens and $3.48 per million output at standard rates, with a promotional rate as low as $0.435 and $0.87 active at times since launch. Even at the standard, undiscounted rate, that is roughly a third the input cost and under a seventh the output cost of Opus 5, and DeepSeek V4 Pro is not a toy model. It ties the closed frontier on several coding benchmarks.
Run the math on a mid-size workload: 300 million tokens a month, split 60/40 input to output, which is a reasonable shape for ticket responses, document summaries, and first-draft extraction. All on Opus 5, that workload costs about $3,900 a month. Routed to DeepSeek V4 Pro at its standard rate, the same volume costs about $731, a drop of more than 80 percent before you even account for any promotional pricing. Keep the hard 10 percent on the frontier model and route the rest, and you are not choosing between quality and cost. You are paying frontier prices only where the frontier ceiling is actually load-bearing.
The market already did this. Most companies have not.
Vercel's data backs this up at the aggregate level: average price per token across its gateway fell 23.2 percent in August alone, the third straight monthly drop, purely from more traffic moving to cheaper models that are good enough for what they are asked to do. That decline did not require a mandate or a vendor begging customers to switch. It happened because enough individual routing decisions, made by engineers who could see their own bill, added up to a market-wide shift. Credit the competition, not a policy, for that curve bending down.
The gap most businesses are missing is not technical. It is that the market-wide 56/14 split is an average of thousands of companies that have already segmented their workloads, pulled up against the ones still sending every request through a single frontier contract because nobody has sat down and split the traffic. If your AI bill has not meaningfully moved in the direction that chart is moving, you are very likely still in the second group. We run exactly this kind of workload audit at Mojo, and it usually takes less time than most businesses expect to find out which 10 percent of their traffic actually needs the expensive model. If you want a second opinion on your own stack, that conversation is free.
Sources
References used in this article. Links also appear alongside the relevant claims.
Let's make it happen.
You don't need a polished brief. A couple of lines about where you want your company to go is plenty, and we'll come back with what we'd tackle first.
We scope the work around your goals and timeline before quoting anything, so you know exactly what you're getting.