· AI

Token Prices Are Falling. The Chips Under Them Just Got More Expensive.

Cheap tokens and expensive chips are both real right now, and only one of those trends is a subsidy that can end.

AMD told partners on September 17 to expect roughly a 10 percent price increase on AI accelerators, Radeon cards, and motherboard chipsets starting in the fourth quarter, citing rising wafer costs at TSMC as the reason (PCWorld). Nine days earlier, on September 10, Oracle co-CEO Clay Magouyrk told investors on an earnings call that not one customer let a GPU reservation lapse last quarter. Every contract due to expire got extended by its existing tenant or picked up by a new one, at a price 20 percent above the old contract, on hardware mostly four years old already (The Motley Fool). Read only that and it sounds like AI compute got expensive again. Read the other half of the ledger and it says the opposite: average price per token fell 23.2 percent in August alone, the third straight monthly drop, according to Vercel's AI Gateway Production Index published September 17 (Vercel). Both numbers are correct. They describe two different markets.

My read: the chip layer and the token layer are pricing two different kinds of scarcity, and a business owner who only watches the API bill will misjudge where next year's real cost pressure sits.

The fair objection is that the hardware data is not as solid as it looks. AMD's number does not come from AMD. It traces back to a supply-chain report out of a distributor intelligence outlet called ChannelGate, amplified through tech press, not an AMD filing or statement. Oracle's 20 percent figure is a genuinely strong data point, but it comes from a company that recently sold roughly 20 billion dollars in stock to fund the same AI buildout it is now citing as proof of scarcity, and the stock rallied on the news, which is exactly the story Oracle wants told. A skeptic can reasonably argue: wait for the new fab capacity due in 2027, and remember that model efficiency, not cheaper chips, is what has been cutting token prices all year, so none of this hardware noise should change how you budget an API contract.

That skepticism deserves a real answer, and the answer is in the details of the same data everyone is citing. AIMultiple's GPU rental index, which tracks listed prices across 75 cloud providers, shows the newest chip tier, the B200, B300, MI300X and RTX 5090 class, rising from a $2.12 median hourly rate in October 2024 to $4.72 in September 2026, a 122 percent jump (AIMultiple). The prior generation, the H100 and A100 class that runs the overwhelming majority of production AI workloads today, moved from $1.35 to $1.45 an hour over the same nearly two years, a 7.4 percent rise that barely covers inflation. Older hardware still, the V100 and P100 class, actually got 46 percent cheaper. The scarcity is not "AI compute" as a category. It is specifically the newest silicon, the tier that frontier labs and hyperscalers like Oracle buy first and in bulk, and that is exactly the tier AMD's TSMC wafer costs hit hardest. The token price war, meanwhile, is being fought on levers that are not that bottleneck: open-weight models, quantization, routing, and caching all cut the number of GPU-hours a given task consumes, which is separate from what a GPU-hour costs to rent. Some of that 23 percent monthly token discount is real efficiency. I think part of it is also market-share subsidy, labs pricing below their own marginal cost to win volume while the newest hardware under them gets pricier, and a subsidy is the one item on this list that is not durable. I cannot prove the exact split between efficiency and subsidy from public data, and neither can anyone else citing these same reports, so treat that split as my judgment, not a fact.

The move for a business signing AI contracts into 2027 is to price each layer on its own terms instead of reading one headline number. If you are buying API access to a frontier model, the discount you are getting is real today, but do not write it into a multi-year plan as if it is guaranteed. A price war has an end date. A wafer shortage does not, on any timeline shorter than a couple of years. If you are pricing your own inference, whether local, open-weight, or rented, price it against the H100 and A100 tier, not the newest chip on the market. At $1.45 to $3.25 an hour that tier still runs the large majority of real production workloads at a cost that has barely moved in two years, while chasing the newest chip at roughly triple that rate buys speed most businesses do not need for the routine share of what actually runs through a model.

This is the exact math we walk clients through before they commit budget to any vendor or hardware tier: which layer is actually scarce for your workload, and which discount you are being handed is durable versus a subsidy with a shelf life. If your 2027 AI budget is still one number instead of a breakdown by layer, let's talk.

Sources

References used in this article. Links also appear alongside the relevant claims.

Let's talk

Tell us what's on your mind.

You don't need a polished brief to reach out. A two-line email about what's bugging you is plenty; we'll tell you straight if we're the right fit, and what we'd tackle first.

We'll scope the work around your workflow, goals, and timeline before quoting anything, so you know what's included before committing.

LocationBoca Raton, Florida
CoverageSouth Florida + remote nationwide
Status Now accepting clients