· AI

Idle GPUs Are the Real AI Shortage, and the Market Just Started Pricing Them

The AI compute shortage was never really about supply. It was about running roughly twenty times more GPU than anyone actually used.

On October 1, the cloud provider Nebius announced it had acquired Inferize, a startup founded in January 2026 that builds software to cut the time it takes an AI model to spin up and scale when demand hits. Nebius calls the problem it solves the "idle GPU tax": the GPU sits there burning money during cold starts, demand spikes, and mid-run weight updates, so platforms keep spare capacity on hand just to hit their service targets. Inferize had a working prototype within three months of founding. Nine months later, it is part of Nebius's Token Factory inference platform. No price was disclosed.

That is a small deal. The problem it points at is not.

A Cast AI analysis of 23,000 Kubernetes clusters, published in April, found average GPU utilization across the sample at 5 percent. CPU utilization averaged 8 percent, memory 20 percent. Put plainly: companies are provisioning about twenty times the GPU capacity they actually use. The same report notes that AWS raised H200 GPU reservation prices 15 percent in January 2026, reportedly the first sustained price increase since EC2 launched in 2006. Idle capacity used to be a rounding error. On GPU hardware, at today's prices, it is a line item that shows up on the income statement.

Here is the steelman for why that 5 percent number is not simply incompetence. Overprovisioning GPUs is the rational choice when the alternative is a latency failure during a traffic spike, and AI traffic is spiky by nature: usage varies by time of day, by region, by whatever went viral that week. A team that right-sizes to average load and gets caught by a surge loses customers and gets paged at 2 a.m. A team that overprovisions just burns cash quietly. Most engineers, reasonably, pick the failure mode that does not page them. That is not laziness. It is a sane response to the incentives they were given.

What changed is the incentive, not the engineering judgment. Inference pricing has been falling fast enough, and margins at the neocloud layer have been thin enough, that "burn cash quietly" stopped being a free option. When Nvidia's Ian Buck told an audience this month that the industry now measures itself in tokens per watt, not GPU count, he was describing the same pressure from the hardware side: every idle GPU-hour is a watt of power and a unit of scarce supply doing nothing, while someone downstream is selling tokens at a thinner and thinner margin on the GPUs that are working. At a 5 percent utilization rate, the real cost of serving one hour of useful inference includes paying for roughly twenty hours of reserved capacity. No regulator wrote that math into existence. A price war did.

It is worth naming the worry a free-market case should take seriously here: is a large cloud buying up the small company that solved the efficiency problem just consolidation with better branding? I do not think so, in this case. Cast AI itself sells a competing optimization product, and so do several others. The technique is not proprietary physics locked behind a patent wall; it is engineering discipline that a nine-month-old startup proved out faster than a well-funded incumbent did on its own. An acquisition at that scale is the market rewarding speed, and any competitor with capital can answer by building or buying the same thing. That is a very different kind of barrier than a compliance regime that only the largest players can afford to clear.

The practical takeaway, if you are running your own GPU spend rather than reading about Nebius's, is to actually measure your utilization before you buy more hardware. Most companies cannot tell you their real number, which is exactly the problem Cast AI found at scale. If you are paying for GPU capacity, reserved or on-demand, to run an open-weight model or a fine-tuned endpoint, the odds are good that you are closer to that 5 percent average than you think. That is the kind of question a cost review exists to answer honestly, before the next invoice. If you want a second set of eyes on where your AI spend is actually going, get in touch.

Sources

References used in this article. Links also appear alongside the relevant claims.

Let's talk

Let's make it happen.

You don't need a polished brief. A couple of lines about where you want your company to go is plenty, and we'll come back with what we'd tackle first.

We scope the work around your goals and timeline before quoting anything, so you know exactly what you're getting.

LocationBoca Raton, Florida
CoverageSouth Florida + remote nationwide
Status Now accepting clients