August 18, 2026 · AI
Gartner's "Inference Paradox" Is Real. The Fix Isn't a New Budget Line.
Cheaper tokens and a bigger AI bill are not a contradiction. They are what happens when you let complexity, not judgment, decide how a task gets solved.
Gartner put out a forecast on August 17 that every business owner running AI agents should read twice: inference costs per agentic workflow will rise more than fivefold through 2028, even as the underlying cost of a token keeps falling. Gartner analyst Will Sommer explains the mechanism this way: a simple chatbot reads a query and returns one probabilistic answer, while an agent has to reason through a problem, negotiate between steps, and check its own work along the way, and each of those extra steps is its own round of token spend. Routing a task to an agentic reasoning model costs at least five times what a basic chatbot exchange costs, by Gartner's estimate, and that multiplier climbs as the workflow gets more elaborate.
Computerworld's writeup adds the sharper numbers behind the headline: Gartner expects token prices to fall 95 percent by 2030, but pegs advanced reasoning agents at up to 150 times the cost per task of a simple chatbot, driven by agents that burn 5 to 30 times more tokens to do equivalent work and by swarms of agents that trigger each other into cascades of extra calls. Gartner's own term for this is the "inference paradox": cheaper unit economics is exactly what encourages people to run bigger, more complicated workflows, so total spend can climb faster than the per-unit savings can offset it. Sommer and co-analyst Sabine Zimmerhansl warn buyers against what they call a token-deflation illusion, the assumption that a provider's price cut shows up in your bill automatically. By their assessment, for a business running agents at real scale, it will not.
That is a legitimate warning, not hype, and it deserves a steelman before I push back on where it leads. Finance teams have seen this movie. Cloud spend in the 2010s ballooned the same way: compute got cheaper per unit, so engineers provisioned more of it, and bills went up anyway because nobody was watching the aggregate. Usage-based AI pricing on top of autonomous agents that can spawn their own sub-tasks is a real setup for the same kind of surprise, and Gartner's recommendations, tiering requests to the cheapest adequate model, moving vendors off flat fees onto usage-based terms, tracking cost against a real outcome instead of a vibe, are sound advice. A business that ignores this and lets agent swarms run unmetered against a credit card is going to get an ugly invoice. That risk is not manufactured.
Here is where I part ways with how this gets framed once it reaches an enterprise audience. The instinct that follows a forecast like this is usually to buy a governance layer: an AI cost-management platform, a chargeback system, a committee that reviews agent budgets before they ship. That is treating a discipline problem as a tooling problem, and it is expensive in exactly the way that keeps enterprise IT bloated and slow. Rising total token consumption is not itself bad news. It is what a technology looks like when it is actually getting used, the same pattern you would see with cloud, with cheap solar, with any input that got radically cheaper and radically more adopted as a result. The problem is not that agents consume more tokens. The problem is consuming more tokens without getting more value back per task, and that is solved with engineering judgment, not a new SaaS subscription.
The fix that actually works is the one that does not require Gartner's org chart: right-size the model to the task before you reach for the agent. Most of what businesses ask an "agent" to do, look up a record, draft a routine email, classify a support ticket, is well within reach of a small, cheap model or a deterministic script, and does not need a reasoning model negotiating with itself across a dozen tool calls. Reserve the expensive agentic reasoning for the sliver of work that genuinely needs multi-step judgment, and route everything else to the cheapest tool that reliably gets it right. That is model routing and right-sizing, not governance theater, and it is the difference between the fivefold cost curve Gartner is warning about and a bill that tracks the value you are actually generating.
If you are running agent workflows and watching the invoice climb faster than the output, that is worth a real look before you buy anything to manage it. Tell me what you have running and I will tell you honestly whether the fix is a smaller model or better judgment about when to reach for a bigger one.
Sources
Every factual claim above is drawn from these independently published sources, linked inline where first referenced.
Tell us what's on your mind.
You don't need a polished brief to reach out. A two-line email about what's bugging you is plenty; we'll tell you straight if we're the right fit, and what we'd tackle first.
We'll scope the work around your workflow, goals, and timeline before quoting anything, so you know what's included before committing.