August 22, 2026 · AI
AT&T Cut AI Costs 56 Percent With Routing. Most Companies Are Doing the Opposite.
One telco proved that routing routine work to cheaper models barely costs you anything in quality, while most of the market quietly walked away from open models for reasons that have nothing to do with quality at all.
AT&T told The Information this week that routing employee coding queries to cheaper models cut the cost of those tasks by as much as 56 percent, with only a 2 percent drop in output quality. Mark Austin, the AT&T vice president who oversees employee AI use, says the company runs roughly 45 billion tokens a day through its internal Ask AT&T platform, using LiteLLM routers that judge how hard a task actually is before deciding whether it needs a frontier model or can go to something cheaper. Austin wants open-source and open-weight models, Nvidia's Nemotron and Meta's Llama among them, to grow from 40 percent of that volume to 60 or 70 percent over the next few years, while keeping the dollars going to Anthropic and OpenAI flat even as total usage climbs.
That is close to the exact case I have made on this blog for months: most of what a business asks AI to do is routine enough that a smaller or open model handles it fine, and paying frontier prices for it is waste. AT&T just ran that experiment at a scale most companies will never touch, and the result held. A 2 percent quality cost for a 56 percent price cut is not a close call on any task that isn't already at the edge of what the model can do.
Here is the fact that complicates a clean "open models win" headline. BigGo's reporting on the same story cites Menlo Ventures data showing that enterprise use of open-source AI actually fell industry-wide, from 19 percent of workloads in 2025 down to 11 percent in 2024 by their count, and that the drop tracked licensing restrictions and governance headaches, not model quality. If open models are closing the capability gap, as Austin himself says they are, narrowing from six to ten months behind frontier models to something tighter, then a shrinking adoption number across the wider market means most companies tried this road and turned back. AT&T's 60 to 70 percent target makes it an outlier at its scale, not a bellwether. That is a fair steelman of the skeptical read: routing sounds simple in a blog post and is a real operational lift in practice, and a lot of IT teams that priced out the governance work concluded it was not worth the savings.
I think that reading gets the diagnosis right and the prescription wrong. The governance overhead is real, but it is a discipline gap, not evidence that open models do not work. Notice what AT&T actually did: it did not grab the cheapest available token indiscriminately. Austin's team explicitly kept DeepSeek and Moonshot's open models off the list on security and compliance grounds, even though Chinese labs are setting the pace on open-weight price and performance right now. That is a company applying its own risk judgment, not a government agency banning a category of model, and it is the difference between prudent vetting and protectionism. The lesson from AT&T is not "use every cheap model you can find." It is "build the two things that make cheap models safe to use: a router that scores task complexity, and a screening step that keeps untrusted models out of the mix." Most of the enterprises in Menlo's declining 11 percent skipped straight to "this is too much work to manage" without building either piece, and gave up savings that were sitting on the table.
You do not need 45 billion tokens a day to benefit from the same logic. A small business running AI for support tickets, internal drafts, and routine coding does not need a telco's infrastructure team to route the boring 80 percent of that work to a cheaper model and reserve the expensive one for what actually requires it. It needs someone who has done the vetting once and built the routing rule, which is a week of engineering, not a standing governance committee. The companies retreating from open models are usually the ones treating this as a compliance project instead of an architecture decision.
If you are paying frontier prices for work that does not need frontier reasoning, tell me what you are running and I will help you figure out where a router and an open model save you real money without you having to build AT&T's governance program to get there.
Sources
Every factual claim above is drawn from these independently published sources, linked inline where first referenced.
Tell us what's on your mind.
You don't need a polished brief to reach out. A two-line email about what's bugging you is plenty; we'll tell you straight if we're the right fit, and what we'd tackle first.
We'll scope the work around your workflow, goals, and timeline before quoting anything, so you know what's included before committing.