September 5, 2026 · AI

Gemini 3.8 Flash Is Cheaper Per Token and Pricier Per Task, and That Gap Is the Real Story

The price war is real and it is working, but if you are budgeting off the per-token number on the announcement page, you are budgeting off the wrong number.

Google shipped Gemini 3.8 Flash on September 2, priced at $0.75 per million input tokens and $3.75 per million output tokens, introductory through the end of the year. Both figures undercut GPT-5.6 Sol at $5 and $30 and Claude Fable 5.1 at $10 and $50, by a wide margin. On paper it looks like the cheapest serious model on the market got cheaper. Independent benchmarking says the opposite happened to what a buyer actually pays.

My position: the per-token price on a model card has stopped being a useful number for budgeting, and businesses that still shop on it are going to keep getting surprised. What matters is cost per completed task, and that number is moving in the other direction even while the sticker price holds still.

Here is the honest case for Google's approach, because it deserves one. Artificial Analysis, the benchmarking firm that runs models through a standardized set of reasoning and coding tasks, put Gemini 3.8 Flash at 59 on its Intelligence Index, up from 56 for the prior Flash model, ranking 17th out of 196 models tracked against a median score of 36. With high reasoning enabled, they found it sitting on the Pareto frontier of intelligence versus cost per task, meaning nothing else on the market currently delivers more capability for less money at that tier. That is a genuine gain, not a marketing chart. If a model thinks longer and calls more tools to get a harder problem right on the first pass, the extra tokens are buying something real: fewer retries, fewer escalations to a human, fewer wrong answers shipped to a customer. Test-time compute has a cost, and sometimes it is worth paying.

But the same benchmarking firm measured what that reasoning actually costs in practice, and the number should reset expectations. Cost per Intelligence Index task came in at $0.58 for Gemini 3.8 Flash on high reasoning, which Artificial Analysis called "the cheapest we've measured at this level of intelligence," and also noted was up roughly 40 percent from Gemini 3.7 Flash despite an unchanged per-token rate. Run the arithmetic backward and the prior model was landing around $0.41 for the same class of task. The reason for the jump: output tokens per task rose about 30 percent, averaging near 48,000 tokens, as the model took more turns and reasoned longer to earn the higher intelligence score. The rate card never moved. The bill did.

This is not a Google problem specifically, and it is not new. It is what happens whenever a vendor competes on reasoning depth instead of raw throughput. A model that "thinks" more to get smarter will, by definition, generate more billable tokens doing it, and a buyer who benchmarks the vendor's rate card instead of their own workload will be the last to notice. The introductory-to-standard pricing structure here, doubling to $1.50 and $7.50 on January 1, is disclosed upfront in the announcement and is standard practice, closer to a cloud reserved-instance discount than a bait and switch. That part is fine. The part that is not fine is treating the promotional per-token number as if it predicts your invoice.

There is a useful counterpoint in the same market. Anthropic had a scheduled price increase for Claude Sonnet 5 set for September 1, moving it from $2 and $10 per million tokens to $3 and $15, and quietly canceled it in August, leaving the launch pricing permanent. Anthropic has not stated a reason publicly, but the timing, weeks after Google and OpenAI both had comparably or more aggressively priced models on the market, makes competitive pressure the more plausible read than a change of heart. That is speculation on my part, not a confirmed motive. What is not speculation is the mechanism: nobody regulated that reversal into existence. A market with real alternatives did the disciplining, which is exactly how competition is supposed to work: badly priced products lose customers to well-priced ones, and the correction happens in weeks, not through a rulemaking docket. The fix for a confusing pricing landscape is more benchmarking firms doing the arithmetic in public, not a mandate standardizing how vendors publish rate cards, which would mostly just raise the compliance bar for the model providers too small to have a pricing team.

The practical move, if you are choosing a model for real production work: stop comparing announcement-page dollar figures across vendors, and instead run your own representative workload through each candidate model and measure the actual token count and completion cost. A model that costs half as much per token can easily cost more per finished job. That comparison is the whole point of picking a model deliberately instead of defaulting to whichever one is loudest this week, which is a chunk of what an AI vendor decision project looks like in practice. If you want a second set of eyes on that math before you commit budget to a model, that is a conversation worth having: get in touch.

Sources

Every factual claim above is drawn from these independently published sources, linked inline where first referenced.

Let's talk

Tell us what's on your mind.

You don't need a polished brief to reach out. A two-line email about what's bugging you is plenty; we'll tell you straight if we're the right fit, and what we'd tackle first.

We'll scope the work around your workflow, goals, and timeline before quoting anything, so you know what's included before committing.

LocationBoca Raton, Florida
CoverageSouth Florida + remote nationwide
Status Now accepting clients