· AI
Sakana Says Its AI Router Is 40 Percent Cheaper. The Router's Own Math Says Otherwise.
The advertised price is the honest number for the tokens you can see. Most of what Fugu actually bills you for, you never see.
Sakana AI shipped two new models on September 11: Fugu Max and Fugu Ultra v2, built on a simple pitch. Instead of one large model answering every question, a smaller orchestrator model reads your prompt, assembles a scratch team of open-weight and specialist models, and stitches their work into one answer, all behind a single OpenAI-compatible endpoint. Sakana's own announcement says Fugu Max hits the best overall score on six benchmarks with output pricing "40 to 60 percent lower" than Claude Sonnet 5, GPT 5.6 Terra, and Kimi K3, priced at $2 per million input tokens and $6 per million output. Fugu Ultra v2 costs more, $5 and $30 per million tokens, and claims best or joint-best scores on five of eight harder benchmarks.
My position: the architecture is a real and useful proof that you do not need one enormous model to get frontier-level output, which is the bet this blog has been making about open models for months. But the price comparison Sakana is advertising is close to fiction for a lot of real workloads, and two independent teams that opened the hood found out why.
Steelman the approach first, because it earns one. Routing a question to the cheapest model capable of answering it, and saving the expensive reasoning for genuinely hard problems, is the correct instinct. It is the same logic behind caching and distillation: most queries are routine, so pay routine prices for them. Sakana trained Fugu to make that routing call itself instead of hand-coding it, and got real results doing so, beating benchmarks with no closed frontier model in its pool at all, by its own account. That is a legitimate engineering achievement, and it should worry anyone selling a single, all-purpose AI subscription as the only serious option.
Here is what the sticker price leaves out. Requesty's reverse-engineering of Fugu Ultra found a simple query carries a fixed overhead of roughly 1,260 orchestration tokens no matter how small the question is, pushing total token consumption to 5.3 to 5.8 times what shows up in the reply. On a harder prompt, comparing Python and Rust, the visible answer ran 2,223 tokens. The total the system processed to produce it: 22,710 tokens, a multiplier near 10x. Sakana's advertised per-token discount is real, but it is a discount on the wrong number if the coordination happening behind the answer runs through roughly ten times as many tokens as the reply implies. The same investigation prompted an individual worker with a debugging system message and got a reply identifying itself as a Gemini model, meaning at least part of what Fugu sells as its own orchestration is buying frontier capacity wholesale and marking it up.
A second hands-on review, published by Paddo, clocked a plain question taking 108 seconds to answer, with 60 percent of the billed tokens going to orchestration the user never reads. Its verdict after users ran the models against real tasks: "Cheaper and faster, not better" than calling a single model directly. That was not the whole picture. One test in the same piece had Fugu Ultra finish a coding task in 22 minutes for $7.32 against Claude Opus at 79 minutes for $37.85, a genuine win. The honest read is that Fugu's economics swing hard by task type, which is exactly the kind of detail a single advertised discount percentage is built to hide.
None of this needed a regulator to surface it. Two independent shops spent their own compute reverse-engineering a vendor's pricing claim within two days of release and published the receipts for free. That is the market doing its actual job. Competition and curiosity caught a pricing claim faster than a mandated disclosure rule could have gotten through a single rulemaking comment period, and any such rule would have applied unevenly across vendors anyway.
The practical takeaway if you are shopping model routers instead of single subscriptions: ask for total tokens consumed per completed task, not price per visible output token, and run your own task mix through it before signing anything longer than a month. That is the same audit worth running for a business choosing between a frontier subscription and a routed or local setup, because the invoice that matters is for the work you actually needed done, not for the tokens a vendor chose to show you. If that math is worth a second pair of eyes before you commit budget to it, that conversation starts at /#contact.
Sources
References used in this article. Links also appear alongside the relevant claims.
Tell us what's on your mind.
You don't need a polished brief to reach out. A two-line email about what's bugging you is plenty; we'll tell you straight if we're the right fit, and what we'd tackle first.
We'll scope the work around your workflow, goals, and timeline before quoting anything, so you know what's included before committing.