· AI

DeepSeek V4.1 Flash Costs 27 Cents a Task. Don't Get It From DeepSeek.

The cost story is real. Getting it through DeepSeek's own app is the one way to also inherit the risk.

DeepSeek released V4.1 Flash on September 10, and the number that matters is 27 cents. That is what Artificial Analysis measured as the model's average cost to complete one task on its Intelligence Index, a composite benchmark of reasoning, coding, and knowledge tests the firm runs against every major model. The model scores 40 on that index, respectable but well behind the frontier: GPT-6 Astra and Claude Fable 5.1 tied for first at 53 in Artificial Analysis's comparison the week before. Raw token pricing backs up the cost story. Off peak, it runs 15 cents per million input tokens and 60 cents per million output tokens, doubling during the daytime demand window. It is a mixture-of-experts model, 552 billion total parameters with only 16 billion active per token, which is most of why it is cheap: most of the network sits idle on any given request.

Here is my position: the interesting story is not that a Chinese lab shipped a cheap model again. DeepSeek has been doing that since V3 landed. The interesting story is that the smart way to use this one has almost nothing to do with using DeepSeek.

Steelman the skeptics first, because part of their case is solid. A model scoring 40 against a frontier 53 is not a free lunch. It is a different tool for different jobs, and treating cost per task as the only number that matters is how a company ends up shipping a support bot that gets basic facts wrong to save fractions of a cent. There is a harder version of the objection too: open weights are not a security guarantee by themselves. Nobody outside DeepSeek has audited what the base model was trained on or whether a checkpoint could carry a hidden behavior. Self-hosting removes the data-in-transit risk. It does not remove provenance risk. That is fair, and my honest answer is the same one every open model deserves: eval it, sandbox it, and do not point it at anything untested.

Where I part ways with the skeptics is on the fallback they usually land on: fine, use DeepSeek, just go through the official app or API. That is the actual mistake. DeepSeek's own privacy policy states plainly that it stores personal data in the People's Republic of China, and China's 2017 National Intelligence Law says, in its own Article 7 language, that "any organization or citizen shall support, assist, and cooperate with the state intelligence work." That is not a hypothetical clause nobody invokes. New York, Texas, and Virginia have already banned DeepSeek from state networks, and South Korea's data protection regulator documented user prompts being transferred to a Chinese vendor without consent. TechTarget's guidance to security teams is blunt: treat unauthorized use of foreign-hosted models, DeepSeek by name, as a live risk, not a theoretical one.

None of that touches the weights themselves. DeepSeek ships V4.1 Flash under an MIT license, so any US-based inference provider is free to host the identical model on its own infrastructure. Fireworks advertises deploying new DeepSeek releases within hours of launch on its own servers, no prompt ever touching a server in China, and it is not the only US-based host building on DeepSeek's open weights. You can get the 27-cent-a-task economics and the full 552-billion-parameter model without a byte leaving US jurisdiction, and without paying DeepSeek anything at all.

That is the real case for open weights, and it is sharper here than almost anywhere else in AI right now. The state bans and the No Adversarial AI Act working its way through Congress, which would bar Chinese AI tools from federal agencies, are trying to solve a genuine problem, data sovereignty, with a blunt instrument, prohibition. The market already solved it better. You do not need an act of Congress to keep your prompts out of Beijing. You need to read the license and pick a different host, which takes about ten minutes and costs nothing extra.

The practical takeaway if you are evaluating models this quarter: separate the model from the vendor. If V4.1 Flash fits a workload, capture the economics through Fireworks, another US-based host, or your own GPUs, and treat DeepSeek's own app the way you would treat any product built by a company whose government can legally compel it to hand over your data, meaning don't feed it anything you would not want read back to you by a stranger. That distinction, model versus vendor, is most of what we walk clients through at Mojo before they commit real spend to either side of it. If you are weighing a cheap model against a safe one and assuming they sit on the same axis, let's talk.

Sources

References used in this article. Links also appear alongside the relevant claims.

Let's talk

Tell us what's on your mind.

You don't need a polished brief to reach out. A two-line email about what's bugging you is plenty; we'll tell you straight if we're the right fit, and what we'd tackle first.

We'll scope the work around your workflow, goals, and timeline before quoting anything, so you know what's included before committing.

LocationBoca Raton, Florida
CoverageSouth Florida + remote nationwide
Status Now accepting clients