August 13, 2026 · AI
DeepSeek Raised Prices Up to 12x. The Number That Matters Is Smaller.
The headline increase is real, but DeepSeek built a business-hours discount into its own price hike, and almost nobody covering the story is running that math.
DeepSeek just did something its own customers did not expect from the company that started the AI price war: it raised prices, sharply, on the models that made it famous for being cheap.
According to DeepSeek's own API pricing page, new rates take effect August 17 at 16:00 UTC. V4-Flash output climbs from $0.28 to as much as $1.32 per million tokens. V4-Pro's cached input rate goes from $0.003625 to $0.044 per million tokens, a jump of roughly 1,114 percent, which is almost certainly where the "up to 1,100 percent" figure reported by TechStartups and WTAQ, via Reuters, came from. Both outlets confirm the range as 50 to 1,100 percent depending on model, token type, and time of day.
That last clause, time of day, is the part getting buried under the scary headline number, and it is the part a business owner actually needs.
Here is why the increase exists at all. XenoSpectrum reported that DeepSeek founder Liang Wenfeng told an investor meeting the company's entire compute fleet amounts to roughly 20,000 H100-equivalent GPUs, most of it acquired in the preceding one to two months. The ultra-cheap V4-Flash model, released July 31, processed 7.22 trillion tokens on OpenRouter in a single week, according to the same report, enough to take the top spot in OpenRouter's rankings outright. On August 4, that demand degraded the API to the point of being nearly unusable before DeepSeek restored service. A company cannot serve 7 trillion tokens a week on 20,000 GPUs at $0.14 per million input tokens forever. Something had to give, and DeepSeek chose price over further outages.
I think that is the right call, and here is the steelman for the harder read before I make that case. A business that built its cost model around V4-Flash at $0.14 got four days' notice of an increase running as high as 12x on some token types. If you signed clients, set your own pricing, or sized a project against that number, you got blindsided, and "the vendor ran out of GPUs" does not make your invoice smaller. An analyst quoted by the South China Morning Post, Michael Guo, went further, questioning whether DeepSeek can afford this move at all, since Meta's Muse Spark and OpenAI's GPT-5.6 Luna are now competitive with DeepSeek on both price and capability. Raising rates while cheaper Western alternatives exist is a real business risk for DeepSeek, not a footnote. If you built a cost strategy on the assumption that the cheapest model today stays the cheapest model forever, this week is the proof that assumption was wrong.
Now the part worth doing the math on. DeepSeek's new structure prices peak hours at double the off-peak rate, and peak hours are set at 01:00 to 04:00 UTC and 06:00 to 10:00 UTC. Convert that to US Eastern time, where most of my clients operate: peak windows land at 9 PM to midnight and 2 AM to 6 AM. Off-peak covers the entire American workday, roughly 6 AM to 9 PM Eastern. A company processing 50 million output tokens a month on V4-Flash during business hours goes from $14 a month to $33 at the new off-peak rate, not the $66 you get by pricing everything at peak. That is still a real increase, more than double, and I am not going to pretend it is not. But it is a different story than the 1,100 percent number every headline is running.
The takeaway is not "DeepSeek is fine" or "DeepSeek is finished." It is that a rate card built for 20,000 GPUs and 7 trillion tokens a week was never going to hold still, and a company that only priced against yesterday's number just found that out at 4 days' notice. If any part of your operation depends on a single AI vendor's current price holding for the next year, this is the week to build a second option into the plan, whether that is a competing model, a self-hosted open-weight model for routine work, or simply scheduling batch jobs into whichever provider's cheap window fits your day. That is the actual discipline behind model routing: not chasing the lowest number on a pricing page, but knowing what your workload costs under three different providers before you need the backup. If you want a second opinion on where your AI spend is exposed to a single vendor's next rate change, that is a conversation worth having before the next price hike, not after it.
Sources
Every factual claim above is drawn from these independently published sources, linked inline where first referenced.
Tell us what's on your mind.
You don't need a polished brief to reach out. A two-line email about what's bugging you is plenty; we'll tell you straight if we're the right fit, and what we'd tackle first.
We'll scope the work around your workflow, goals, and timeline before quoting anything, so you know what's included before committing.