DeepSeek is raising its API prices and adding surge pricing. Here is what you will actually pay

With the launch of its new V4 models, DeepSeek is ending the flat, ultra-cheap pricing widely credited with helping pull AI prices down. From 16:00 UTC on 16 August it charges peak and off-peak rates, and output on its cheapest model now costs more than double off-peak and nearly five times the old rate during a seven-hour daily peak window. Here is exactly what changes, and how to pay the lower rate.

DeepSeek is raising its API prices and adding surge pricing. Here is what you will actually pay
TL;DR

DeepSeek is overhauling its API pricing with its new V4 models. From 16:00 UTC on 16 August 2026 it charges peak and off-peak rates, with off-peak set at half of peak. V4-Flash output rises from a flat $0.28 per million tokens to $0.66 off-peak or $1.32 at peak; V4-Pro output goes from $0.87 to $1.98 or $3.96. Peak covers just seven hours a day.

DeepSeek built its reputation on being shockingly cheap. That era is ending. Alongside the release of its new V4 model lineup, the company is tearing up its flat, low pricing and replacing it with a peak and off-peak system that, at its worst, costs several times what developers were paying the week before. The new rates took effect at 16:00 UTC on 16 August. Here is what actually changed, and where the increases hurt most.

What is DeepSeek changing?

The core change is a move from one flat price per model to time-based pricing. DeepSeek now charges a higher peak rate and a lower off-peak rate, with off-peak set at half of peak, in its own words to "allocate resources more reasonably" and push developers to run work outside the busiest hours.

The change arrived alongside the general availability of DeepSeek's V4 lineup, split into a cheaper V4-Flash and a more capable V4-Pro. Both now carry separate peak and off-peak prices for input and output tokens, and both cost more than the flat rates they replace.

How much more will you pay?

Here are the old flat prices against the new peak and off-peak rates, per million tokens, taken from DeepSeek's official pricing page:

Per 1M tokensOld (flat)New off-peakNew peak
V4-Flash, output$0.28$0.66$1.32
V4-Flash, input (cache miss)$0.14$0.22$0.44
V4-Pro, output$0.87$1.98$3.96
V4-Pro, input (cache miss)$0.435$0.66$1.32

On the number most developers watch, output tokens, that is roughly 2.3 to 2.4 times the old price off-peak and about 4.5 to 4.7 times at peak. The steepest increase is not in this table: on cache-hit input tokens, the cheapest tier of all, the peak rate is roughly twelve times the old one, which is where the "up to 1,100 percent" figures in the coverage come from. Those tiny per-token numbers add up quickly at scale.

When are the peak hours, and how do you avoid them?

DeepSeek's pricing page defines peak hours as 01:00 to 04:00 and 06:00 to 10:00 UTC. Everything else is off-peak. That detail matters: peak is only seven hours out of every twenty-four, so most of the day will bill at the lower off-peak rate.

A workload that can be scheduled, batch jobs, evaluations, offline data processing, can sidestep the peak windows entirely and pay the off-peak rate, which is half of peak, though still more than double the old flat price. It is interactive, real-time traffic, the kind that has to run whenever users show up, that will feel the full peak price. If your usage is flexible, the practical takeaway is simple: keep it out of those two windows.

Why is DeepSeek doing this?

DeepSeek frames it as capacity management. Its cheap models have been hugely popular, and reporting on the change ties it to demand straining the company's compute. There is history here too: DeepSeek's promotional pricing had been due to expire earlier in the year, and the company had said it would make the low rates permanent before reversing course.

The blunt reading, which DeepSeek does not offer, is that serving frontier-class models at prices this low was never going to last, and the compute bill is now being handed to developers. It also dents one of DeepSeek's main selling points. Much of its draw was being dramatically cheaper than the big American labs. The new rates are still low in absolute terms, especially off-peak, but the "almost free" framing no longer holds.

What it means for the AI price war

DeepSeek is widely credited with helping trigger a race to the bottom on AI prices, with rivals cutting their own rates to keep up. A price rise from that same company is a notable turn. It does not mean prices are about to jump everywhere, DeepSeek's off-peak rates are still low, but it is a sign that the underlying economics are starting to bite. Someone has to pay for the compute, and "cheap forever" was always more a marketing position than a law of physics.

For the wider question of whether the money behind all this adds up, see our look at whether AI is a bubble. For where the tools stand today, see our ranking of the best AI coding tools.

Frequently asked questions

How much are DeepSeek's new API prices?

Per million tokens, V4-Flash output is $0.66 off-peak and $1.32 at peak (up from a flat $0.28), and V4-Pro output is $1.98 off-peak and $3.96 at peak (up from $0.87). Cache-miss input is $0.22 / $0.44 for V4-Flash and $0.66 / $1.32 for V4-Pro. Off-peak is always half the peak rate.

When did DeepSeek's new pricing take effect?

At 16:00 UTC on 16 August 2026, alongside the launch of its V4 model lineup.

What are DeepSeek's peak hours?

01:00 to 04:00 and 06:00 to 10:00 UTC, seven hours a day in total. All other hours are billed at the lower off-peak rate.

Why is DeepSeek raising prices?

DeepSeek says it is to "allocate resources more reasonably" and to encourage developers to schedule work outside the busiest hours. Reporting links the move to surging demand straining its compute capacity.

Is DeepSeek still cheap?

In absolute terms its off-peak rates are still low, but the era of flat, ultra-low pricing is over. Output now costs just over two to nearly five times what it did before, depending on the model and the time of day.