Pricing verified: 17 August 2026. Compiled from DeepSeek’s official API documentation, read after the change took effect, and compared against archived captures of the same page taken on 6, 12 and 14 August 2026 — before it. We do not hold a DeepSeek API account and have not benchmarked these models ourselves. Every rate below is a published vendor number, and every percentage, cost projection and timezone conversion built on top of those rates is our own arithmetic and is labelled where it appears.

Compiled and fact-checked by the AI Tools Worth editorial team. Corrections: contact page.

The increase is live as of yesterday

DeepSeek’s new API pricing took effect at 16:00 UTC on 16 August 2026. It is not a proposal or a scheduled change — it is the rate you are being billed at right now.

Two things changed at once. Every published rate went up, and the flat per-token price was replaced by peak and off-peak billing, where the rate you pay depends on what time of day your request lands. DeepSeek’s documentation now states only: “Off-peak rates are half of the peak rates. Peak hours are 01:00 – 04:00 and 06:00 – 10:00 UTC (all other hours are off-peak).”

The warning ran for at least ten days. The 6 August capture of the pricing page already carried the note: “We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected. Please plan your usage accordingly.” By the 14 August capture that had been replaced with the exact new table and the 16:00 UTC effective time. The 13 August announcement of the DeepSeek-V4-Pro-0813 release put it more gently: DeepSeek said it was “updating our API pricing and introducing peak and off-peak rates. Off-peak rates are 50% lower than peak, enabling more flexible workload scheduling.”

Nothing else on the spec sheet moved. Both models keep a 1M context window, a 384K maximum output, and unchanged concurrency limits of 2,500 requests for V4-Flash and 500 for V4-Pro. This is a price change and only a price change.

The full before and after rate card

All figures per million tokens, in USD, exactly as published by DeepSeek. The “before” column is the flat rate that applied until 16:00 UTC on 16 August; it is taken from the archived captures of the same documentation page.

ModelRateBefore (flat)Now, off-peakNow, peak
deepseek-v4-flashInput, cache hit$0.0028$0.007$0.014
Input, cache miss$0.14$0.22$0.44
Output$0.28$0.66$1.32
deepseek-v4-proInput, cache hit$0.003625$0.022$0.044
Input, cache miss$0.435$0.66$1.32
Output$0.87$1.98$3.96

One structural detail worth noticing: V4-Pro’s off-peak cache-miss input and output rates — $0.66 and $1.98 — are identical to V4-Flash’s peak rates of $0.66 and $1.98 for cache-miss input and output respectively. Scheduling a V4-Pro job into off-peak hours now costs the same per token as running V4-Flash at peak.

Where the “52% to 1,114%” range actually lands

The increase has been widely reported as a range running to roughly 1,100%, which is accurate but not very useful on its own, because the two ends of that range describe completely different line items. The table below is our arithmetic on the published rates above, showing every rate change individually.

ModelRateIncrease, off-peakIncrease, peak
V4-FlashInput, cache hit+150.0%+400.0%
Input, cache miss+57.1%+214.3%
Output+135.7%+371.4%
V4-ProInput, cache hit+506.9%+1,113.8%
Input, cache miss+51.7%+203.4%
Output+127.6%+355.2%

The floor and the ceiling both belong to V4-Pro. The smallest increase in the whole table is cache-miss input off-peak, up 51.7% from $0.435 to $0.66. The largest is cache-hit input at peak, up 1,113.8% from $0.003625 to $0.044 — a twelvefold rise on what was previously the cheapest line item DeepSeek sold.

That 1,114% figure is real, but it applies to a rate that starts at a third of a cent. Nobody’s bill went up twelvefold. What your bill actually did depends on the mix below.

Prompt caching made this worse, not better

The usual intuition is that heavy prompt caching insulates you from a price rise. Here it does the opposite, because cache-hit input rose far more steeply in percentage terms than anything else. The more of your spend sat in cached tokens, the larger your increase.

Three worked examples on V4-Pro and V4-Flash, all our arithmetic on the published rates, all at a monthly volume of 100M input and 20M output tokens.

WorkloadBeforeNow, all off-peakNow, all peak
V4-Flash, no caching
100M cache-miss in, 20M out
$19.60$35.20 (+79.6%)$70.40 (+259.2%)
V4-Pro, no caching
100M cache-miss in, 20M out
$60.90$105.60 (+73.4%)$211.20 (+246.8%)
V4-Pro, 90% cache hit rate
90M cached + 10M cache-miss in, 20M out
$22.08$48.18 (+118.2%)$96.36 (+336.5%)

Compare the second and third rows. Identical token volumes on the identical model; the only difference is that one caches 90% of its input and the other caches none. The uncached workload went up 73.4% off-peak. The cached one went up 118.2%.

The cached workload is still much cheaper in absolute terms — $48.18 against $105.60 — so caching remains worth doing. But if you are trying to reconcile a bill against a forecast, a heavily cached agent pipeline will have overshot by considerably more than a simple percentage applied across the board would predict.

“Peak” is only 7 hours of the 24

The word peak implies most of the working day. It does not. Peak is 01:00–04:00 and 06:00–10:00 UTC — three hours plus four hours, 7 hours in total. The remaining 17 hours, just over 70% of the day, are off-peak, including the two-hour gap from 04:00 to 06:00 UTC that sits between the two peak blocks.

If your traffic is spread evenly around the clock and you change nothing, roughly 71% of it now bills at the off-peak rate. The default outcome is closer to the off-peak column than the peak one.

An observation, offered as our inference rather than a vendor statement. Converted to Beijing time, UTC+8, the peak windows are 09:00–12:00 and 14:00–18:00 — a standard Chinese working day, with the 12:00–14:00 lunch break falling neatly into the off-peak gap. DeepSeek has published no rationale for the window boundaries, so this is a reading of the arithmetic, not a claim about intent. It does suggest the windows track domestic demand on DeepSeek’s own infrastructure rather than anything about your workload.

Which rate you pay depends on where you are

Because the windows are fixed in UTC, the same 9-to-5 workload lands in completely different rate bands depending on the timezone it runs from. The conversions below are ours, for August 2026, with northern-hemisphere daylight saving in effect where it applies.

TimezonePeak windows, local timeEffect on a 09:00–17:00 working day
US Pacific (PDT, UTC−7)18:00–21:00 and 23:00–03:00Entirely off-peak
US Eastern (EDT, UTC−4)21:00–00:00 and 02:00–06:00Entirely off-peak
UK (BST, UTC+1)02:00–05:00 and 07:00–11:0009:00–11:00 peak — 2 hours of 8
Central Europe (CEST, UTC+2)03:00–06:00 and 08:00–12:0009:00–12:00 peak — 3 hours of 8
India (IST, UTC+5:30)06:30–09:30 and 11:30–15:30Roughly half the day peak
China (CST, UTC+8)09:00–12:00 and 14:00–18:00Almost entirely peak
Japan (JST, UTC+9)10:00–13:00 and 15:00–19:00Most of the day peak

The practical upshot for a US-based team is worth stating plainly: a workload driven by North American office hours never touches the peak rate at all. Interactive daytime traffic from either US coast bills at off-peak throughout, and the relevant increase is the off-peak column — roughly 52% to 507% by line item, not the headline figure.

A European team has the opposite problem. The 08:00–12:00 CEST block is the busiest part of a working morning and it is now the expensive one. Overnight batch work run from Europe, by contrast, lands mostly in off-peak.

How DeepSeek now compares with GPT-5.6

DeepSeek’s position has always rested on undercutting the US labs by a wide margin. After this change, that is still true of its frontier model and no longer true of its cheap one.

OpenAI’s published GPT-5.6 rates, standard service tier, short-context pricing, read the same day: Luna at $0.20 input / $0.02 cached input / $1.20 output, and Terra at $2.00 / $0.20 / $12.00 per million tokens.

ComparisonOff-peakPeak
V4-Flash cache-miss input vs Luna ($0.20)$0.22 — 10% more expensive$0.44 — 120% more expensive
V4-Flash output vs Luna ($1.20)$0.66 — 45% cheaper$1.32 — 10% more expensive
V4-Flash cached input vs Luna ($0.02)$0.007 — 65% cheaper$0.014 — 30% cheaper
V4-Pro cache-miss input vs Terra ($2.00)$0.66 — 67% cheaper$1.32 — 34% cheaper
V4-Pro output vs Terra ($12.00)$1.98 — 83.5% cheaper$3.96 — 67% cheaper

Read the first two rows carefully, because they are the headline. DeepSeek V4-Flash is now more expensive than GPT-5.6 Luna on input tokens at every hour of the day, and more expensive on output tokens during peak hours. Its remaining advantage over Luna is cached input, where it is still 65% cheaper off-peak, and output during the 17 off-peak hours.

V4-Pro is a different story. Even at the peak rate it undercuts Terra by 34% on input and 67% on output, and off-peak it is cheaper by 67% and 83.5%. The frontier-tier discount survived the increase largely intact.

What this comparison is not. It is a price comparison and nothing more. We have not evaluated whether V4-Flash and GPT-5.6 Luna produce comparable output on any task, and a cheaper model that needs two attempts is not cheaper. We have also used OpenAI’s standard short-context tier; OpenAI charges more above its context threshold and less on Batch and Flex, while DeepSeek publishes one context tier at 1M tokens.

Timeline

DateWhat happened
6 August 2026Pricing page already warns of “a significant increase” with no figures or date (archived capture)
12 August 2026Same warning still in place, rates unchanged at $0.14 / $0.28 and $0.435 / $0.87 (archived capture)
13 August 2026DeepSeek-V4-Pro-0813 released; announcement confirms peak/off-peak billing is coming
14 August 2026Pricing page now shows the exact new table and the 16:00 UTC effective time (archived capture)
16 August 2026, 16:00 UTCNew rates take effect
17 August 2026Our check: the new table is live and the old flat rates are gone from the page

Frequently asked questions

How much did DeepSeek raise its API prices?

Between 51.7% and 1,113.8% depending on the exact line item, model and time of day (our arithmetic on DeepSeek’s published rates). The smallest increase is V4-Pro cache-miss input off-peak, from $0.435 to $0.66. The largest is V4-Pro cache-hit input at peak, from $0.003625 to $0.044. A typical uncached workload run entirely off-peak went up around 73–80%.

When did the new DeepSeek pricing take effect?

16:00 UTC on 16 August 2026, as stated on DeepSeek’s own pricing documentation before the change. The new rates were live when we checked on 17 August.

What are DeepSeek’s peak and off-peak hours?

Peak is 01:00–04:00 and 06:00–10:00 UTC. Everything else is off-peak, which is 17 of the 24 hours. Off-peak rates are exactly half of peak rates.

What does deepseek-v4-pro cost now?

Off-peak: $0.022 per million cached input tokens, $0.66 per million cache-miss input tokens, and $1.98 per million output tokens. Peak rates are double those: $0.044, $1.32 and $3.96.

What does deepseek-v4-flash cost now?

Off-peak: $0.007 cached input, $0.22 cache-miss input, $0.66 output, per million tokens. Peak: $0.014, $0.44 and $1.32.

Is DeepSeek still cheaper than OpenAI?

For the frontier model, clearly yes — V4-Pro undercuts GPT-5.6 Terra by 67–83.5% off-peak and 34–67% at peak. For the cheap model, no longer on input: V4-Flash cache-miss input is $0.22 off-peak against Luna’s $0.20, and $0.44 at peak. This is a price comparison only and says nothing about output quality on your workload.

Does prompt caching still save money on DeepSeek?

Yes in absolute terms, but the increase hit cached tokens hardest. On our worked example, a 90%-cached V4-Pro workload still costs less than half the uncached equivalent, while its percentage increase was larger — 118.2% against 73.4% off-peak.

Can I avoid the peak rate?

If you can schedule work, yes — 17 hours of every day are off-peak, and DeepSeek’s own framing of the change points at “more flexible workload scheduling.” For teams working North American hours, interactive daytime traffic already falls entirely in off-peak with no changes at all.

What we could not verify

Gaps we could not close, listed here rather than papered over.

  • Whether billing follows the time a request starts or the time it completes — DeepSeek does not say. For a long generation crossing a boundary at 10:00 UTC, the difference is a doubled or halved rate, and we found no published rule.
  • Whether existing balances or granted credits get any transitional treatment — no statement in the announcement or the pricing page.
  • DeepSeek’s stated reason for the increase — the announcement describes the mechanism and its scheduling benefit but gives no rationale for the size of the rise.
  • The pricing table in the 13 August announcement post — it is published as an image we could not read. The figures we use come from the documentation page, which carried the same rates in text.
  • Whether the peak windows are fixed — DeepSeek reserves the right to adjust prices and has not said whether the hours themselves are stable.
  • Whether DeepSeek’s consumer app or web subscription pricing changed — we checked the API documentation only.
  • Whether resellers and aggregators have passed the increase through — third-party platforms set their own rates and we did not survey them.
  • Our Beijing-hours reading of the peak windows — the UTC+8 conversion is arithmetic, but the inference that the windows track the Chinese working day is ours and DeepSeek has not confirmed it.

Sources and method

Compiled 17 August 2026 from DeepSeek’s official API pricing documentation, read after the change took effect and fetched with a cache-busting query string to confirm we were not served a stale copy. The “before” rates come from three archived captures of the same page, dated 6, 12 and 14 August 2026, which let us establish both the old figures and the exact wording of the advance notice. The effective time and the peak/off-peak definition are quoted from the 14 August capture, which published them ahead of the change, and from the live page, which now carries the definition alone. The 13 August DeepSeek-V4-Pro-0813 release announcement supplied the vendor’s own framing of the change. OpenAI’s GPT-5.6 comparison rates were read the same day from OpenAI’s published API pricing documentation, standard tier, short-context. All percentages, cost projections and timezone conversions are our own arithmetic on those published rates and are labelled where they appear. Prices change — verify against the vendor before you commit.

THE 5-MINUTE AI BRIEF
Know which AI tools are actually worth it — in one weekly email

Assessed verdicts, real price changes and the launches that matter. No hype, no spam — unsubscribe anytime.

Free forever. We never share your email. By the AI Tools Worth editorial team.
THE 5-MINUTE AI BRIEF
Weekly verdicts on AI tools worth paying for — free, no hype