DEV Community

daniel
daniel

Posted on Originally published at ai-info.fastget.link

Your DeepSeek V4 Bill Can Swing 2x by the Hour — the Actual Math Across Four APIs

Full disclosure up front: I run ai-info.fastget.link, where I track AI API pricing and publish the DeepSeek V4 cost calculator this post is based on. All rates below were checked against official pricing pages and live dashboards on 2 Sep 2026 — if you're reading this later, treat them as a snapshot, not gospel. No affiliate links anywhere.

The pricing page is lying to you (sort of)

DeepSeek doesn't have one price for V4. It has four, and which one you pay depends on decisions most developers never consciously make.

Take V4 Flash, the cheap workhorse model. DeepSeek's official list rate is $0.33 input / $0.99 output per million tokens — except that's not what you'll actually pay. DeepSeek runs peak/off-peak billing: off-peak hours cost ×0.67 of list ($0.22/$0.66), peak hours cost ×1.33 ($0.44/$1.32). So the same API call costs $0.22/M input at 3 AM and $0.44/M at 2 PM. Same model, same tokens, 2× apart. (The exact windows and how they map to your timezone are in my off-peak pricing breakdown.)

And that's just DeepSeek's own portal. The same model ID is resold elsewhere:

Provider Flash (in/out per M) Pro (in/out per M)
DeepSeek official (list) $0.33 / $0.99 $0.99 / $2.97
DeepSeek official, off-peak (×0.67) $0.22 / $0.66 $0.66 / $1.98
DeepSeek official, peak (×1.33) $0.44 / $1.32 $1.32 / $3.96
OpenRouter (cheapest host) $0.05 / $0.16 $0.58 / $1.74
Novita flat rate, see dashboard flat rate, see dashboard

Yes, you read the OpenRouter row right. The cheapest host on OpenRouter serves deepseek-v4-flash-0731 at $0.05/$0.16 — roughly a quarter of DeepSeek's off-peak rate. For Pro (deepseek-v4-pro-0813), the cheapest route is Alibaba at $0.58/$1.74, which actually undercuts DeepSeek's own off-peak rate.

So which number ends up on your invoice? That depends on three choices:

1. Official vs. reseller is a cache trade, not just a price trade

The per-token table says "reseller wins." But look at cache pricing before you switch: DeepSeek official sells cache hits at $0.007–$0.044/M — dramatically cheaper than anything resellers offer. If your app reuses long prompts (RAG with a fat system prompt, agents with big tool schemas, long-running chat threads), your real bill at DeepSeek official can come in far below the naive per-token estimate, sometimes below the reseller's headline rate.

Rule of thumb I've settled on: high cache-hit workload → official; stateless, prompt-diverse workload → the cheapest reseller route.

2. Peak windows are a scheduling problem, and scheduling is free

DeepSeek's off-peak discount is the cheapest optimization in the entire AI API market, because it costs zero code quality. Batch jobs, eval runs, embeddings backfills — none of them care whether they run at 3 AM. If half your Flash tokens are batch-shaped, moving them off-peak cuts that half's cost by half ($0.22 vs $0.44 per million input — a 2× spread).

The Pro model makes this more dramatic: official Pro ranges from $1.98/M output off-peak to $3.96/M at peak. Your nightly reasoning batch and your midday interactive traffic are not the same product and shouldn't share a price.

3. Watch the model ID, not the marketing name

One trap worth its own section: on OpenRouter, the undated deepseek-v4-pro ID routes from $0.87/$1.74 — above DeepSeek's own off-peak rate — while the dated deepseek-v4-pro-0813 starts at $0.58/$1.74 via Alibaba. Same model, different ID, different routing pool, different floor price. Pick by model ID and check which host actually serves the route; OpenRouter's "cheapest host" changes and per-route pricing isn't uniform.

Running your own numbers

The calculator on my site takes your monthly input/output token volumes and returns totals for both models across providers, with the peak/off-peak midpoint math built in. To give you scale: at a fairly typical solo-project load of 10M input / 5M output tokens per month on Flash, the spread between OpenRouter's cheapest host and DeepSeek's peak-hour official rate is the difference between roughly $1 and $11 a month. On Pro, the same spread is tens of dollars.

One caveat on the table: Krater.ai is subscription-based ($20/mo for 1,500 credits across 350+ models), so it's excluded from per-token comparisons. If your usage is spiky rather than steady, a subscription can beat all the per-token rates above — but that's a different math problem.

The short version

  • The "DeepSeek V4 price" is a range, not a number: ×0.67 off-peak to ×1.33 peak on official, and resellers undercut even the off-peak rate by 4× on Flash.
  • Cache-heavy workloads belong on DeepSeek official ($0.007–$0.044/M cache hits); stateless workloads should shop routes.
  • Schedule batch work into off-peak windows — it's the only API optimization that costs nothing.
  • On OpenRouter, quote the dated model ID and check the host.

Prices shift frequently on every one of these providers — I re-verify against official pages and dashboards regularly, and the calculator gets updated when they move. If a number here looks stale, the calculator page is the live copy.

Cross-posted from ai-info.fastget.link, my site on AI API pricing and decisions — the canonical version of this post lives there.

Top comments (0)