The pricing page now splits DeepSeek's lineup into two live routes, deepseek-flash and deepseek-v4-pro, and they are not aliases for each other. One serves V4.1 Flash, the other serves V4-Pro-0813. Different prices, different concurrency ceilings, different vision support. If you are wiring up a new integration, pick the ID that matches the workload instead of trusting whatever the old alias resolved to last quarter.
Here is what I look at before moving traffic.
Spec comparison: the architecture flipped
V4 Pro is a very large sparse MoE. V4.1 Flash is an asymmetric causal encoder-decoder MoE, where input processing activates fewer parameters than output generation. That split is deliberate: prefill and decode have different compute profiles, so there is no reason to run both stages through the same active-parameter budget.
| Spec | V4.1 Flash | V4 Pro 0813 |
|---|---|---|
| Architecture | Causal encoder-decoder MoE | Sparse MoE |
| Total parameters | 552B backbone; Hugging Face lists 763B model size | 1.6T backbone; Hugging Face lists about 1.7T model size |
| Active parameters | 8B input / 16B output | 49B per token |
| Context window | 1M tokens | 1M tokens |
| Max output | 384K tokens | 384K tokens |
| Thinking and non-thinking modes | Supported | Supported |
| Native image understanding | Supported | Not supported |
| Documented concurrency | 2,500 | 500 on the current configuration |
| API status | Active as deepseek-flash
|
Active as deepseek-v4-pro (V4-Pro-0813) |
Two numbers people keep conflating
The 552B in DeepSeek's model card counts backbone parameters. The 763B on Hugging Face is model size, a different accounting scope. Do not treat them as interchangeable totals in a capacity plan.
Similarly, the V4.1 Flash model card recommends max_tokens ≥ 256K for local inference. The API pricing page states a 384K maximum output for both models. A recommended local setting is not an enforced API ceiling.
Benchmarks: the win is execution-heavy, not universal
The published table favors Flash almost everywhere that involves running code, driving a terminal, or chaining tools. Percentage deltas below are my own arithmetic on the published scores; I skip them where the metric is a rating.
| Benchmark | V4.1 Flash | V4 Pro 0813 | Delta |
|---|---|---|---|
| Terminal-Bench 2.1 | 90.6 | 87.9 | +2.7 points, +3.1% |
| Terminal-Bench 3.0 | 30.0 | 11.8 | +18.2 points, +154.2% |
| Terminal-Bench 4.0 | 31.2 | 12.4 | +18.8 points, +151.6% |
| DeepSWE v1.1 | 74.2 | 62.7 | +11.5 points, +18.3% |
| ProgramBench | 20.3 | 15.5 | +4.8 points, +31.0% |
| NL2Repo-Bench | 64.0 | 61.5 | +2.5 points, +4.1% |
| CyberGym | 88.1 | 83.3 | +4.8 points, +5.8% |
| HLE with tools | 63.9 | 60.0 | +3.9 points, +6.5% |
| Automation-Bench | 54.8 | 43.2 | +11.6 points, +26.9% |
| Agents' Last Exam | 31.8 | 25.7 | +6.1 points, +23.7% |
| Codeforces rating | 3,471 | 3,348 | +123 rating points |
| GPQA Diamond | 90.9 | 92.4 | -1.5 points |
| HLE without tools | 36.8 (39.1 on text subset) | 42.7 on text subset | -3.6 points on the comparable text subset |
Read that as two different products rather than a ranking. Flash owns terminal work, repo-level coding, automation, security execution, and tool-augmented reasoning. Pro keeps its edge on closed-book reasoning, specifically GPQA Diamond and HLE without tools. Judge your workload by task mix, not by a single headline number.
Why the serving cost drops
Asymmetric activation
8B active for input, 16B for output, against 49B per token for V4 Pro. Prefill gets cheap, decode keeps enough capacity to stay useful.
KV cache
DeepSeek reports V4.1 Flash uses one quarter of the HBM and one eighth of the SSD of the previous generation's KV cache. The published chart puts global KV cache at 890 bytes per token versus 3,514 bytes for V4 Flash. Long-context services feel this immediately, because that footprint is what limits how many concurrent sessions fit on a node.
Vision and concurrency
Native image understanding plus a documented concurrency of 2,500, five times the current V4 Pro figure of 500. If you are building screenshot-driven agents, document extraction, or visual troubleshooting, this is the part that changes the design, not the benchmark table.
Pricing per 1M tokens
Off-peak / peak, from the current official schedule:
| Route | Cache-hit input | Cache-miss input | Output |
|---|---|---|---|
| V4.1 Flash | $0.003 / $0.006 | $0.15 / $0.30 | $0.60 / $1.20 |
| V4 Pro | $0.022 / $0.044 | $0.66 / $1.32 | $1.98 / $3.96 |
Roughly 4.4x on input and 3.3x on output. Prices move, so pin your budget model to the live page rather than to this table.
Migration
Use explicit model IDs
from openai import OpenAI
client = OpenAI(
api_key="YOUR_DEEPSEEK_API_KEY",
base_url="https://api.deepseek.com",
)
response = client.chat.completions.create(
model="deepseek-flash", # or "deepseek-v4-pro"
messages=[{"role": "user", "content": "Review this migration plan."}],
)
print(response.choices[0].message.content)
deepseek-v4-flash and deepseek-v4-flash-vision-exp still resolve to V4.1 Flash and bill at the new Flash price, but explicit naming keeps your dashboards and future rollbacks auditable. Use deepseek-v4-pro only when you intend to hit V4-Pro-0813.
If you already route through a unified multi-model endpoint such as CometAPI, the model string is the only switch to flip, though the evaluation work below still applies.
Re-evaluate, do not just smoke-test
- Run representative prompts in both thinking and non-thinking modes. Compare task success, token counts, and latency, not just HTTP 200.
- Retest tool schemas, JSON output, Responses API behavior, prefix completion, and FIM where you depend on them.
- Add image-input tests if the product will use native vision.
- Re-baseline budgets at both peak and off-peak rates. Do not carry V4 Pro assumptions forward.
- Track prompt-cache hit rate. At scale it dominates input cost.
Record what resolved
The deepseek-v4-pro route currently selects V4-Pro-0813. Log the resolved version, evaluation date, prompt set, and pricing window alongside your results so a future routing change does not invalidate the comparison silently.
Workload routing cheat sheet
| Workload | Pick | Why |
|---|---|---|
| Coding agents, repo maintenance | V4.1 Flash | DeepSWE, ProgramBench, NL2Repo, terminal scores |
| Tool-driven automation | V4.1 Flash | Automation-Bench, HLE with tools, agent scores |
| Image-aware assistants | V4.1 Flash | Native image understanding |
| High-throughput or cost-sensitive serving | V4.1 Flash | 2,500 concurrency, much lower token prices |
| Historical V4 Pro reproduction | V4 Pro | The route currently identifies V4-Pro-0813 |
| Pure closed-book reasoning | Test on your domain data | Pro stays ahead on GPQA Diamond and HLE without tools |
Traps I have hit
- Endpoint compatibility is not output compatibility. Reasoning paths, tool selection, token usage, and formatting all shift. Re-run production evals before trusting old thresholds.
- The two model IDs are not aliases. Do not swap them by config drift.
- Legacy V4 Flash IDs still work and bill at the new rate, so a forgotten environment variable will quietly move you to V4.1 Flash. That is usually what you want, but you should know it happened.
- Concurrency limits are documented, not guaranteed by your account tier. Check before designing a 2,500-way queue.
Originally published at cometapi.com

Top comments (0)