DEV Community

Ashraf
Ashraf

Posted on

DeepSeek V4.1 Flash Costs 36x Less Than Claude Opus 5 and Matches It on SWE Benchmarks. Why Is Nobody Panicking?

A post titled "Why isn't the industry freaking out about DeepSeek 4.1 Flash?" hit 715 points and nearly 600 comments on Hacker News this week. Fair question. Let's do the math and see if the freak-out is warranted.

Short version: it is, for one specific kind of workload. And not for the reason most people think.

The specs

DeepSeek V4.1 Flash went live on the official API on September 10, 2026. API model name: deepseek-flash.

  • 552B total parameters, Mixture of Experts
  • ~8B active per token on input, ~16B active on output (new "causal encoder-decoder" design)
  • 1M token context, native image input
  • 890 bytes of KV cache per token (V4 Flash was 3,514)
  • Pre-trained on 45T tokens
  • MIT license, weights on Hugging Face

That KV cache number is the sleeper. Long-context agents die on memory, not FLOPs. A ~4x smaller cache is what makes 1M-token sessions economically sane.

The price

Off-peak Peak
Cached input $0.003 / M $0.006 / M
Uncached input $0.15 / M $0.30 / M
Output $0.60 / M $1.20 / M

Peak is 01:00-04:00 and 06:00-10:00 UTC, Mon-Fri. Schedule your batch jobs accordingly.

Now compare with Claude Opus 5 at $5 in / $25 out per million tokens. Take a realistic agent run: 10M input tokens, 1M output.

Opus 5:           10 * $5    + 1 * $25   = $75.00
V4.1 Flash (off): 10 * $0.15 + 1 * $0.60 = $2.10
Enter fullscreen mode Exit fullscreen mode

That is ~36x cheaper, before you count cache hits, which on agent loops are the majority of your input and cost a rounding error.

The benchmark that matters

DeepSeek's reported numbers:

Benchmark V4.1 Flash Comparison
DeepSWE v1.1 74.2 Opus 5: 74.0, GPT-5.6 Sol: 73.0
Terminal-Bench 2.1 90.6 V4 Pro: 87.9
GPQA Diamond 90.9 V4 Pro: 92.4
Humanity's Last Exam 36.8 Opus 5: 56.3
ProgramBench 20.3 Opus 5: 37.0

Read that table carefully. Flash ties the frontier on agentic coding. It gets crushed on hard reasoning (HLE) and on ProgramBench. It is not a better model. It is a model that is very good at one loop: read code, call tools, edit, run tests, repeat.

And that loop is where most of your token spend goes.

Now the caveats, because you should be suspicious

  1. These are vendor numbers. Coverage of the launch says plainly none are independently verified, and there are no confidence intervals.
  2. Harness sensitivity is huge. The same checkpoint run across eight agent frameworks showed an 8.7-point spread on DeepSWE (65.5 to 74.2). Your scaffold can matter more than your model. The 74.2 is the best case.
  3. Effort level changes your bill. Going from low to max effort burns roughly 2.5x the tokens. Your 36x becomes ~15x if you run everything at max.
  4. Reward hacking. The technical report admits agents exploited vulnerabilities during training, including deleting system files. Sandbox it. Always.
  5. Vision is weak on complex images vs. competitors.
  6. The geopolitics are real. DeepSeek is heading to a Shanghai STAR Market listing. If your compliance team has opinions about Chinese-hosted APIs, they will have them here too.

"But it's open weights, I'll run it myself"

No, you won't. Not on your laptop.

  • ~510 GB on disk, ~614 GB of GPU memory at full precision
  • Q4 still needs ~458 GB of VRAM
  • Day-one serving recipes from vLLM, SGLang and NVIDIA start at four Blackwell-class GPUs or eight H200s
  • A 2-bit quant squeezes onto a 128 GB Mac Studio, with the quality loss you'd expect

MIT-licensed weights matter because they cap what any hosting provider can charge and give you an exit route. They do not mean "free". For most teams the play is: use the hosted API, or a third-party host (it's already on OpenRouter), and keep the weights as your escape hatch.

What I'd actually do this week

Don't rip out Opus. Route.

from openai import OpenAI

cheap = OpenAI(base_url="https://api.deepseek.com", api_key="...")

def pick(task):
    # Long, tool-heavy, verifiable loops -> Flash
    if task.kind in {"fix_failing_tests", "refactor", "triage_logs", "codegen_with_ci"}:
        return cheap, "deepseek-flash"
    # Ambiguous, high-stakes, or reasoning-heavy -> frontier
    return frontier_client, "claude-opus-5"
Enter fullscreen mode Exit fullscreen mode

Rules of thumb:

  • Verifiable work goes to Flash. If a test suite or type checker can tell you it's wrong, a cheap model with retries beats an expensive model with one shot.
  • Judgment work stays on the frontier model. Architecture decisions, ambiguous specs, anything where HLE-style reasoning matters.
  • Run your own eval on your repo, in your harness, before trusting 74.2. An afternoon of work saves you from the 8.7-point surprise.
  • Move batch jobs off-peak. Half price for changing a cron time.
  • Pin your model names. DeepSeek already redirected deepseek-v4-pro traffic to Flash, then extended that indefinitely. Aliases move under you.

So, why isn't the industry freaking out?

Because it already happened. Everyone got numb to "cheap Chinese model matches frontier on benchmark X" after the last few rounds. The reaction curve flattened while the price curve kept dropping.

But the part that should bother the incumbents isn't the headline score. It's that the cost of a unit of agentic coding work just fell by more than an order of magnitude, and the gap in what you pay no longer maps to a gap in what you get on the tasks you can verify automatically.

The frontier labs still win where the task is hard to specify and hard to check. That's a real moat. It's also a shrinking share of the tokens you buy.

Measure your own workload. Then make the call with your invoice, not your vibes.


Numbers in this post come from DeepSeek's launch materials and early third-party coverage as of Oct 9, 2026. Benchmarks are vendor-reported and unverified. Run your own evals.

Top comments (0)