DEV Community

jamilxt
jamilxt

Posted on

DeepSeek 4.1 Flash Costs $0.003 per Million Tokens. Here Is the Two-Tier Router It Makes Possible.

If you run AI agents on cron jobs the way I do, you already know where the bill comes from. It is not the output. It is the re-sent context. Every turn of a coding agent re-reads the whole conversation, so a session that runs for an afternoon can push tens of millions of input tokens through an API even if the model only writes a few thousand lines back. Input pricing is the entire game, and last week the input price fell off a cliff.

A post titled Why isn't the industry freaking out about DeepSeek 4.1 Flash? has been sitting near the top of Hacker News with 829 points and over 700 comments. The author's claim is simple: DeepSeek 4.1 Flash is good enough that he stopped thinking about which model he was using, and his all-day agent sessions rarely exceed one dollar. The number underneath that claim is the real story: DeepSeek prices cached input on this model at $0.003 per million tokens off-peak, and $0.006 at peak. Flat frontier pricing is around $3 per million input tokens. That is not a discount. That is three orders of magnitude.

This piece is not a recap of that thread. It is the thing the thread implies: a two-tier model router you can actually build, with the math that justifies it worked out line by line. One caveat up front. I have not benchmarked 4.1 Flash against Opus on my own workloads, and this article does not claim I did. The cost figures below are arithmetic from published pricing. The routing design is the part you can copy.

What 4.1 Flash actually is

The model shipped on September 10, 2026. The specs that matter for cost:

  • 552B parameter MoE backbone with a 1M token context window, which is large enough to hold a long agent session without truncation.
  • Open weights under the MIT license, so the same routing math works against a self-hosted deployment or any of the dozen providers on OpenRouter, where third-party hosts currently sell input between roughly $0.014 and $0.30 per million tokens depending on provider and cache behavior.
  • Native multimodal input and aggressive context caching, which is where the pricing gets interesting.

The pricing structure, per DeepSeek's API docs and coverage from VentureBeat:

  • Cached input, off-peak: $0.003 per million tokens
  • Uncached input, off-peak: $0.15 per million tokens
  • Output, off-peak: $0.60 per million tokens
  • Peak pricing: exactly 2x each of those numbers

Two details worth pausing on. First, cache hits are 50x cheaper than cache misses on this model, at both peak and off-peak. The pricing is deliberately shaped to reward you for keeping long-lived context, which is precisely what agents do. Second, DeepSeek compressed the KV cache by a factor the blog post puts at roughly 437x versus their V1 architecture. Holding that cache in GPU memory is one of the biggest costs of serving long sessions, and it is the structural reason the price can be this low. Off-peak windows are published on DeepSeek's pricing page; batch work that can wait for those hours gets the bottom rate automatically.

The math on a real agent session

Let me make "orders of magnitude" concrete. Here is a worked example with stated assumptions so you can adjust them. It is arithmetic, not a measured bill.

The scenario. A coding agent session runs 30 turns. Context grows by about 40k tokens per turn as files, diffs, and tool output accumulate, and every turn re-sends the full history. The model writes back about 600 tokens per turn.

  • Total input tokens re-sent: 18.6 million
  • Total output tokens: 18 thousand

On flat frontier pricing at $3 per million input and $15 per million output, that session costs $56.07. This is the number that shows up on bills and gets written up as horror stories.

On 4.1 Flash at peak with a 95 percent cache hit rate, which is realistic when the conversation prefix is stable: 17.67M cached tokens at $0.006, plus 0.93M uncached at $0.30, plus output at $1.20 per million. The session costs $0.41. Run it in the off-peak window and it drops to $0.20.

The same arithmetic over a month of 1,000 such sessions: about $407 on 4.1 Flash peak versus about $56,000 on flat frontier pricing. Even if you distrust the benchmark claims entirely, the pricing gap is so wide that the model can be meaningfully worse and still win most workloads outright.

The extreme case is worth knowing because it changes what you dare to build. 50 million input tokens that fully hit the cache cost $0.15 off-peak, against $150 at flat frontier input rates. Exactly 1,000x. Tasks that were economically absurd, like re-reading an entire codebase on every CI run, re-summarizing a documentation set nightly, or letting an exploratory agent wander for hours, become rounding errors. The dgt.is author's framing matches this: with no marginal cost anxiety, you stop rationing the boring tasks.

The design: a two-tier router

Cheap capacity does not remove the frontier model. It changes its job. The pattern that falls out of the numbers:

  • Tier 1, the worker: a cheap high-cache model runs everything. Exploration, test loops, drafting, bulk refactors, the 90 percent of calls where you need competence, not brilliance.
  • Tier 2, the reviewer: the frontier model sees only the distilled result. Final review of a diff, architecture decisions, the one call per session where a subtle bug costs more than a thousand cheap calls.

The cost inversion is the point: the expensive model stops being your default and becomes your exception, and its input shrinks because it reviews outcomes instead of conversations. Here is a minimal router, the policy shape rather than production code:

from dataclasses import dataclass

@dataclass
class ModelTier:
    name: str
    model: str
    max_input_tokens: int

TIERS = [
    ModelTier("worker",   "deepseek-v4.1-flash", 1_000_000),
    ModelTier("reviewer", "claude-opus-5-5",       200_000),
]

def route(task: dict) -> ModelTier:
    # Escalate on stakes, not on prompt size
    if task.get("kind") in ("final_review", "architecture", "security"):
        return TIERS[1]
    # Escalate after repeated cheap failures, not before
    if task.get("failed_attempts", 0) >= 2:
        return TIERS[1]
    return TIERS[0]

def run_task(task: dict, llm_call):
    tier = route(task)
    result = llm_call(tier.model, task["prompt"])
    attempts = 0
    while not result.ok and attempts < 2:
        attempts += 1
        result = llm_call(tier.model, task["prompt"], retry=True)
    if not result.ok:
        fallback = route({**task, "failed_attempts": 2})
        if fallback.name != tier.name:
            result = llm_call(fallback.model, task["prompt"])
    return result
Enter fullscreen mode Exit fullscreen mode

Three design decisions in that snippet matter more than the code:

  • Route on stakes and failure count, never on prompt size. Sending a big context to the cheap model costs almost nothing on a cache hit. Sending a routine task to the expensive model costs 50 to 1,000x more than it needed to.
  • Escalate after two cheap failures. Most retries fail for the same reason the first attempt did. Burning two cents before spending fifty proves the task is actually hard.
  • Shrink the reviewer's input. Hand tier 2 the diff and the test results, not the whole session transcript. A review pass over 20k tokens is a rounding error even at frontier rates.

Where this breaks down

Honest limits, because the HN enthusiasm deserves a counterweight:

  • Quality claims are still largely self-reported. The dgt.is post is one developer's subjective experience plus vendor benchmarks. Third-party checks like the Artificial Analysis run that scored the V4 Flash line at 52 on their intelligence index are useful, but "could not tell which model I was using" is an anecdote, not a benchmark. Run your own eval on your own tasks before moving anything critical down-tier.
  • Cache-hit economics require prefix discipline. The $0.003 rate assumes your re-sent context is actually stable. Agents that rewrite system prompts mid-session, inject timestamps into messages, or shuffle tool results destroy their own hit rate and fall back to the 50x more expensive miss rate. Keep the prompt prefix byte-identical and append, never reorder.
  • Data residency is a real consideration. DeepSeek's API processes requests in China, and the same HN thread relitigates the training-data controversy around distilled Claude outputs. For personal projects and cost-sensitive batch work the tradeoff may be easy. For proprietary code under compliance obligations, it may be a hard no, or the reason to point the same router at a self-hosted copy of the MIT-licensed weights.
  • Rate limits and provider variance exist. OpenRouter's provider table shows throughput from roughly 60 to 330 tokens per second and uptimes from 94 to 100 percent across hosts. A router worth building pins a primary and fallback provider per tier instead of assuming one endpoint.

The checklist I would actually apply

The save-this part, distilled:

  • Audit your last bill by direction of tokens. If input dominates, cache-friendly pricing matters more than output price or benchmark rank. Most agent bills are 90 percent+ input.
  • Move exploratory and bulk work to the cheap tier first. Test loops, codebase Q&A, nightly summarization, UI monkey testing. These are the tasks the dgt.is author stopped rationing, and they are also the tasks where a miss costs you a retry, not a production bug.
  • Keep the frontier model as a reviewer, not a default. One escalation per session, fed a summary, is enough to catch most edge cases at two orders of magnitude less cost than doing everything there.
  • Protect your cache hit rate like a resource. Stable prompt prefixes, appended history, pinned tool schemas. A dropped hit rate is a silent 50x price increase.
  • Re-check this math quarterly. The frontier labs are responding; the same caching techniques are why Opus-class models quietly got cheaper too. A router built on a fixed price gap rots. A router built on a routing policy does not.

The deeper shift is not one model's price. It is that "which model" is becoming a routing decision, a line in a config file, instead of a subscription decision, a monthly bill. The developers who internalize that first are the ones whose agent budgets look reasonable next year while everyone else posts screenshots.


I write about developer tools, backend engineering, and the economics of running AI agents every week. Subscribe, it is free, and it means the next one shows up without you hunting for it.

Are you routing between model tiers in your own agent stack, or paying flat frontier rates for everything? If you have run DeepSeek 4.1 Flash on real work, what broke first? I am collecting failure modes before I trust it with more of my cron jobs.

Top comments (1)

Collapse
 
manojkagitha profile image
Manoj Kumar Kagitha •

Prefix discipline is the detail I'd bold. We run an LLM gateway with per-tenant metering, and cache hits decide the bill more than the model choice, so we log the hit rate per tenant to catch a quiet prompt change before it 50x's the cost.