LLM Pricing Digest: DeepSeek Quietly Drops a Free Flash Model
Weekly snapshot from LLM Price Watch — tracking Claude, GPT, Gemini, DeepSeek, and Grok pricing daily.
This was a quiet week on the pricing front. No price cuts, no increases, nothing moved across our five tracked providers. If you're budgeting for an existing integration, your numbers are the same as they were seven days ago.
The one thing worth noting: DeepSeek shipped a new model.
DeepSeek V4 Flash (0731) — Free Tier
We spotted deepseek/deepseek-v4-flash-0731 on September 18th, and it's available on the free tier via OpenRouter. That's the part that actually matters if you're evaluating it for a project.
Here's what I can tell you from the data we have: it's a "flash" variant, which in the current model-naming landscape generally signals a smaller, faster model optimized for lower latency and cost rather than maximum capability. The 0731 date stamp suggests it was trained on or finalized around July 31st of this year. Beyond that, I'm not going to speculate about benchmark numbers or capability claims — we track pricing, not vibes.
What I can say practically: if you've been experimenting with DeepSeek's models and want to test this one without burning through API credits, the free tier is a reasonable way to do that. Run it against your actual use case — summarization, classification, code completion, whatever you're building — and see if the quality holds up before committing.
Why Free-Tier Models Are Worth Watching
Free tiers on OpenRouter tend to come with rate limits and sometimes different routing than paid endpoints, so they're not always a direct proxy for production performance. But they're genuinely useful for:
- Prototyping before you decide which model to pay for
- Regression testing in CI pipelines where cost adds up fast
- Side-by-side evals without having to pre-authorize spend
DeepSeek has been aggressive about making their models accessible, and that's continued here. If you're already using DeepSeek V3 or their R1 series for anything, this flash variant is at least worth a quick benchmark on your workload.
The Bigger Picture This Week
No price changes across Claude, GPT, Gemini, or Grok either. The market has been relatively stable for a few weeks now after a busy stretch of cuts earlier this year. That's not a complaint — stable pricing makes it easier to plan infrastructure costs — but it does mean this digest is shorter than usual.
If you're in the middle of a model selection decision right now, the calculus hasn't shifted this week. The relative cost differences between providers are the same as last week. DeepSeek models remain meaningfully cheaper than OpenAI and Anthropic at the high end, Gemini Flash variants are still the go-to for high-volume, cost-sensitive workloads, and Claude holds its ground for tasks where output quality and instruction-following matter most.
What I'm Watching
A few things I'll be tracking over the next couple of weeks:
- Whether DeepSeek V4 Flash gets more detailed specs or a paid pricing tier announced
- Any movement on Grok pricing, which has been static for a while
- Whether Gemini or OpenAI respond to each other's pricing — historically those two tend to leapfrog
That's it for this week. Short digest, but accurate is more useful than padded.
For live, daily-updated pricing across all five providers, check out llmpricewatch.com. We track input/output token costs and flag changes as soon as they hit.
Top comments (0)