<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Anders Rasmussen</title>
    <description>The latest articles on DEV Community by Anders Rasmussen (@adrasmussen).</description>
    <link>https://dev.to/adrasmussen</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4067961%2F95c580d5-2d5d-457c-982b-072a37ed7c4c.png</url>
      <title>DEV Community: Anders Rasmussen</title>
      <link>https://dev.to/adrasmussen</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/adrasmussen"/>
    <language>en</language>
    <item>
      <title>DeepSeek Cuts Prices in Half Mid-Week: This Week in LLM Pricing</title>
      <dc:creator>Anders Rasmussen</dc:creator>
      <pubDate>Sun, 11 Oct 2026 10:01:44 +0000</pubDate>
      <link>https://dev.to/adrasmussen/deepseek-cuts-prices-in-half-mid-week-this-week-in-llm-pricing-okd</link>
      <guid>https://dev.to/adrasmussen/deepseek-cuts-prices-in-half-mid-week-this-week-in-llm-pricing-okd</guid>
      <description>&lt;p&gt;The big story this week is straightforward: DeepSeek halved its prices on two models mid-week. If you're routing any significant volume through either of them, this is worth paying attention to.&lt;/p&gt;

&lt;h2&gt;
  
  
  The DeepSeek Price Cuts
&lt;/h2&gt;

&lt;p&gt;Our tracker caught two snapshots this week for &lt;code&gt;deepseek-v4-1-flash&lt;/code&gt; and &lt;code&gt;deepseek-v4-pro&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;deepseek-v4-1-flash&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Oct 5: $0.30 input / $1.20 output per million tokens&lt;/li&gt;
&lt;li&gt;Oct 10: $0.15 input / $0.60 output per million tokens&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;deepseek-v4-pro&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Oct 5: $1.32 input / $3.96 output per million tokens&lt;/li&gt;
&lt;li&gt;Oct 10: $0.66 input / $1.98 output per million tokens&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That's a clean 50% reduction across the board for both models, applied between October 5th and October 10th. Not a rounding artifact — both models dropped by exactly half on both input and output.&lt;/p&gt;

&lt;p&gt;What does this mean practically? If you've been holding off on &lt;code&gt;deepseek-v4-pro&lt;/code&gt; because the cost felt hard to justify against something like GPT-4o or Gemini 1.5 Pro, the math changes now. At $0.66/$1.98, it sits in a much more competitive band for mid-tier inference tasks. And &lt;code&gt;deepseek-v4-1-flash&lt;/code&gt; at $0.15 input is genuinely cheap for high-volume, lower-stakes work — summarization, classification, first-pass drafts.&lt;/p&gt;

&lt;p&gt;DeepSeek has form here. They've used aggressive pricing as a market share play before, and this follows the same pattern. Whether the prices stay here or drop further is hard to predict, but right now they represent real value if the model quality holds up for your use case.&lt;/p&gt;

&lt;h2&gt;
  
  
  New Models Spotted
&lt;/h2&gt;

&lt;p&gt;Three new model IDs appeared in our tracker this week.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;anthropic/claude-haiku-5.5&lt;/strong&gt; and &lt;strong&gt;anthropic/claude-haiku-5.5:batch&lt;/strong&gt; both showed up on October 8th. The batch variant follows Anthropic's established pattern of offering asynchronous batch processing at reduced rates. Haiku has always been Anthropic's speed-and-cost tier, and a 5.5 version suggests incremental improvements over whatever 5.x came before. No pricing anomalies to flag on these yet — we'll track them going forward.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;google/gemini-nano-banana-2.1&lt;/strong&gt; appeared on October 7th from Google. The "nano" and "banana" naming suggests this is either an experimental model or a specialized lightweight variant — Google has used codename-style identifiers before for models not yet in full public release. Worth watching but nothing concrete to say about it yet beyond the fact that it exists in the API.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Practical Takeaway
&lt;/h2&gt;

&lt;p&gt;If you have a cost-sensitive workload and you haven't benchmarked DeepSeek V4 Flash recently, now's a reasonable time to do it. At $0.15 per million input tokens, the failure cost of running a quick eval is minimal. For higher-capability needs, V4 Pro at $0.66 input puts it in a range where it's worth comparing directly against your current go-to.&lt;/p&gt;

&lt;p&gt;The new Anthropic and Google models are worth keeping an eye on as they mature — particularly Claude Haiku 5.5 if you're already in the Anthropic ecosystem and care about the cost/performance tradeoff at the smaller end.&lt;/p&gt;

&lt;p&gt;As always, prices on third-party routers like OpenRouter can lag official announcements or differ slightly from direct API pricing, so verify before committing to anything at scale.&lt;/p&gt;




&lt;p&gt;I track these changes weekly over at &lt;a href="https://llmpricewatch.com/" rel="noopener noreferrer"&gt;LLM Price Watch&lt;/a&gt;, which checks live pricing across Claude, GPT, Gemini, DeepSeek, and Grok daily. Worth bookmarking if you're making model selection decisions based on cost.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>api</category>
      <category>webdev</category>
    </item>
    <item>
      <title>LLM Pricing Digest: Grok-4 Gets a Price Drop, and a Wave of New Models From Qwen, Xiaomi, and Z-AI</title>
      <dc:creator>Anders Rasmussen</dc:creator>
      <pubDate>Sun, 04 Oct 2026 10:02:21 +0000</pubDate>
      <link>https://dev.to/adrasmussen/llm-pricing-digest-grok-4-gets-a-price-drop-and-a-wave-of-new-models-from-qwen-xiaomi-and-z-ai-44mp</link>
      <guid>https://dev.to/adrasmussen/llm-pricing-digest-grok-4-gets-a-price-drop-and-a-wave-of-new-models-from-qwen-xiaomi-and-z-ai-44mp</guid>
      <description>&lt;h1&gt;
  
  
  LLM Pricing Digest — Week of September 29, 2026
&lt;/h1&gt;

&lt;p&gt;&lt;em&gt;Weekly data from &lt;a href="https://llmpricewatch.com/" rel="noopener noreferrer"&gt;LLM Price Watch&lt;/a&gt;, which tracks live pricing for Claude, GPT, Gemini, DeepSeek, Grok, and more.&lt;/em&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  The One Real Price Change: Grok-4
&lt;/h2&gt;

&lt;p&gt;This week had one confirmed price movement: &lt;strong&gt;grok-4&lt;/strong&gt; dropped to &lt;strong&gt;$2/million input tokens and $6/million output tokens&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;To put that in context, Grok-4 launched as a frontier-tier model, and pricing at that level used to reliably mean $10–$15+ per million on input. At $2 in and $6 out, it's now sitting in territory that makes it genuinely worth comparing against mid-tier options for production workloads — not just for experimentation.&lt;/p&gt;

&lt;p&gt;Practically speaking: if you've been routing complex reasoning tasks to GPT-4o or Claude Sonnet and paying ~$5–$15/M input depending on your mix, Grok-4 at $2 input is worth a benchmark run. Output-heavy workloads (long generations, document drafting) will feel the $6/M output price more, but that's still competitive for a model at this capability tier.&lt;/p&gt;

&lt;p&gt;No price changes detected for Claude, GPT, or Gemini models this week.&lt;/p&gt;




&lt;h2&gt;
  
  
  New Models: Qwen, Xiaomi, and Z-AI All Showed Up at Once
&lt;/h2&gt;

&lt;p&gt;Fifteen new models appeared in our tracker on September 30th. They all landed the same day, which usually means a coordinated release push rather than a gradual rollout. Here's the breakdown:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Qwen (5 models)&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;qwen/qwen3.8-max-prime&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;qwen/qwen3.8-omni-flash&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;qwen/qwen3.8-max-0902&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;qwen/qwen3.8-flash&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;qwen/qwen3.8-27b&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;qwen/qwen3.8-27b:free&lt;/code&gt; ← free tier variant&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Qwen continues to ship variants aggressively. The naming here follows their established pattern: max = higher capability, flash = faster/cheaper, omni = multimodal. The 27B free tier is notable — free access to a 27B model is useful for prototyping or low-volume use cases where cost is the primary constraint.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Xiaomi (3 models)&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;xiaomi/mimo-v2.6-pro-ultraspeed&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;xiaomi/mimo-v2.6-flash&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;xiaomi/mimo-v2.6-pro&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Xiaomi's MiMo line is newer to the router ecosystem. Three variants in one drop — ultraspeed, flash, and pro — suggests they're covering the same speed/quality tradeoff spectrum everyone else is. No pricing data to share yet beyond what's in the tracker.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Z-AI / GLM-5.3 (7 models)&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;z-ai/glm-5.3-prime&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;z-ai/glm-5.3-flashx&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;z-ai/glm-5.3-flash&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;z-ai/glm-5.3-flash:batch&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;z-ai/glm-5.3&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;z-ai/glm-5.3:batch&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Z-AI's GLM series (from Zhipu AI) is the most prolific of the three this week, with seven variants including batch-mode endpoints. Batch variants matter if you're running offline pipelines — they're typically priced lower in exchange for slower turnaround. Worth checking if you have async workloads that don't need real-time responses.&lt;/p&gt;




&lt;h2&gt;
  
  
  What This Week Means Practically
&lt;/h2&gt;

&lt;p&gt;If you're actively managing API costs right now:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Re-evaluate Grok-4&lt;/strong&gt; if you dismissed it earlier on price. $2 input is a different conversation than where it started.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The Qwen 27B free tier&lt;/strong&gt; is a legitimate option for dev/test environments or low-stakes internal tools.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GLM-5.3 batch endpoints&lt;/strong&gt; are worth a look if you're doing bulk processing — batch pricing tends to undercut synchronous endpoints meaningfully.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Xiaomi MiMo&lt;/strong&gt; is one to watch but probably wait for more community benchmarks before routing production traffic there.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The broader trend continues to hold: the number of capable models available through routing APIs keeps growing, and the middle of the market (the $1–$6/M input range) is getting crowded. That's good for developers with flexibility in model choice.&lt;/p&gt;




&lt;p&gt;Full pricing data, historical charts, and alerts are at &lt;strong&gt;&lt;a href="https://llmpricewatch.com/" rel="noopener noreferrer"&gt;llmpricewatch.com&lt;/a&gt;&lt;/strong&gt;. The tracker checks OpenRouter daily, so any new price moves on these models will show up there first.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>api</category>
      <category>webdev</category>
    </item>
    <item>
      <title>DeepSeek V3 Gets Cheaper Again, Plus Eight GPT-6 Variants and More Hit the API</title>
      <dc:creator>Anders Rasmussen</dc:creator>
      <pubDate>Sun, 27 Sep 2026 10:02:06 +0000</pubDate>
      <link>https://dev.to/adrasmussen/deepseek-v3-gets-cheaper-again-plus-eight-gpt-6-variants-and-more-hit-the-api-1j4h</link>
      <guid>https://dev.to/adrasmussen/deepseek-v3-gets-cheaper-again-plus-eight-gpt-6-variants-and-more-hit-the-api-1j4h</guid>
      <description>&lt;p&gt;Every week I pull the actual price data from our tracker at &lt;a href="https://llmpricewatch.com/" rel="noopener noreferrer"&gt;LLM Price Watch&lt;/a&gt; and report what changed. This week there's a real price drop worth knowing about, plus a flood of new models that showed up on OpenRouter overnight.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Price Change: DeepSeek V3
&lt;/h2&gt;

&lt;p&gt;DeepSeek V3 dropped again. As of September 24th, it's sitting at &lt;strong&gt;$0.27 per million input tokens and $0.41 per million output tokens&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;To put that in practical terms: if you're running a mid-sized RAG pipeline pushing a few hundred million tokens a month, you're looking at real money saved versus most of the alternatives. DeepSeek has been on a consistent downward trajectory on pricing, and V3 in particular has become a go-to recommendation for teams that need solid general-purpose performance without paying frontier model prices.&lt;/p&gt;

&lt;p&gt;For comparison, if you were previously paying $0.38/M input (roughly where V3 was a few months back), the new rate is a meaningful cut. Output-heavy workloads — think long-form generation, detailed summarization, agentic loops — benefit most here since output tokens are typically where costs pile up.&lt;/p&gt;

&lt;p&gt;If you haven't benchmarked DeepSeek V3 for your use case recently, it's worth a look now.&lt;/p&gt;

&lt;h2&gt;
  
  
  New Models: OpenAI's GPT-6 Family
&lt;/h2&gt;

&lt;p&gt;Eight new OpenAI model IDs appeared on OpenRouter on September 23rd:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;gpt-6-luna-pro&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;gpt-6-luna-pro:batch&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;gpt-6-luna&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;gpt-6-luna:batch&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;gpt-6-sol-pro&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;gpt-6-sol-pro:batch&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;gpt-6-sol&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;gpt-6-sol:batch&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Our tracker spotted all eight at the same timestamp, which suggests a coordinated rollout rather than a staged one. The naming pattern — Luna and Sol as sub-variants, with Pro tiers and Batch variants for each — looks like OpenAI is following the same tiering playbook they used with the GPT-4o family.&lt;/p&gt;

&lt;p&gt;I don't have pricing data for these yet (they appeared as new model IDs without confirmed per-token rates in our snapshot), so I'd treat any numbers you see floating around as unverified until they stabilize. Worth watching.&lt;/p&gt;

&lt;h2&gt;
  
  
  New Models: Anthropic Claude Opus 5.5
&lt;/h2&gt;

&lt;p&gt;Two more new IDs from Anthropic:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;claude-opus-5.5&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;claude-opus-5.5:batch&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Opus has historically been Anthropic's most capable and most expensive tier, so a point-release here is interesting. Again, no confirmed pricing in our snapshot yet — our tracker caught the model IDs before rate data was available.&lt;/p&gt;

&lt;h2&gt;
  
  
  New Models: DeepSeek V4.1 Flash Batch and Grok 4.7
&lt;/h2&gt;

&lt;p&gt;DeepSeek also snuck in a batch variant of what appears to be a V4.1 Flash model (&lt;code&gt;deepseek-v4.1-flash:batch&lt;/code&gt;), first seen September 23rd. No pricing confirmed yet, but the "flash" naming convention from DeepSeek has previously meant a smaller, faster, cheaper variant — so this could be interesting for high-volume, latency-tolerant workloads.&lt;/p&gt;

&lt;p&gt;And on the xAI side, &lt;code&gt;grok-4.7&lt;/code&gt; appeared on September 22nd — a day ahead of the others. No pricing data in our snapshot for this one either.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to Actually Do With This
&lt;/h2&gt;

&lt;p&gt;Here's the practical summary:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;If you're using DeepSeek V3&lt;/strong&gt;, check your current rate against $0.27/$0.41 — you may already be on the new pricing depending on your provider, but it's worth verifying.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The GPT-6 and Claude Opus 5.5 launches&lt;/strong&gt; are real, but hold off on budgeting around them until pricing stabilizes and benchmarks are out.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;DeepSeek V4.1 Flash batch and Grok 4.7&lt;/strong&gt; are names to file away. Watch for pricing.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;I'll update as confirmed rates come in. The tracker runs daily checks, so if anything shifts mid-week you can catch it live.&lt;/p&gt;




&lt;p&gt;Full live pricing for Claude, GPT, Gemini, DeepSeek, and Grok is at &lt;a href="https://llmpricewatch.com/" rel="noopener noreferrer"&gt;llmpricewatch.com&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>api</category>
      <category>webdev</category>
    </item>
    <item>
      <title>LLM Pricing Digest: DeepSeek Quietly Drops a Free Flash Model</title>
      <dc:creator>Anders Rasmussen</dc:creator>
      <pubDate>Mon, 21 Sep 2026 08:41:18 +0000</pubDate>
      <link>https://dev.to/adrasmussen/llm-pricing-digest-deepseek-quietly-drops-a-free-flash-model-4hk8</link>
      <guid>https://dev.to/adrasmussen/llm-pricing-digest-deepseek-quietly-drops-a-free-flash-model-4hk8</guid>
      <description>&lt;h1&gt;
  
  
  LLM Pricing Digest: DeepSeek Quietly Drops a Free Flash Model
&lt;/h1&gt;

&lt;p&gt;&lt;em&gt;Weekly snapshot from &lt;a href="https://llmpricewatch.com/" rel="noopener noreferrer"&gt;LLM Price Watch&lt;/a&gt; — tracking Claude, GPT, Gemini, DeepSeek, and Grok pricing daily.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;This was a quiet week on the pricing front. No price cuts, no increases, nothing moved across our five tracked providers. If you're budgeting for an existing integration, your numbers are the same as they were seven days ago.&lt;/p&gt;

&lt;p&gt;The one thing worth noting: DeepSeek shipped a new model.&lt;/p&gt;

&lt;h2&gt;
  
  
  DeepSeek V4 Flash (0731) — Free Tier
&lt;/h2&gt;

&lt;p&gt;We spotted &lt;code&gt;deepseek/deepseek-v4-flash-0731&lt;/code&gt; on September 18th, and it's available on the free tier via OpenRouter. That's the part that actually matters if you're evaluating it for a project.&lt;/p&gt;

&lt;p&gt;Here's what I can tell you from the data we have: it's a "flash" variant, which in the current model-naming landscape generally signals a smaller, faster model optimized for lower latency and cost rather than maximum capability. The &lt;code&gt;0731&lt;/code&gt; date stamp suggests it was trained on or finalized around July 31st of this year. Beyond that, I'm not going to speculate about benchmark numbers or capability claims — we track pricing, not vibes.&lt;/p&gt;

&lt;p&gt;What I &lt;em&gt;can&lt;/em&gt; say practically: if you've been experimenting with DeepSeek's models and want to test this one without burning through API credits, the free tier is a reasonable way to do that. Run it against your actual use case — summarization, classification, code completion, whatever you're building — and see if the quality holds up before committing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Free-Tier Models Are Worth Watching
&lt;/h2&gt;

&lt;p&gt;Free tiers on OpenRouter tend to come with rate limits and sometimes different routing than paid endpoints, so they're not always a direct proxy for production performance. But they're genuinely useful for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Prototyping&lt;/strong&gt; before you decide which model to pay for&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Regression testing&lt;/strong&gt; in CI pipelines where cost adds up fast&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Side-by-side evals&lt;/strong&gt; without having to pre-authorize spend&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;DeepSeek has been aggressive about making their models accessible, and that's continued here. If you're already using DeepSeek V3 or their R1 series for anything, this flash variant is at least worth a quick benchmark on your workload.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Bigger Picture This Week
&lt;/h2&gt;

&lt;p&gt;No price changes across Claude, GPT, Gemini, or Grok either. The market has been relatively stable for a few weeks now after a busy stretch of cuts earlier this year. That's not a complaint — stable pricing makes it easier to plan infrastructure costs — but it does mean this digest is shorter than usual.&lt;/p&gt;

&lt;p&gt;If you're in the middle of a model selection decision right now, the calculus hasn't shifted this week. The relative cost differences between providers are the same as last week. DeepSeek models remain meaningfully cheaper than OpenAI and Anthropic at the high end, Gemini Flash variants are still the go-to for high-volume, cost-sensitive workloads, and Claude holds its ground for tasks where output quality and instruction-following matter most.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'm Watching
&lt;/h2&gt;

&lt;p&gt;A few things I'll be tracking over the next couple of weeks:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Whether DeepSeek V4 Flash gets more detailed specs or a paid pricing tier announced&lt;/li&gt;
&lt;li&gt;Any movement on Grok pricing, which has been static for a while&lt;/li&gt;
&lt;li&gt;Whether Gemini or OpenAI respond to each other's pricing — historically those two tend to leapfrog&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That's it for this week. Short digest, but accurate is more useful than padded.&lt;/p&gt;




&lt;p&gt;For live, daily-updated pricing across all five providers, check out &lt;a href="https://llmpricewatch.com/" rel="noopener noreferrer"&gt;llmpricewatch.com&lt;/a&gt;. We track input/output token costs and flag changes as soon as they hit.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>api</category>
      <category>webdev</category>
    </item>
    <item>
      <title>LLM Pricing Digest: DeepSeek Quietly Drops Two New Flash Models</title>
      <dc:creator>Anders Rasmussen</dc:creator>
      <pubDate>Sun, 13 Sep 2026 10:02:23 +0000</pubDate>
      <link>https://dev.to/adrasmussen/llm-pricing-digest-deepseek-quietly-drops-two-new-flash-models-191g</link>
      <guid>https://dev.to/adrasmussen/llm-pricing-digest-deepseek-quietly-drops-two-new-flash-models-191g</guid>
      <description>&lt;h1&gt;
  
  
  LLM Pricing Digest: DeepSeek Quietly Drops Two New Flash Models
&lt;/h1&gt;

&lt;p&gt;&lt;em&gt;Weekly roundup from &lt;a href="https://llmpricewatch.com/" rel="noopener noreferrer"&gt;LLM Price Watch&lt;/a&gt; — tracking Claude, GPT, Gemini, DeepSeek, and Grok pricing daily.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;This week's data is straightforward: no price changes across any of our five tracked providers. Zero. If you've got a budget locked in for GPT-4o, Claude Sonnet, or Gemini Flash, nothing moved on you.&lt;/p&gt;

&lt;p&gt;What did happen is DeepSeek showed up with two new models on OpenRouter, and they're worth knowing about.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's New from DeepSeek
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;deepseek/deepseek-v4.1-flash&lt;/code&gt;&lt;/strong&gt; — spotted September 10th.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;deepseek/deepseek-v4-flash-vision-exp:batch&lt;/code&gt;&lt;/strong&gt; — spotted September 9th, a day earlier.&lt;/p&gt;

&lt;p&gt;Let's talk about what these names actually tell us.&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;v4.1-flash&lt;/code&gt; model follows DeepSeek's now-familiar naming pattern: a Flash tier is their lower-latency, lower-cost option, comparable in positioning to Gemini Flash or GPT-4o mini. If previous DeepSeek Flash releases are any guide, this is aimed squarely at high-volume, cost-sensitive workloads — think classification, summarization, structured extraction, anything where you're firing off thousands of calls and cost-per-token matters more than raw capability.&lt;/p&gt;

&lt;p&gt;The second model is more interesting: &lt;code&gt;deepseek-v4-flash-vision-exp&lt;/code&gt; with a &lt;code&gt;:batch&lt;/code&gt; suffix. Two things stand out here. First, the &lt;code&gt;vision&lt;/code&gt; tag means multimodal input — image understanding on top of text. Second, the &lt;code&gt;exp&lt;/code&gt; label means experimental, so treat it accordingly: don't build a production dependency on it yet, but it's worth testing. The &lt;code&gt;:batch&lt;/code&gt; variant on OpenRouter typically means asynchronous batch processing at a reduced rate — good for offline pipelines where you don't need a real-time response.&lt;/p&gt;

&lt;h2&gt;
  
  
  What This Means Practically
&lt;/h2&gt;

&lt;p&gt;If you're currently using DeepSeek for text-only tasks and you've been waiting for a vision-capable option at Flash pricing, this is the moment to run some tests. Vision at Flash cost is a genuinely useful combination — it opens up document parsing, screenshot analysis, and image classification use cases that previously required bumping up to a pricier tier or switching providers entirely.&lt;/p&gt;

&lt;p&gt;The batch variant specifically is worth a look if you have any async workloads. Batch endpoints typically offer meaningful discounts over synchronous calls, and if your pipeline can tolerate latency (nightly report generation, bulk document review, etc.), you can often cut costs significantly just by routing to the batch endpoint.&lt;/p&gt;

&lt;p&gt;That said, both models are new and one is explicitly experimental. Before committing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Run your own evals on representative tasks from your actual workload&lt;/li&gt;
&lt;li&gt;Check the current pricing on OpenRouter directly — new models sometimes launch at promotional rates that adjust within weeks&lt;/li&gt;
&lt;li&gt;Keep an eye on the &lt;code&gt;exp&lt;/code&gt; label; experimental models can change, be deprecated, or graduate to stable naming without much notice&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The Quiet Week Is Actually Fine
&lt;/h2&gt;

&lt;p&gt;No price changes across Claude, GPT, Gemini, and Grok this week. That's a stable environment for planning. If you've been putting off a cost analysis because prices keep shifting, this is a decent window to benchmark and lock in assumptions.&lt;/p&gt;

&lt;p&gt;DeepSeek continues to move fast on model releases. Two models in two days suggests an active release cadence heading into Q4. Whether that means more aggressive pricing pressure on the other providers is something we'll be watching.&lt;/p&gt;




&lt;p&gt;I track price changes and new model releases daily across all five major providers. Full live data at &lt;strong&gt;&lt;a href="https://llmpricewatch.com/" rel="noopener noreferrer"&gt;llmpricewatch.com&lt;/a&gt;&lt;/strong&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>api</category>
      <category>webdev</category>
    </item>
    <item>
      <title>GPT-6 Astra, Claude Fable, Gemini 3.8: A Busy Week for New LLM Releases</title>
      <dc:creator>Anders Rasmussen</dc:creator>
      <pubDate>Tue, 08 Sep 2026 10:37:56 +0000</pubDate>
      <link>https://dev.to/adrasmussen/gpt-6-astra-claude-fable-gemini-38-a-busy-week-for-new-llm-releases-12d9</link>
      <guid>https://dev.to/adrasmussen/gpt-6-astra-claude-fable-gemini-38-a-busy-week-for-new-llm-releases-12d9</guid>
      <description>&lt;h1&gt;
  
  
  GPT-6 Astra, Claude Fable, Gemini 3.8: A Busy Week for New LLM Releases
&lt;/h1&gt;

&lt;p&gt;&lt;em&gt;Weekly LLM pricing digest — week of September 5, 2026&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;No price changes to report this week across the five providers we track (OpenAI, Anthropic, Google, DeepSeek, and xAI). But the model catalog grew noticeably in just a few days, so let's walk through what actually showed up and what it might mean if you're picking a model right now.&lt;/p&gt;




&lt;h2&gt;
  
  
  OpenAI: GPT-6 Astra lands with a Pro tier and batch variants
&lt;/h2&gt;

&lt;p&gt;The biggest news this week is that OpenAI dropped two new models on September 5th: &lt;strong&gt;gpt-6-astra&lt;/strong&gt; and &lt;strong&gt;gpt-6-astra-pro&lt;/strong&gt;, each with a corresponding &lt;code&gt;:batch&lt;/code&gt; variant.&lt;/p&gt;

&lt;p&gt;The naming convention here is familiar — the &lt;code&gt;:batch&lt;/code&gt; versions are typically lower-cost, asynchronous endpoints suited for workloads where you don't need a real-time response. If you're doing large-scale document processing, evals, or anything that can tolerate a delay, batch is almost always the right call financially.&lt;/p&gt;

&lt;p&gt;The split between a base &lt;code&gt;astra&lt;/code&gt; and an &lt;code&gt;astra-pro&lt;/code&gt; tier suggests OpenAI is continuing the pattern of offering a more capable (and presumably more expensive) variant alongside a leaner everyday option. Worth checking current pricing before assuming which one fits your use case — the gap between base and pro tiers has varied a lot across model generations.&lt;/p&gt;




&lt;h2&gt;
  
  
  Anthropic: Claude Fable 5.1
&lt;/h2&gt;

&lt;p&gt;Anthropic's new entry, &lt;strong&gt;claude-fable-5.1&lt;/strong&gt;, appeared on September 2nd, again with a batch variant.&lt;/p&gt;

&lt;p&gt;The "Fable" name is new in Anthropic's lineup. Whether this sits above, below, or alongside the Sonnet/Opus/Haiku naming structure isn't something our tracker can tell you — we detect models as they appear on OpenRouter, not their internal capability tiers. What we do know is it showed up with batch support from day one, which suggests Anthropic is treating batch as a first-class offering rather than an afterthought.&lt;/p&gt;

&lt;p&gt;If you're currently using Claude for any high-volume async workload, it's worth benchmarking fable-5.1 against whatever you're running now.&lt;/p&gt;




&lt;h2&gt;
  
  
  Google: Gemini 3.8 Flash
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;gemini-3.8-flash&lt;/strong&gt; arrived on September 3rd, with batch support included.&lt;/p&gt;

&lt;p&gt;The Flash line has consistently been Google's price-performance workhorse — fast, cheap, and good enough for a wide range of tasks. 3.8 continuing that series is a reasonable assumption, though again, actual pricing on our tracker is what should drive your decision, not the name alone.&lt;/p&gt;

&lt;p&gt;If you're already using Gemini 2.x Flash for something latency-sensitive or cost-sensitive, this is probably worth a quick test.&lt;/p&gt;




&lt;h2&gt;
  
  
  xAI: Grok 4.3 Batch
&lt;/h2&gt;

&lt;p&gt;xAI added &lt;strong&gt;grok-4.3:batch&lt;/strong&gt; on September 4th. This appears to be a batch variant of an existing model rather than a completely new release, which fits the pattern of providers progressively rolling out async options for their model lineup.&lt;/p&gt;




&lt;h2&gt;
  
  
  What this week means practically
&lt;/h2&gt;

&lt;p&gt;Five new model IDs in four days across four different providers, with zero price changes on existing models. A few things stand out:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Batch variants are becoming standard.&lt;/strong&gt; Every new model this week came with a batch option. If you're not already routing eligible workloads to batch endpoints, you're likely leaving money on the table.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The catalog is expanding faster than pricing is changing.&lt;/strong&gt; That means the cost landscape for existing models is stable right now, but the choice of &lt;em&gt;which&lt;/em&gt; model to use is getting more complex. It's worth revisiting your model selection if you haven't done so recently.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;None of these models have a pricing history yet.&lt;/strong&gt; First-week pricing on a new model sometimes gets revised. It's not unusual to see adjustments in the first few weeks after launch.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;We'll be tracking all of these as they settle in. If prices shift on any of the new models — or on anything else across our five providers — it'll show up in next week's digest.&lt;/p&gt;




&lt;p&gt;For live pricing on all of these models and more, check &lt;a href="https://llmpricewatch.com/" rel="noopener noreferrer"&gt;LLM Price Watch&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;If you want the real story of how this tracker (and its sister site, StackIndex AI) got built — domain traps, design mistakes, the works — it's now a course: &lt;a href="https://whop.com/checkout/plan_SvLh118mBrURJ" rel="noopener noreferrer"&gt;StackIndex Academy&lt;/a&gt;. First 3 chapters are free, no signup.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>api</category>
      <category>webdev</category>
    </item>
    <item>
      <title>GPT-6 Astra, Claude Fable, Gemini 3.8: A Busy Week for New LLM Releases (No Price Changes Though)</title>
      <dc:creator>Anders Rasmussen</dc:creator>
      <pubDate>Sun, 06 Sep 2026 10:01:47 +0000</pubDate>
      <link>https://dev.to/adrasmussen/gpt-6-astra-claude-fable-gemini-38-a-busy-week-for-new-llm-releases-no-price-changes-though-493h</link>
      <guid>https://dev.to/adrasmussen/gpt-6-astra-claude-fable-gemini-38-a-busy-week-for-new-llm-releases-no-price-changes-though-493h</guid>
      <description>&lt;h1&gt;
  
  
  GPT-6 Astra, Claude Fable, Gemini 3.8: A Busy Week for New LLM Releases
&lt;/h1&gt;

&lt;p&gt;&lt;em&gt;Weekly digest from &lt;a href="https://llmpricewatch.com/" rel="noopener noreferrer"&gt;LLM Price Watch&lt;/a&gt; — tracking live pricing across OpenAI, Anthropic, Google, DeepSeek, and xAI.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;No price changes this week across any of the five providers we track. Zero. But the model catalog? That got noticeably busier. Between September 2nd and 5th, we spotted nine new model IDs across four providers. Here's what showed up and what's worth paying attention to.&lt;/p&gt;

&lt;h2&gt;
  
  
  OpenAI: GPT-6 Astra (and a Pro tier)
&lt;/h2&gt;

&lt;p&gt;The biggest names to land on our radar are &lt;code&gt;openai/gpt-6-astra&lt;/code&gt; and &lt;code&gt;openai/gpt-6-astra-pro&lt;/code&gt;, both spotted on September 5th. Both also have batch variants (&lt;code&gt;gpt-6-astra:batch&lt;/code&gt; and &lt;code&gt;gpt-6-astra-pro:batch&lt;/code&gt;).&lt;/p&gt;

&lt;p&gt;The GPT-6 naming is a notable jump — we've been sitting with GPT-4 variants for a long time in practical use. The "Astra" label has been associated with Google's multimodal work in the past, so it's an interesting name choice from OpenAI. The Pro tier suggests a familiar tiered pricing structure is coming, similar to what we've seen with other model families.&lt;/p&gt;

&lt;p&gt;Practically speaking: if you're building something today, I'd hold off on committing to these until pricing is confirmed and benchmarks are out. Batch variants being available at launch is a good sign — batch processing has been one of the more reliable ways to cut costs on longer, async workloads.&lt;/p&gt;

&lt;h2&gt;
  
  
  xAI: Grok 4.3 Batch
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;x-ai/grok-4.3:batch&lt;/code&gt; appeared on September 4th. This is a batch-only variant, no standard inference version detected yet. xAI has been iterating fairly quickly on the Grok 4 line, and adding batch support makes sense if they're targeting workloads where latency isn't critical — think evals, document processing, data pipelines.&lt;/p&gt;

&lt;h2&gt;
  
  
  Google: Gemini 3.8 Flash
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;google/gemini-3.8-flash&lt;/code&gt; and its batch counterpart showed up on September 3rd. The Flash line from Google has generally been their cost-efficient, faster-response tier — good for high-volume use cases where you don't need the full Gemini Pro/Ultra capability. Gemini 3.8 continuing that Flash tradition makes sense. If past Flash pricing holds, this could be one of the cheaper options in the new generation lineup once pricing is confirmed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Anthropic: Claude Fable 5.1
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;anthropic/claude-fable-5.1&lt;/code&gt; and &lt;code&gt;anthropic/claude-fable-5.1:batch&lt;/code&gt; appeared on September 2nd, making it the earliest arrival in this week's batch.&lt;/p&gt;

&lt;p&gt;The "Fable" name is new for Anthropic — we haven't tracked a Fable line before. It's not clear yet where this sits relative to Sonnet, Haiku, or Opus in terms of capability or price point. The name might suggest a model tuned for narrative or creative tasks, but that's speculation. Worth watching for the official positioning.&lt;/p&gt;

&lt;h2&gt;
  
  
  What This Week Actually Means
&lt;/h2&gt;

&lt;p&gt;With no price changes, your current cost modeling stays the same. Nothing got cheaper or more expensive in the existing lineup.&lt;/p&gt;

&lt;p&gt;The main thing to flag is that four providers dropped new model IDs in a four-day window. That's a lot of surface area appearing at once. A few practical notes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Don't route production traffic to these yet&lt;/strong&gt; unless you've tested them. New model IDs appearing in the API doesn't always mean they're fully stable or that pricing is final.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Batch variants at launch&lt;/strong&gt; is genuinely useful — it signals the providers are thinking about cost-sensitive workloads from day one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The GPT-6 naming&lt;/strong&gt; is the most significant signal here. If OpenAI is moving to a 6.x generation label, expect Anthropic and Google to respond with their own positioning in the coming weeks.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I'll be watching for actual pricing data on all of these. Once numbers are confirmed, I'll do a proper comparison.&lt;/p&gt;




&lt;p&gt;For live pricing across all tracked models and providers, check &lt;a href="https://llmpricewatch.com/" rel="noopener noreferrer"&gt;llmpricewatch.com&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>api</category>
      <category>webdev</category>
    </item>
    <item>
      <title>GPT-6 Astra, Claude Fable, Gemini 3.8: A Busy Week for New LLM Releases</title>
      <dc:creator>Anders Rasmussen</dc:creator>
      <pubDate>Sat, 05 Sep 2026 19:04:05 +0000</pubDate>
      <link>https://dev.to/adrasmussen/gpt-6-astra-claude-fable-gemini-38-a-busy-week-for-new-llm-releases-n60</link>
      <guid>https://dev.to/adrasmussen/gpt-6-astra-claude-fable-gemini-38-a-busy-week-for-new-llm-releases-n60</guid>
      <description>&lt;h1&gt;
  
  
  GPT-6 Astra, Claude Fable, Gemini 3.8: A Busy Week for New LLM Releases
&lt;/h1&gt;

&lt;p&gt;&lt;em&gt;Weekly digest from &lt;a href="https://llmpricewatch.com/" rel="noopener noreferrer"&gt;LLM Price Watch&lt;/a&gt; — tracking real pricing data across OpenAI, Anthropic, Google, DeepSeek, and xAI.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;No price changes to report this week across our tracked providers. But the model catalog grew significantly — nine new model IDs showed up in our tracker between September 2nd and 5th, 2026, spanning all four of the big players we watch. Here's what appeared and what's worth noting practically.&lt;/p&gt;

&lt;h2&gt;
  
  
  OpenAI: GPT-6 Astra (and a Pro tier)
&lt;/h2&gt;

&lt;p&gt;The biggest news by name alone: OpenAI dropped two new model families on September 5th — &lt;code&gt;gpt-6-astra&lt;/code&gt; and &lt;code&gt;gpt-6-astra-pro&lt;/code&gt;, each with a corresponding &lt;code&gt;:batch&lt;/code&gt; variant.&lt;/p&gt;

&lt;p&gt;The batch variants are the immediate practical signal here. OpenAI has consistently priced batch inference at a meaningful discount (typically around 50% off standard rates), so if you're running non-latency-sensitive workloads — evaluations, document processing, bulk classification — the batch endpoints are worth watching closely once pricing is confirmed.&lt;/p&gt;

&lt;p&gt;The split between a base &lt;code&gt;gpt-6-astra&lt;/code&gt; and &lt;code&gt;gpt-6-astra-pro&lt;/code&gt; suggests a tiered capability structure, similar to how the GPT-4o / GPT-4o-mini split played out. Whether the Pro variant is meaningfully better for reasoning-heavy tasks or just a marketing distinction remains to be seen. I'd hold off on routing production traffic until pricing and benchmark data are clearer.&lt;/p&gt;

&lt;h2&gt;
  
  
  xAI: Grok 4.3 Batch
&lt;/h2&gt;

&lt;p&gt;On September 4th, &lt;code&gt;grok-4.3:batch&lt;/code&gt; appeared — notably, just the batch variant, with no corresponding standard model showing up in our tracker this week. That's an interesting release pattern. It could mean the base &lt;code&gt;grok-4.3&lt;/code&gt; was already present in the catalog, or that xAI is leading with batch access before a broader rollout. Either way, if you're already using Grok models for async workloads, this is worth checking against your current setup.&lt;/p&gt;

&lt;h2&gt;
  
  
  Google: Gemini 3.8 Flash
&lt;/h2&gt;

&lt;p&gt;September 3rd brought &lt;code&gt;gemini-3.8-flash&lt;/code&gt; and its batch counterpart. The Flash line from Google has generally been their price-performance sweet spot — faster and cheaper than Pro variants, with quality that's often sufficient for structured output tasks, summarization, and classification.&lt;/p&gt;

&lt;p&gt;If the 3.8 Flash follows the same pricing trajectory as earlier Flash models, it could become a solid default for high-volume, cost-sensitive applications. The batch variant showing up simultaneously with the standard model is a good sign — it suggests Google is treating batch access as a first-class feature rather than an afterthought.&lt;/p&gt;

&lt;h2&gt;
  
  
  Anthropic: Claude Fable 5.1
&lt;/h2&gt;

&lt;p&gt;The most intriguing naming this week: &lt;code&gt;claude-fable-5.1&lt;/code&gt; appeared on September 2nd, again with a batch variant. "Fable" isn't a naming convention we've seen from Anthropic before — their previous lines have been Haiku, Sonnet, and Opus. A new name likely means a new positioning, possibly a specialized or fine-tuned variant rather than a general-purpose tier.&lt;/p&gt;

&lt;p&gt;Without more documentation from Anthropic, it's hard to know where Fable sits in their lineup — whether it's meant to replace something, complement the existing tiers, or target a specific use case. The 5.1 version number suggests it's not a ground-up new model but an iteration of something in the Claude 5 family.&lt;/p&gt;

&lt;h2&gt;
  
  
  What This Means Practically
&lt;/h2&gt;

&lt;p&gt;A week with nine new model IDs and zero price changes is a week for watching, not necessarily acting. A few things I'd suggest:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Don't rush to migrate&lt;/strong&gt; to any of these until pricing is confirmed and published. New model IDs on OpenRouter don't always come with finalized pricing on day one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Batch variants are worth bookmarking&lt;/strong&gt; across all four providers. If your workload tolerates async processing, batch endpoints typically offer the best cost efficiency available.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The Anthropic "Fable" naming&lt;/strong&gt; is unusual enough to keep an eye on — it may signal a meaningful capability shift or specialization that affects how you'd evaluate it against Sonnet or Haiku.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I'll be tracking pricing as it settles across all nine of these models in next week's digest.&lt;/p&gt;




&lt;p&gt;For live pricing and model tracking across Claude, GPT, Gemini, DeepSeek, and Grok, visit &lt;strong&gt;&lt;a href="https://llmpricewatch.com/" rel="noopener noreferrer"&gt;llmpricewatch.com&lt;/a&gt;&lt;/strong&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>api</category>
      <category>webdev</category>
    </item>
    <item>
      <title>New Batch Models From DeepSeek, Google, and OpenAI: LLM Pricing Digest</title>
      <dc:creator>Anders Rasmussen</dc:creator>
      <pubDate>Sun, 30 Aug 2026 10:01:38 +0000</pubDate>
      <link>https://dev.to/adrasmussen/new-batch-models-from-deepseek-google-and-openai-llm-pricing-digest-1pok</link>
      <guid>https://dev.to/adrasmussen/new-batch-models-from-deepseek-google-and-openai-llm-pricing-digest-1pok</guid>
      <description>&lt;h2&gt;
  
  
  LLM Price Watch Weekly Digest
&lt;/h2&gt;

&lt;p&gt;No price changes this week across the five providers we track — Claude, GPT, Gemini, DeepSeek, and Grok all held steady. But there was a quiet flurry of new model IDs showing up on OpenRouter, all of them batch variants, and worth knowing about if you're running async workloads.&lt;/p&gt;

&lt;p&gt;Here's what we spotted on August 29th.&lt;/p&gt;




&lt;h3&gt;
  
  
  DeepSeek: Two New Batch Models
&lt;/h3&gt;

&lt;p&gt;DeepSeek added two new batch-mode models:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;deepseek/deepseek-v4-pro-0813:batch&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;deepseek/deepseek-v4-flash-0731:batch&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The naming convention tells you a few things. The &lt;code&gt;-pro&lt;/code&gt; vs &lt;code&gt;-flash&lt;/code&gt; split mirrors what we've seen before — pro for heavier reasoning tasks, flash for faster and cheaper throughput. The date suffixes (&lt;code&gt;0813&lt;/code&gt;, &lt;code&gt;0731&lt;/code&gt;) suggest these are specific checkpoints, which is useful if you care about reproducibility or are comparing outputs across versions.&lt;/p&gt;

&lt;p&gt;Batch mode generally means you submit a job and get results back asynchronously, usually at a lower cost per token than the real-time endpoint. If you're doing large-scale data processing, eval runs, or document analysis where latency doesn't matter, batch is almost always the right call. Worth checking whether DeepSeek's batch pricing lands meaningfully below their synchronous rates.&lt;/p&gt;




&lt;h3&gt;
  
  
  Google: Gemma 4 31B Batch
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;google/gemma-4-31b-it:batch&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Gemma 4 at 31B is a reasonable mid-size open-weights model. The &lt;code&gt;it&lt;/code&gt; tag means instruction-tuned. Seeing it show up as a batch endpoint on OpenRouter suggests Google is continuing to expand the Gemma family's API availability beyond just the hosted Gemini models. For teams that want Google's open-weights lineage without committing to Gemini pricing, this is worth a look — especially for batch jobs where you can afford to wait.&lt;/p&gt;




&lt;h3&gt;
  
  
  OpenAI: Two "OSS" Models
&lt;/h3&gt;

&lt;p&gt;These two caught my eye:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;openai/gpt-oss-120b:batch&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;openai/gpt-oss-20b:batch&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The &lt;code&gt;oss&lt;/code&gt; label is interesting. OpenAI hasn't broadly publicized models under that naming scheme, so I'm not going to speculate too much about what's behind it. What's observable is that there's a 120B and a 20B variant, both batch-only for now, both appearing on OpenRouter. The size gap between them is large — 120B puts it in flagship territory, 20B is more of an efficient workhorse. If these are genuinely open or open-weight models from OpenAI, that would be notable, but I'd wait for official documentation before drawing conclusions.&lt;/p&gt;

&lt;p&gt;For now, treat them as "spotted in the wild" and monitor whether pricing and model cards show up.&lt;/p&gt;




&lt;h3&gt;
  
  
  What This Week Means Practically
&lt;/h3&gt;

&lt;p&gt;If you're actively choosing models right now:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Batch is underused.&lt;/strong&gt; Most developers I talk to default to synchronous endpoints even when their use case has no real latency requirement. If you're running nightly pipelines, processing large document sets, or doing offline evals, switching to batch can cut costs significantly. This week's additions give you more batch options across providers.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;DeepSeek continues to be prolific.&lt;/strong&gt; Two new models in one week, both with dated checkpoints, suggests active development. If you're relying on DeepSeek for production workloads, pinning to a specific checkpoint (like &lt;code&gt;0813&lt;/code&gt;) is probably smarter than using a floating alias.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;OpenAI's OSS naming is worth watching.&lt;/strong&gt; No firm conclusions yet, but if these turn out to be open-weight models with competitive pricing, that changes the calculus for teams currently choosing between self-hosting and API access.&lt;/p&gt;

&lt;p&gt;No pricing drama this week — just new surface area to explore. I'll flag it here if anything changes on the cost side.&lt;/p&gt;




&lt;p&gt;We track live pricing and new model appearances for Claude, GPT, Gemini, DeepSeek, and Grok at &lt;strong&gt;&lt;a href="https://llmpricewatch.com/" rel="noopener noreferrer"&gt;llmpricewatch.com&lt;/a&gt;&lt;/strong&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>api</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>LLM Pricing This Week: DeepSeek Quietly Drops a Vision Model</title>
      <dc:creator>Anders Rasmussen</dc:creator>
      <pubDate>Sun, 23 Aug 2026 10:02:11 +0000</pubDate>
      <link>https://dev.to/adrasmussen/llm-pricing-this-week-deepseek-quietly-drops-a-vision-model-13eo</link>
      <guid>https://dev.to/adrasmussen/llm-pricing-this-week-deepseek-quietly-drops-a-vision-model-13eo</guid>
      <description>&lt;h1&gt;
  
  
  LLM Pricing This Week: DeepSeek Quietly Drops a Vision Model
&lt;/h1&gt;

&lt;p&gt;&lt;em&gt;Weekly digest from &lt;a href="https://llmpricewatch.com/" rel="noopener noreferrer"&gt;LLM Price Watch&lt;/a&gt; — we track live pricing for Claude, GPT, Gemini, DeepSeek, and Grok so you don't have to.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;It was a quiet week on the pricing front. No cuts, no hikes, nothing moved across the five providers we track. If you're mid-project and budgeting around current rates, your spreadsheet is still valid.&lt;/p&gt;

&lt;p&gt;The only thing worth noting this week is a new model showing up in the wild.&lt;/p&gt;

&lt;h2&gt;
  
  
  New: deepseek/deepseek-v4-flash-vision-exp
&lt;/h2&gt;

&lt;p&gt;On August 22nd, our tracker spotted &lt;code&gt;deepseek/deepseek-v4-flash-vision-exp&lt;/code&gt; appearing on OpenRouter for the first time. It's sitting under DeepSeek's lineup, and the name tells you a few things:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;v4&lt;/strong&gt; — this is positioned as their fourth-generation base&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;flash&lt;/strong&gt; — expect a speed/cost optimized variant rather than a full-capability flagship&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;vision&lt;/strong&gt; — multimodal, so it can handle image inputs alongside text&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;exp&lt;/strong&gt; — experimental, meaning DeepSeek hasn't called this production-ready yet&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The &lt;code&gt;exp&lt;/code&gt; tag is worth paying attention to. Experimental models from any provider can change behavior, pricing, or disappear entirely without much notice. If you're evaluating this for something that needs to stay stable, treat it as a preview rather than a dependency.&lt;/p&gt;

&lt;p&gt;That said, DeepSeek's "flash" tier has generally meant aggressive pricing in the past, and a vision-capable model in that tier is interesting. A lot of use cases — document parsing, receipt extraction, basic image captioning — don't need the heaviest multimodal model available. If DeepSeek prices this the way they've priced their other flash variants, it could be worth benchmarking against GPT-4o mini or Gemini Flash for vision tasks.&lt;/p&gt;

&lt;p&gt;We don't have confirmed pricing locked in for this model yet since it just appeared. We'll update the tracker as that settles.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Bigger Picture This Week
&lt;/h2&gt;

&lt;p&gt;Honestly, a week with no price changes and one experimental model drop is pretty normal. The pace of model releases has been fast enough over the past year that it's easy to assume something significant happens every week, but sometimes the answer is just: nothing moved, here's what's new.&lt;/p&gt;

&lt;p&gt;A few things I'd keep in mind heading into next week:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;DeepSeek continues to ship fast.&lt;/strong&gt; Whether or not &lt;code&gt;deepseek-v4-flash-vision-exp&lt;/code&gt; turns into a production model worth using, the cadence from DeepSeek has been consistent. They tend to iterate in public, so experimental models often do graduate to stable relatively quickly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Vision is getting commoditized.&lt;/strong&gt; A year ago, image input was a premium feature. Now it's showing up in flash/lite tiers across providers. If you're paying for a heavier model primarily because you need vision and assumed cheaper options didn't have it, it's worth re-checking the landscape.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;No price changes this week doesn't mean no price pressure.&lt;/strong&gt; Several providers have cut prices in 2025, and competition hasn't slowed. Stable prices one week doesn't signal a plateau — it just means nothing happened this particular week.&lt;/p&gt;




&lt;p&gt;That's the week. One new experimental model from DeepSeek, everything else held steady. If you want to keep an eye on pricing as it shifts — especially for DeepSeek's new addition once pricing confirms — the tracker is running daily checks.&lt;/p&gt;

&lt;p&gt;→ &lt;a href="https://llmpricewatch.com/" rel="noopener noreferrer"&gt;llmpricewatch.com&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>api</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>New Models From Google, DeepSeek, and xAI This Week — LLM Price Watch Digest</title>
      <dc:creator>Anders Rasmussen</dc:creator>
      <pubDate>Sun, 16 Aug 2026 10:01:33 +0000</pubDate>
      <link>https://dev.to/adrasmussen/new-models-from-google-deepseek-and-xai-this-week-llm-price-watch-digest-37lp</link>
      <guid>https://dev.to/adrasmussen/new-models-from-google-deepseek-and-xai-this-week-llm-price-watch-digest-37lp</guid>
      <description>&lt;h2&gt;
  
  
  LLM Price Watch Weekly Digest — Week of August 14, 2026
&lt;/h2&gt;

&lt;p&gt;No price changes to report this week across Claude, GPT, Gemini, DeepSeek, and Grok. That's actually notable in itself — it's been a period of relative pricing stability. But it wasn't a quiet week model-wise. Four new models showed up in our tracker across three providers, and they're worth knowing about if you're evaluating what to build on.&lt;/p&gt;




&lt;h3&gt;
  
  
  Google: Gemini 3.7 Flash (and a Batch Variant)
&lt;/h3&gt;

&lt;p&gt;The two biggest additions are &lt;code&gt;google/gemini-3.7-flash&lt;/code&gt; and &lt;code&gt;google/gemini-3.7-flash:batch&lt;/code&gt;, both first spotted on August 14th.&lt;/p&gt;

&lt;p&gt;Gemini Flash has been Google's go-to for cost-efficient, lower-latency tasks — the kind of workloads where you need volume without burning through budget. The 3.7 iteration landing now suggests Google is continuing to iterate on that line in parallel with their heavier models.&lt;/p&gt;

&lt;p&gt;The batch variant is the one I'd pay attention to if you're running anything offline — document processing, evals, bulk classification, that sort of thing. Batch endpoints typically come with a meaningful price discount in exchange for higher latency, so if your use case doesn't need a real-time response, it's worth checking the pricing before defaulting to the standard endpoint. We'll have the confirmed pricing posted on LLM Price Watch as soon as it's stable in the tracker.&lt;/p&gt;




&lt;h3&gt;
  
  
  DeepSeek: v4 Pro 0813
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;deepseek/deepseek-v4-pro-0813&lt;/code&gt; showed up on August 13th. The &lt;code&gt;0813&lt;/code&gt; suffix is a date stamp — a common convention DeepSeek uses to version checkpoint releases. This appears to be a refreshed checkpoint of DeepSeek v4 Pro rather than an entirely new architecture.&lt;/p&gt;

&lt;p&gt;DeepSeek has been one of the more interesting providers to watch from a price-to-performance standpoint over the past year. Their models have consistently undercut comparable Western models on cost. If you've been using an earlier v4 Pro checkpoint, it's worth running a quick eval against this one to see if quality has shifted. Sometimes these point releases are minor; sometimes they're not.&lt;/p&gt;




&lt;h3&gt;
  
  
  xAI: Grok 4.6
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;x-ai/grok-4.6&lt;/code&gt; also appeared on August 13th. xAI has been incrementing Grok's version numbers at a reasonable clip. Grok 4.6 sits between major releases, so this reads like a refinement update — likely improvements to instruction following or reasoning rather than a fundamentally different model.&lt;/p&gt;

&lt;p&gt;Grok has carved out a niche for users who want strong general-purpose performance and are already in the xAI ecosystem. Pricing on this one we'll confirm as the data firms up.&lt;/p&gt;




&lt;h3&gt;
  
  
  What This Week Means Practically
&lt;/h3&gt;

&lt;p&gt;If you're in the middle of a model selection decision right now, here's how I'd think about these additions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;High-volume, latency-tolerant workloads&lt;/strong&gt;: Look at &lt;code&gt;gemini-3.7-flash:batch&lt;/code&gt; once pricing is confirmed. Batch endpoints are consistently underused by developers who could benefit from them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost-sensitive inference&lt;/strong&gt;: Keep an eye on the DeepSeek v4 Pro 0813 pricing. DeepSeek tends to be aggressive here and a checkpoint update sometimes comes with a price adjustment.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;General-purpose API work&lt;/strong&gt;: Grok 4.6 is worth a quick benchmark if you're already evaluating xAI models.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The lack of price changes this week means your existing cost projections are still valid — no surprises there. But new model versions mean your performance baselines might shift if providers quietly improve quality at the same price point.&lt;/p&gt;




&lt;p&gt;I track all of this daily at &lt;strong&gt;&lt;a href="https://llmpricewatch.com/" rel="noopener noreferrer"&gt;LLM Price Watch&lt;/a&gt;&lt;/strong&gt; — live pricing for Claude, GPT, Gemini, DeepSeek, and Grok in one place, with alerts when something changes. Worth bookmarking if you're making cost-sensitive model decisions.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>api</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Two New Models Just Dropped: DeepSeek V4 Pro and Grok 4.6 — LLM Pricing Digest</title>
      <dc:creator>Anders Rasmussen</dc:creator>
      <pubDate>Thu, 13 Aug 2026 14:47:08 +0000</pubDate>
      <link>https://dev.to/adrasmussen/two-new-models-just-dropped-deepseek-v4-pro-and-grok-46-llm-pricing-digest-8ee</link>
      <guid>https://dev.to/adrasmussen/two-new-models-just-dropped-deepseek-v4-pro-and-grok-46-llm-pricing-digest-8ee</guid>
      <description>&lt;h1&gt;
  
  
  Two New Models Just Dropped: DeepSeek V4 Pro and Grok 4.6 — LLM Pricing Digest
&lt;/h1&gt;

&lt;p&gt;&lt;em&gt;Weekly roundup from &lt;a href="https://llmpricewatch.com/" rel="noopener noreferrer"&gt;LLM Price Watch&lt;/a&gt; — we track live pricing across Claude, GPT, Gemini, DeepSeek, and Grok so you don't have to.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;No price cuts or hikes to report this week across our tracked providers. What we did catch: two new model IDs appearing in the wild on August 13th — one from DeepSeek and one from xAI. Here's what we know so far.&lt;/p&gt;

&lt;h2&gt;
  
  
  DeepSeek V4 Pro (0813)
&lt;/h2&gt;

&lt;p&gt;Model ID: &lt;code&gt;deepseek/deepseek-v4-pro-0813&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;DeepSeek continues their habit of date-stamping model releases, which is actually a useful practice — it makes it easy to tell at a glance which checkpoint you're running. The "V4 Pro" naming suggests this sits above their previous V3 line, though we don't yet have benchmark comparisons or confirmed context window specs to share. The &lt;code&gt;0813&lt;/code&gt; suffix matches the date it first appeared in our tracker.&lt;/p&gt;

&lt;p&gt;DeepSeek has consistently been one of the more cost-competitive options in our index, so if V4 Pro follows that pattern, it's worth watching. We'll update pricing on the site as soon as it's confirmed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Practical take:&lt;/strong&gt; If you're currently running DeepSeek V3 in production and cost-efficiency is a priority, keep an eye on how V4 Pro prices in. DeepSeek's previous generational jumps have held the line on price while improving capability — but don't assume that until the numbers are public.&lt;/p&gt;

&lt;h2&gt;
  
  
  Grok 4.6
&lt;/h2&gt;

&lt;p&gt;Model ID: &lt;code&gt;x-ai/grok-4.6&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;xAI is pushing another point release with Grok 4.6, spotted the same day as the DeepSeek drop. Grok 4 was already in our tracker, so this is an incremental update rather than a major new generation. Point releases from xAI have sometimes come with quiet capability improvements or adjusted rate limits, so it's worth checking if you're an active Grok user.&lt;/p&gt;

&lt;p&gt;Pricing for 4.6 isn't confirmed in our system yet. Grok 4 has sat at a premium compared to some of the DeepSeek options, so if 4.6 comes in at the same price point, the question is whether the incremental improvements justify staying on that tier versus alternatives.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Practical take:&lt;/strong&gt; If you're already using Grok 4 via API, test 4.6 on your actual use cases before migrating. Point releases don't always move the needle on the tasks that matter most to your specific workload.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Bigger Picture This Week
&lt;/h2&gt;

&lt;p&gt;Two new models, zero price changes. That's actually a notable signal in itself. We're in a period where the model release cadence remains fast, but the pricing floor seems to have stabilized — at least for this week. The race to the bottom on token pricing that defined much of early 2025 appears to have leveled off, at least temporarily.&lt;/p&gt;

&lt;p&gt;For anyone building on top of these APIs right now, the practical advice is the same as always: don't hard-code model names if you can avoid it, because the &lt;code&gt;deepseek-v4-pro-0813&lt;/code&gt; you integrate today may be superseded by a &lt;code&gt;-0901&lt;/code&gt; variant before your next sprint is done. Abstract your model selection layer where possible.&lt;/p&gt;

&lt;p&gt;We'll be watching both of these new models closely over the coming days as pricing details get confirmed and early benchmarks surface.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I track these changes daily at *&lt;/em&gt;&lt;a href="https://llmpricewatch.com/" rel="noopener noreferrer"&gt;llmpricewatch.com&lt;/a&gt;** — live pricing for Claude, GPT, Gemini, DeepSeek, and Grok, updated automatically. If you want alerts when something changes, the site has you covered.*&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>api</category>
      <category>webdev</category>
    </item>
  </channel>
</rss>
