<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: 4663437Mehdi</title>
    <description>The latest articles on DEV Community by 4663437Mehdi (@4663437mehdi).</description>
    <link>https://dev.to/4663437mehdi</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3933985%2Fd12e89b2-cc42-404c-b494-5ebd7577086c.png</url>
      <title>DEV Community: 4663437Mehdi</title>
      <link>https://dev.to/4663437mehdi</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/4663437mehdi"/>
    <language>en</language>
    <item>
      <title>The Token Ledger – 2026-08-06</title>
      <dc:creator>4663437Mehdi</dc:creator>
      <pubDate>Thu, 06 Aug 2026 09:50:33 +0000</pubDate>
      <link>https://dev.to/4663437mehdi/the-token-ledger-2026-08-06-478n</link>
      <guid>https://dev.to/4663437mehdi/the-token-ledger-2026-08-06-478n</guid>
      <description>&lt;h1&gt;
  
  
  The Token Ledger – 2026-08-06
&lt;/h1&gt;

&lt;p&gt;&lt;strong&gt;Most cost‑impacting change:&lt;/strong&gt; Qwen: Qwen3.6 27B saw a major price increase. Prompt rose from $0.289 to $0.60 per 1M tokens (+$0.311); completion rose from $2.40 to $3.60 per 1M tokens (+$1.20). Total cost per 1M tokens is now $4.20, up $1.511. Teams running large‑scale inference or fine‑tuning on this model should reassess budget allocations.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Added models&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Meta: Muse Spark 1.2 – prompt $1.25 / 1M, completion $4.25 / 1M. Ideal for developers needing long‑context (1M tokens) generative tasks.&lt;/li&gt;
&lt;li&gt;inclusionAI: Ling-3.0-flash – prompt $0.075 / 1M, completion $0.22 / 1M. Suitable for low‑latency, cost‑sensitive applications with moderate context (128K).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Price changes&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;MoonshotAI: Kimi K2.7 Code – prompt down $0.03 to $0.70 / 1M; completion unchanged at $3.50 / 1M. Minor savings for code‑generation workloads.&lt;/li&gt;
&lt;li&gt;DeepSeek: DeepSeek V4 Flash 0423 – prompt down $0.0518 to $0.0882 / 1M; completion down $0.1036 to $0.1764 / 1M. Total reduction $0.1554 / 1M; beneficial for high‑volume flash inference.&lt;/li&gt;
&lt;li&gt;Z.ai: GLM 5.1 – prompt down $0.014 to $0.952 / 1M; completion down $0.044 to $2.992 / 1M. Small overall cut ($0.058 / 1M) for balanced prompt/completion use.&lt;/li&gt;
&lt;li&gt;MiniMax: MiniMax M2.5 – prompt up $0.07 to $0.22 / 1M; completion steady at $0.90 / 1M. Slight cost rise for users of this model.&lt;/li&gt;
&lt;li&gt;Qwen: Qwen3 235B A22B Instruct 2507 – prompt down $0.0595 to $0.09 / 1M; completion down $0.048 to $0.55 / 1M. Total saving $0.1075 / 1M, relevant for large‑model deployments.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Cheapest models today&lt;/strong&gt; (for reference): inclusionAI: Ling-2.6-flash ($0.01 / 1M prompt, $0.03 / 1M completion), IBM: Granite 4.0 Micro ($0.017 / 1M prompt, $0.112 / 1M completion), Mistral: Mistral Nemo ($0.019 / 1M prompt, $0.03 / 1M completion).&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://4663437Mehdi.github.io/token-ledger/entry.html?d=2026-08-06" rel="noopener noreferrer"&gt;The Token Ledger&lt;/a&gt;. Subscribe for the daily digest.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>api</category>
      <category>news</category>
    </item>
    <item>
      <title>The Token Ledger – 2026-08-05</title>
      <dc:creator>4663437Mehdi</dc:creator>
      <pubDate>Wed, 05 Aug 2026 09:47:02 +0000</pubDate>
      <link>https://dev.to/4663437mehdi/the-token-ledger-2026-08-05-422f</link>
      <guid>https://dev.to/4663437mehdi/the-token-ledger-2026-08-05-422f</guid>
      <description>&lt;h1&gt;
  
  
  The Token Ledger – 2026-08-05
&lt;/h1&gt;

&lt;h2&gt;
  
  
  MiniMax: MiniMax M2.7
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;What changed:&lt;/strong&gt; Prompt and completion prices increased.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Numbers (per 1M tokens):&lt;/strong&gt;

&lt;ul&gt;
&lt;li&gt;Prompt: $0.25 → $0.27 (+$0.02)
&lt;/li&gt;
&lt;li&gt;Completion: $1.00 → $1.08 (+$0.08)
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Who should care:&lt;/strong&gt; Teams running high‑volume completion workloads on MiniMax; the $0.08/M rise adds $80 per 1B tokens.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Mistral: Mistral Small 3.2 24B
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;What changed:&lt;/strong&gt; Prompt and completion prices increased.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Numbers (per 1M tokens):&lt;/strong&gt;

&lt;ul&gt;
&lt;li&gt;Prompt: $0.075 → $0.09375 (+$0.01875)
&lt;/li&gt;
&lt;li&gt;Completion: $0.20 → $0.25 (+$0.05)
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Who should care:&lt;/strong&gt; Users of Mistral’s 24B instruct model; the $0.05/M completion uplift adds $50 per 1B tokens.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;No models were added or removed today. Total models tracked: 338.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://4663437Mehdi.github.io/token-ledger/entry.html?d=2026-08-05" rel="noopener noreferrer"&gt;The Token Ledger&lt;/a&gt;. Subscribe for the daily digest.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>api</category>
      <category>news</category>
    </item>
    <item>
      <title>The Token Ledger Digest – 2026-08-04</title>
      <dc:creator>4663437Mehdi</dc:creator>
      <pubDate>Tue, 04 Aug 2026 09:49:18 +0000</pubDate>
      <link>https://dev.to/4663437mehdi/the-token-ledger-digest-2026-08-04-4gnb</link>
      <guid>https://dev.to/4663437mehdi/the-token-ledger-digest-2026-08-04-4gnb</guid>
      <description>&lt;h1&gt;
  
  
  The Token Ledger Digest – 2026-08-04
&lt;/h1&gt;

&lt;p&gt;&lt;strong&gt;Z.ai: GLM 5.2&lt;/strong&gt; – Completion price fell from &lt;strong&gt;$3.74&lt;/strong&gt; to &lt;strong&gt;$2.42&lt;/strong&gt; per 1M tokens (‑$1.32/M). Prompt price dropped from &lt;strong&gt;$1.19&lt;/strong&gt; to &lt;strong&gt;$0.76&lt;/strong&gt; per 1M tokens (‑$0.43/M). &lt;em&gt;Who should care:&lt;/em&gt; Teams running cost‑sensitive, long‑completion workloads on Z.ai.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Qwen: Qwen3.5‑122B‑A10B&lt;/strong&gt; – Prompt ↓ $0.40→$0.26/M (‑$0.14/M). Completion ↓ $3.20→$2.08/M (‑$1.12/M). &lt;em&gt;Who should care:&lt;/em&gt; Users of large‑scale Qwen models seeking lower generation cost.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Qwen: Qwen2.5 VL 72B Instruct&lt;/strong&gt; – Prompt ↓ $0.80→$0.25/M (‑$0.55/M). Completion ↓ $1.00→$0.75/M (‑$0.25/M). &lt;em&gt;Who should care:&lt;/em&gt; Vision‑language apps needing cheaper prompt processing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;MoonshotAI: Kimi K2.6&lt;/strong&gt; – Prompt ↓ $0.60→$0.589/M (‑$0.011/M). Completion ↓ $3.41→$2.48/M (‑$0.93/M). &lt;em&gt;Who should care:&lt;/em&gt; Developers balancing prompt and completion costs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Qwen: Qwen3.6 27B&lt;/strong&gt; – Prompt ↓ $0.30→$0.289/M (‑$0.011/M). Completion ↑ $2.00→$2.40/M (+$0.40/M). &lt;em&gt;Who should care:&lt;/em&gt; Slightly cheaper prompts but higher completion expense.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Meta: Llama 3.3 70B Instruct&lt;/strong&gt; – Prompt ↓ $0.13→$0.10/M (‑$0.03/M). Completion ↓ $0.40→$0.32/M (‑$0.08/M). &lt;em&gt;Who should care:&lt;/em&gt; General‑purpose Llama users seeing modest savings.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Qwen: Qwen3 Next 80B A3B Instruct&lt;/strong&gt; – Prompt ↓ $0.10→$0.09/M (‑$0.01/M). Completion unchanged at $1.10/M. &lt;em&gt;Who should care:&lt;/em&gt; Minor prompt‑cost reduction.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Qwen: Qwen3 Coder 30B A3B Instruct&lt;/strong&gt; – Prompt unchanged $0.07/M. Completion ↓ $0.28→$0.27/M (‑$0.01/M). &lt;em&gt;Who should care:&lt;/em&gt; Tiny completion‑cost trim.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Qwen: Qwen3 VL 30B A3B Instruct&lt;/strong&gt; – Prompt ↑ $0.13→$0.15/M (+$0.02/M). Completion ↑ $0.52→$0.60/M (+$0.0. &lt;em&gt;Who should care:&lt;/em&gt; Slight costlier.&lt;/p&gt;

&lt;p&gt;*&lt;em&gt;M0.08/M). *Who should care:&lt;/em&gt; Small increase for vision‑language tasks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;MythoMax 13B&lt;/strong&gt; – Prompt ↑ $0.06→$0.08/M (+$0.02/M). Completion ↑ $0.06→$0.11/M (+$0.05/M). &lt;em&gt;Who should care:&lt;/em&gt; Minor cost rise for this model.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Added:&lt;/strong&gt; &lt;strong&gt;Qwen: Qwen3.8 Max&lt;/strong&gt; – Prompt $2.00/M, Completion $6.00/M, 1M‑token context. &lt;em&gt;Who should care:&lt;/em&gt; New high‑capacity option for enterprises needing massive context windows.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://4663437Mehdi.github.io/token-ledger/entry.html?d=2026-08-04" rel="noopener noreferrer"&gt;The Token Ledger&lt;/a&gt;. Subscribe for the daily digest.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>api</category>
      <category>news</category>
    </item>
    <item>
      <title>The Token Ledger Digest – 2026-08-03</title>
      <dc:creator>4663437Mehdi</dc:creator>
      <pubDate>Mon, 03 Aug 2026 10:50:55 +0000</pubDate>
      <link>https://dev.to/4663437mehdi/the-token-ledger-digest-2026-08-03-4gc4</link>
      <guid>https://dev.to/4663437mehdi/the-token-ledger-digest-2026-08-03-4gc4</guid>
      <description>&lt;h1&gt;
  
  
  The Token Ledger Digest – 2026-08-03
&lt;/h1&gt;

&lt;p&gt;&lt;strong&gt;Most cost‑impacting change:&lt;/strong&gt; Z.ai’s GLM 5.2 saw a sharp price rise, raising both prompt and completion costs by roughly $0.90‑$2.85 per 1M tokens.&lt;/p&gt;

&lt;h2&gt;
  
  
  Price changes
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;What changed&lt;/th&gt;
&lt;th&gt;Old price (/1M)&lt;/th&gt;
&lt;th&gt;New price (/1M)&lt;/th&gt;
&lt;th&gt;Who should care&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Z.ai: GLM 5.2&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Prompt ↑ from $0.28 to $1.19; Completion ↑ from $0.89 to $3.74&lt;/td&gt;
&lt;td&gt;$0.28 / $0.89&lt;/td&gt;
&lt;td&gt;$1.19 / $3.74&lt;/td&gt;
&lt;td&gt;Teams using GLM 5.2 for high‑volume generation; budget reassessment needed.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Qwen: Qwen3.5‑122B‑A10B&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Prompt ↑ $0.26 → $0.40; Completion ↑ $2.08 → $3.20&lt;/td&gt;
&lt;td&gt;$0.26 / $2.08&lt;/td&gt;
&lt;td&gt;$0.40 / $3.20&lt;/td&gt;
&lt;td&gt;Developers balancing quality vs cost for mid‑size LLMs.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Qwen: Qwen3 VL 235B A22B Thinking&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Prompt ↑ $0.40 → $0.98; Completion ↓ $4.00 → $3.95&lt;/td&gt;
&lt;td&gt;$0.40 / $4.00&lt;/td&gt;
&lt;td&gt;$0.98 / $3.95&lt;/td&gt;
&lt;td&gt;Vision‑language users; prompt cost up slightly, completion marginally cheaper.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Qwen: Qwen3 235B A22B Instruct 2507&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Prompt ↑ $0.09 → $0.15; Completion ↑ $0.55 → $0.60&lt;/td&gt;
&lt;td&gt;$0.09 / $0.55&lt;/td&gt;
&lt;td&gt;$0.15 / $0.60&lt;/td&gt;
&lt;td&gt;Cost‑sensitive applications using this instruct variant.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;OpenAI: GPT‑5.6 Luna Pro&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;No meaningful change (prompt $0.10, completion $0.60 both unchanged)&lt;/td&gt;
&lt;td&gt;$0.10 / $0.60&lt;/td&gt;
&lt;td&gt;$0.10 / $0.60&lt;/td&gt;
&lt;td&gt;No action required.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;OpenAI: GPT‑5.6 Luna&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Same as above&lt;/td&gt;
&lt;td&gt;$0.10 / $0.60&lt;/td&gt;
&lt;td&gt;$0.10 / $0.60&lt;/td&gt;
&lt;td&gt;No action required.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;OpenAI: GPT‑5.6 Terra Pro&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;No meaningful change (prompt $1.00, completion $6.00)&lt;/td&gt;
&lt;td&gt;$1.00 / $6.00&lt;/td&gt;
&lt;td&gt;$1.00 / $6.00&lt;/td&gt;
&lt;td&gt;No action required.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;OpenAI: GPT‑5.6 Terra&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Same as above&lt;/td&gt;
&lt;td&gt;$1.00 / $6.00&lt;/td&gt;
&lt;td&gt;$1.00 / $6.00&lt;/td&gt;
&lt;td&gt;No action required.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Per‑token prices were multiplied by 1,000,000 and rounded to two decimal places for readability.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Cheapest models today (per‑million tokens)
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;inclusionAI: Ling‑2.6‑flash&lt;/strong&gt; – Prompt $0.01, Completion $0.03
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mistral: Mistral Nemo&lt;/strong&gt; – Prompt $0.02, Completion $0.03
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;IBM: Granite 4.0 Micro&lt;/strong&gt; – Prompt $0.02, Completion $0.11
&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;No models were added or removed today. Total models tracked: 337.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://4663437Mehdi.github.io/token-ledger/entry.html?d=2026-08-03" rel="noopener noreferrer"&gt;The Token Ledger&lt;/a&gt;. Subscribe for the daily digest.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>api</category>
      <category>news</category>
    </item>
    <item>
      <title>The Token Ledger Digest – 2026-08-02</title>
      <dc:creator>4663437Mehdi</dc:creator>
      <pubDate>Sun, 02 Aug 2026 09:19:16 +0000</pubDate>
      <link>https://dev.to/4663437mehdi/the-token-ledger-digest-2026-08-02-2bem</link>
      <guid>https://dev.to/4663437mehdi/the-token-ledger-digest-2026-08-02-2bem</guid>
      <description>&lt;h1&gt;
  
  
  The Token Ledger Digest – 2026-08-02
&lt;/h1&gt;

&lt;p&gt;&lt;strong&gt;Price Drop – Z.ai: GLM 5.2&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Prompt:&lt;/strong&gt; fell from $0.76 /1M to $0.28 /1M (‑63%).
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Completion:&lt;/strong&gt; fell from $2.39 /1M to $0.89 /1M (‑63%).
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Who should care:&lt;/strong&gt; Teams running large‑scale inference or fine‑tuning on GLM 5.2 see per‑token costs cut by roughly two‑thirds.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Price Drop – DeepSeek: DeepSeek V4 Flash 0731&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Prompt:&lt;/strong&gt; dropped from $0.14 /1M to $0.09 /1M (‑36%).
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Completion:&lt;/strong&gt; dropped from $0.28 /1M to $0.18 /1M (‑36%).
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Who should care:&lt;/strong&gt; Users of the 0731 checkpoint benefit from lower latency‑cost trade‑offs for flash‑style workloads.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Price Increase – NVIDIA: Nemotron 3 Ultra&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Prompt:&lt;/strong&gt; rose from $0.50 /1M to $0.60 /1M (+20%).
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Completion:&lt;/strong&gt; rose from $2.20 /1M to $3.60 /1M (+64%).
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Who should care:&lt;/strong&gt; Cost‑sensitive applications relying on Nemotron 3 Ultra for completion‑heavy tasks will see noticeably higher bills.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Added Model – DeepSeek V4 Flash Latest&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Prompt:&lt;/strong&gt; $0.09 /1M
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Completion:&lt;/strong&gt; $0.18 /1M
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Context:&lt;/strong&gt; 1,048,576 tokens
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Who should care:&lt;/strong&gt; Developers needing ultra‑long context with flash‑level pricing can now access this newest variant.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;Total models tracked: 337.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://4663437Mehdi.github.io/token-ledger/entry.html?d=2026-08-02" rel="noopener noreferrer"&gt;The Token Ledger&lt;/a&gt;. Subscribe for the daily digest.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>api</category>
      <category>news</category>
    </item>
    <item>
      <title>Token Ledger Digest – 2026-08-01</title>
      <dc:creator>4663437Mehdi</dc:creator>
      <pubDate>Sat, 01 Aug 2026 09:19:05 +0000</pubDate>
      <link>https://dev.to/4663437mehdi/token-ledger-digest-2026-08-01-3bd6</link>
      <guid>https://dev.to/4663437mehdi/token-ledger-digest-2026-08-01-3bd6</guid>
      <description>&lt;h1&gt;
  
  
  Token Ledger Digest – 2026-08-01
&lt;/h1&gt;

&lt;p&gt;&lt;strong&gt;Most impactful change&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Z.ai: GLM 5.2&lt;/strong&gt; – Prompt price fell from $1.232 / 1M to $0.760 / 1M; completion price fell from $3.872 / 1M to $2.389 / 1M.
&lt;em&gt;Who should care:&lt;/em&gt; Cost‑sensitive applications using this model see ~48% lower prompt and ~38% lower completion costs.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Added&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Thinking Machines: Inkling Small&lt;/strong&gt; – New model with 524 k context, prompt $0.50 / 1M, completion $1.20 / 1M.
&lt;em&gt;Who should care:&lt;/em&gt; Teams needing long‑context generation at low cost.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Removed&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;30 batch models&lt;/strong&gt; were deleted, including Google Gemini 3.6 Flash batch, Anthropic Claude Sonnet 5 batch, and OpenAI GPT‑5.5 batch. Their prompt prices ranged from $0.15 / 1M to $2.50 / 1M and completion prices from $1.25 / 1M to $25.00 / 1M.
&lt;em&gt;Who should care:&lt;/em&gt; Users relying on batch pricing must migrate to alternative endpoints or non‑batch versions.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Other price changes&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;NVIDIA: Nemotron 3 Ultra&lt;/strong&gt; – Prompt $0.60 → $0.50 / 1M; completion $3.60 → $2.20 / 1M.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MoonshotAI: Kimi K2.6&lt;/strong&gt; – Prompt $0.95 → $0.60 / 1M; completion $4.00 → $3.41 / 1M.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Qwen: Qwen3 Coder 30B A3B Instruct&lt;/strong&gt; – Prompt unchanged $0.07 / 1M; completion rose $0.27 → $0.28 / 1M.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mistral: Mistral Small 3.2 24B&lt;/strong&gt; – Prompt $0.10 → $0.075 / 1M; completion $0.30 → $0.20 / 1M.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Cheapest models today&lt;/strong&gt; (for reference)  &lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;inclusionAI: Ling‑2.6‑flash – $0.01 / 1M prompt, $0.03 / 1M completion
&lt;/li&gt;
&lt;li&gt;IBM: Granite 4.0 Micro – $0.017 / 1M prompt, $0.112 / 1M completion
&lt;/li&gt;
&lt;li&gt;Mistral: Mistral Nemo – $0.019 / 1M prompt, $0.03 / 1M completion
&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Total models tracked: 336.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://4663437Mehdi.github.io/token-ledger/entry.html?d=2026-08-01" rel="noopener noreferrer"&gt;The Token Ledger&lt;/a&gt;. Subscribe for the daily digest.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>api</category>
      <category>news</category>
    </item>
    <item>
      <title>Token Ledger Digest – 2026-07-31</title>
      <dc:creator>4663437Mehdi</dc:creator>
      <pubDate>Fri, 31 Jul 2026 09:52:25 +0000</pubDate>
      <link>https://dev.to/4663437mehdi/token-ledger-digest-2026-07-31-3b67</link>
      <guid>https://dev.to/4663437mehdi/token-ledger-digest-2026-07-31-3b67</guid>
      <description>&lt;h1&gt;
  
  
  Token Ledger Digest – 2026-07-31
&lt;/h1&gt;

&lt;h2&gt;
  
  
  Added
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;DeepSeek: DeepSeek V4 Flash 0731&lt;/strong&gt; – New model with 1M‑token context. Prompt $0.14/1M, completion $0.28/1M. &lt;em&gt;Who should care:&lt;/em&gt; Teams needing long‑context generation at low cost.
&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Removed
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;OpenAI: o3 Deep Research&lt;/strong&gt; – Prompt $10.00/1M, completion $40.00/1M.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OpenAI: o4 Mini Deep Research&lt;/strong&gt; – Prompt $2.00/1M, completion $8.00/1M.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OpenAI: GPT-5 Codex&lt;/strong&gt; – Prompt $1.25/1M, completion $10.00/1M.
&lt;em&gt;Who should care:&lt;/em&gt; Any workflows relying on these models must migrate to alternatives.
&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Price Changes
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Old Prompt ($/1M)&lt;/th&gt;
&lt;th&gt;New Prompt ($/1M)&lt;/th&gt;
&lt;th&gt;Δ Prompt&lt;/th&gt;
&lt;th&gt;Old Compl. ($/1M)&lt;/th&gt;
&lt;th&gt;New Compl. ($/1M)&lt;/th&gt;
&lt;th&gt;Δ Compl.&lt;/th&gt;
&lt;th&gt;Who should care&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Poolside: Laguna S 2.1&lt;/td&gt;
&lt;td&gt;0.10&lt;/td&gt;
&lt;td&gt;0.09&lt;/td&gt;
&lt;td&gt;–0.01&lt;/td&gt;
&lt;td&gt;0.20&lt;/td&gt;
&lt;td&gt;0.18&lt;/td&gt;
&lt;td&gt;–0.02&lt;/td&gt;
&lt;td&gt;Cost‑sensitive inference on Laguna S.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenAI: GPT-5.6 Luna Pro&lt;/td&gt;
&lt;td&gt;0.50&lt;/td&gt;
&lt;td&gt;0.10&lt;/td&gt;
&lt;td&gt;–0.40&lt;/td&gt;
&lt;td&gt;3.00&lt;/td&gt;
&lt;td&gt;0.60&lt;/td&gt;
&lt;td&gt;–2.40&lt;/td&gt;
&lt;td&gt;Heavy users of Luna Pro see ~80% cost cut.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenAI: GPT-5.6 Luna&lt;/td&gt;
&lt;td&gt;0.50&lt;/td&gt;
&lt;td&gt;0.10&lt;/td&gt;
&lt;td&gt;–0.40&lt;/td&gt;
&lt;td&gt;3.00&lt;/td&gt;
&lt;td&gt;0.60&lt;/td&gt;
&lt;td&gt;–2.40&lt;/td&gt;
&lt;td&gt;Same impact as Luna Pro.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenAI: GPT-5.6 Terra Pro&lt;/td&gt;
&lt;td&gt;1.25&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;–0.25&lt;/td&gt;
&lt;td&gt;7.50&lt;/td&gt;
&lt;td&gt;6.00&lt;/td&gt;
&lt;td&gt;–1.50&lt;/td&gt;
&lt;td&gt;Terra Pro workloads become cheaper.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenAI: GPT-5.6 Terra&lt;/td&gt;
&lt;td&gt;1.25&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;–0.25&lt;/td&gt;
&lt;td&gt;7.50&lt;/td&gt;
&lt;td&gt;6.00&lt;/td&gt;
&lt;td&gt;–1.50&lt;/td&gt;
&lt;td&gt;Same as Terra Pro.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Z.ai: GLM 5.2&lt;/td&gt;
&lt;td&gt;0.70&lt;/td&gt;
&lt;td&gt;1.23&lt;/td&gt;
&lt;td&gt;+0.53&lt;/td&gt;
&lt;td&gt;2.20&lt;/td&gt;
&lt;td&gt;3.87&lt;/td&gt;
&lt;td&gt;+1.67&lt;/td&gt;
&lt;td&gt;GLM 5.2 now significantly more expensive.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;NVIDIA: Nemotron 3 Ultra&lt;/td&gt;
&lt;td&gt;0.50&lt;/td&gt;
&lt;td&gt;0.60&lt;/td&gt;
&lt;td&gt;+0.10&lt;/td&gt;
&lt;td&gt;2.20&lt;/td&gt;
&lt;td&gt;3.60&lt;/td&gt;
&lt;td&gt;+1.40&lt;/td&gt;
&lt;td&gt;Ultra‑scale users face higher spend.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MoonshotAI Kimi Latest&lt;/td&gt;
&lt;td&gt;2.90&lt;/td&gt;
&lt;td&gt;2.90&lt;/td&gt;
&lt;td&gt;0.00&lt;/td&gt;
&lt;td&gt;15.00&lt;/td&gt;
&lt;td&gt;14.00&lt;/td&gt;
&lt;td&gt;–1.00&lt;/td&gt;
&lt;td&gt;Slight completion‑cost relief.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MoonshotAI: Kimi K2.6&lt;/td&gt;
&lt;td&gt;0.65&lt;/td&gt;
&lt;td&gt;0.95&lt;/td&gt;
&lt;td&gt;+0.30&lt;/td&gt;
&lt;td&gt;2.72&lt;/td&gt;
&lt;td&gt;4.00&lt;/td&gt;
&lt;td&gt;+1.28&lt;/td&gt;
&lt;td&gt;Kimi K2.6 cost rises noticeably.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen: Qwen3 VL 30B A3B Instruct&lt;/td&gt;
&lt;td&gt;0.15&lt;/td&gt;
&lt;td&gt;0.13&lt;/td&gt;
&lt;td&gt;–0.02&lt;/td&gt;
&lt;td&gt;0.60&lt;/td&gt;
&lt;td&gt;0.52&lt;/td&gt;
&lt;td&gt;–0.08&lt;/td&gt;
&lt;td&gt;Minor savings for vision‑language tasks.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen: Qwen3 235B A22B Thinking 2507&lt;/td&gt;
&lt;td&gt;0.30&lt;/td&gt;
&lt;td&gt;0.23&lt;/td&gt;
&lt;td&gt;–0.07&lt;/td&gt;
&lt;td&gt;3.00&lt;/td&gt;
&lt;td&gt;2.30&lt;/td&gt;
&lt;td&gt;–0.70&lt;/td&gt;
&lt;td&gt;Thinking model becomes cheaper.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Most cost‑impacting change:&lt;/strong&gt; The OpenAI GPT-5.6 Luna Pro and Luna models dropped prompt pricing from $0.50 to $0.10/1M and completion from $3.00 to $0.60/1M – an ~80% reduction, saving up to $2.80 per million tokens.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cheapest Models Today
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;inclusionAI: Ling-2.6-flash – $0.01/1M prompt, $0.03/1M completion
&lt;/li&gt;
&lt;li&gt;IBM: Granite 4.0 Micro – $0.017/1M prompt, $0.112/1M completion
&lt;/li&gt;
&lt;li&gt;Mistral: Mistral Nemo – $0.019/1M prompt, $0.03/1M completion
&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;em&gt;Total models tracked: 365.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://4663437Mehdi.github.io/token-ledger/entry.html?d=2026-07-31" rel="noopener noreferrer"&gt;The Token Ledger&lt;/a&gt;. Subscribe for the daily digest.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>api</category>
      <category>news</category>
    </item>
    <item>
      <title>The Token Ledger Digest – 2026-07-30</title>
      <dc:creator>4663437Mehdi</dc:creator>
      <pubDate>Thu, 30 Jul 2026 09:39:59 +0000</pubDate>
      <link>https://dev.to/4663437mehdi/the-token-ledger-digest-2026-07-30-3djl</link>
      <guid>https://dev.to/4663437mehdi/the-token-ledger-digest-2026-07-30-3djl</guid>
      <description>&lt;h1&gt;
  
  
  The Token Ledger Digest – 2026-07-30
&lt;/h1&gt;

&lt;p&gt;&lt;strong&gt;Biggest cost impact:&lt;/strong&gt; DeepSeek DeepSeek V3 – prompt price rose from $0.20 to $0.26 / 1M tokens (+$0.06), completion from $0.80 to $1.03 / 1M tokens (+$0.23). Total increase ≈ $0.29 / 1M tokens.&lt;br&gt;&lt;br&gt;
&lt;em&gt;Who should care:&lt;/em&gt; Teams running high‑volume chat or code‑gen workloads on DeepSeek V3; budget forecasts need upward adjustment.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Other price changes:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Z.ai GLM 5.2&lt;/strong&gt; – prompt ↓ $0.74 → $0.70 / 1M (-$0.05); completion ↓ $2.34 → $2.20 / 1M (-$0.14). Total –$0.19 / 1M.&lt;br&gt;&lt;br&gt;
&lt;em&gt;Who should care:&lt;/em&gt; Users of GLM 5.2 seeing modest savings.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;MoonshotAI Kimi Latest&lt;/strong&gt; – prompt ↓ $3.00 → $2.90 / 1M (-$0.10); completion unchanged at $15.00 / 1M. Total –$0.10 / 1M.&lt;br&gt;&lt;br&gt;
&lt;em&gt;Who should care:&lt;/em&gt; Cost‑sensitive prompt‑heavy applications.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Google Gemma 4 31B&lt;/strong&gt; – prompt ↓ $0.14 → $0.10 / 1M (-$0.04); completion ↓ $0.40 → $0.34 / 1M (-$0.06). Total –$0.10 / 1M.&lt;br&gt;&lt;br&gt;
&lt;em&gt;Who should care:&lt;/em&gt; Developers using Gemma 4 for latency‑critical tasks.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Qwen Qwen3 Coder Next&lt;/strong&gt; – prompt ↓ $0.18 → $0.12 / 1M (-$0.06); completion ↓ $0.90 → $0.80 / 1M (-$0.10). Total –$0.16 / 1M.&lt;br&gt;&lt;br&gt;
&lt;em&gt;Who should care:&lt;/em&gt; Coding assistants benefiting from lower token cost.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Qwen Qwen3 VL 30B A3B Instruct&lt;/strong&gt; – prompt ↑ $0.13 → $0.15 / 1M (+$0.02); completion ↑ $0.52 → $0.60 / 1M (+$0.08). Total +$0.10 / 1M.&lt;br&gt;&lt;br&gt;
&lt;em&gt;Who should care:&lt;/em&gt; Vision‑language workloads seeing a slight cost rise.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Qwen Qwen2.5 7B Instruct&lt;/strong&gt; – prompt ↑ $0.04 → $0.10 / 1M (+$0.06); completion ↑ $0.10 → $0.20 / 1M (+$0.10). Total +$0.16 / 1M.&lt;br&gt;&lt;br&gt;
&lt;em&gt;Who should care:&lt;/em&gt; Light‑weight inference services now noticeably more expensive.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;No models were added or removed today.&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
&lt;em&gt;Cheapest models available:&lt;/em&gt; inclusionAI Ling‑2.6‑flash ($0.01 / 1M prompt, $0.03 / 1M completion), IBM Granite 4.0 Micro ($0.017 / 1M prompt, $0.112 / 1M completion), Mistral Nemo ($0.019 / 1M prompt, $0.03 / 1M completion).  &lt;/p&gt;

&lt;p&gt;Total models tracked: 367.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://4663437Mehdi.github.io/token-ledger/entry.html?d=2026-07-30" rel="noopener noreferrer"&gt;The Token Ledger&lt;/a&gt;. Subscribe for the daily digest.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>api</category>
      <category>news</category>
    </item>
    <item>
      <title>Token Ledger Digest – 2026-07-29</title>
      <dc:creator>4663437Mehdi</dc:creator>
      <pubDate>Wed, 29 Jul 2026 09:48:30 +0000</pubDate>
      <link>https://dev.to/4663437mehdi/token-ledger-digest-2026-07-29-22hc</link>
      <guid>https://dev.to/4663437mehdi/token-ledger-digest-2026-07-29-22hc</guid>
      <description>&lt;h1&gt;
  
  
  Token Ledger Digest – 2026-07-29
&lt;/h1&gt;

&lt;p&gt;The biggest cost impact today is a price cut for NVIDIA Nemotron 3 Ultra, dropping completion from $3.60 to $2.20 per 1M tokens (‑$1.40/1M).&lt;/p&gt;

&lt;h2&gt;
  
  
  Added (28 models)
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Google Gemini 3.6 Flash (batch): added; prompt $0.75/1M, completion $3.75/1M; teams needing batch inference.
&lt;/li&gt;
&lt;li&gt;Google Gemini 3.5 Flash Lite (batch): added; prompt $0.15/1M, completion $1.25/1M; teams needing batch inference.
&lt;/li&gt;
&lt;li&gt;Anthropic Claude Sonnet 5 (batch): added; prompt $1.00/1M, completion $5.00/1M; teams needing batch inference.
&lt;/li&gt;
&lt;li&gt;Anthropic Claude Fable 5 (batch): added; prompt $5.00/1M, completion $25.00/1M; teams needing batch inference.
&lt;/li&gt;
&lt;li&gt;MiniMax MiniMax M3 (batch): added; prompt $0.15/1M, completion $0.60/1M; teams needing batch inference.
&lt;/li&gt;
&lt;li&gt;Anthropic Claude Opus 4.8 (batch): added; prompt $2.50/1M, completion $12.50/1M; teams needing batch inference.
&lt;/li&gt;
&lt;li&gt;Google Gemini 3.5 Flash (batch): added; prompt $0.75/1M, completion $4.50/1M; teams needing batch inference.
&lt;/li&gt;
&lt;li&gt;Google Gemini 3.1 Flash Lite (batch): added; prompt $0.125/1M, completion $0.75/1M; teams needing batch inference.
&lt;/li&gt;
&lt;li&gt;OpenAI GPT-5.5 (batch): added; prompt $2.50/1M, completion $15.00/1M; teams needing batch inference.
&lt;/li&gt;
&lt;li&gt;Anthropic Claude Opus 4.7 (batch): added; prompt $2.50/1M, completion $12.50/1M; teams needing batch inference.
&lt;/li&gt;
&lt;li&gt;OpenAI GPT-5.4 Nano (batch): added; prompt $0.10/1M, completion $0.625/1M; teams needing batch inference.
&lt;/li&gt;
&lt;li&gt;OpenAI GPT-5.4 Mini (batch): added; prompt $0.375/1M, completion $2.25/1M; teams needing batch inference.
&lt;/li&gt;
&lt;li&gt;OpenAI GPT-5.4 (batch): added; prompt $1.25/1M, completion $7.50/1M; teams needing batch inference.
&lt;/li&gt;
&lt;li&gt;Google Gemini 3.1 Pro Preview (batch): added; prompt $1.00/1M, completion $6.00/1M; teams needing batch inference.
&lt;/li&gt;
&lt;li&gt;Anthropic Claude Opus 4.6 (batch): added; prompt $2.50/1M, completion $12.50/1M; teams needing batch inference.
&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Removed (2 models)
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Poolside Laguna M.1: removed; prompt $0.20/1M, completion $0.40/1M; users of this model.
&lt;/li&gt;
&lt;li&gt;Poolside Laguna M.1 (free): removed; prompt $0.00/1M, completion $0.00/1M; users of the free tier.
&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Price Changes (13 models)
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Z.ai GLM 5.2: prompt ↓$0.7686→$0.7448/1M, completion ↓$2.4156→$2.3408/1M; developers using this model.
&lt;/li&gt;
&lt;li&gt;NVIDIA Nemotron 3 Ultra: prompt ↓$0.60→$0.50/1M, completion ↓$3.60→$2.20/1M; developers using this model.
&lt;/li&gt;
&lt;li&gt;Qwen Qwen3.6 Max Preview: prompt ↓$1.04→$1.027/1M, completion ↓$6.24→$6.162/1M; developers using this model.
&lt;/li&gt;
&lt;li&gt;Google Gemma 4 26B A4B: prompt ↓$0.14→$0.07/1M, completion ↓$0.42→$0.34/1M; developers using this model.
&lt;/li&gt;
&lt;li&gt;Qwen Qwen3 Coder Next: prompt ↑$0.11→$0.18/1M, completion ↑$0.80→$0.90/1M; developers using this model.
&lt;/li&gt;
&lt;li&gt;Qwen Qwen3 VL 8B Thinking: prompt ↑$0.117→$0.18/1M, completion ↑$1.365→$2.10/1M; developers using this model.
&lt;/li&gt;
&lt;li&gt;Qwen Qwen3 VL 30B A3B Thinking: prompt ↑$0.13→$0.20/1M, completion ↑$1.56→$2.40/1M; developers using this model.
&lt;/li&gt;
&lt;li&gt;Qwen Qwen3 VL 30B A3B Instruct: prompt ↓$0.15→$0.13/1M, completion ↓$0.60→$0.52/1M; developers using this model.
&lt;/li&gt;
&lt;li&gt;Qwen Qwen3 VL 235B A22B Thinking: prompt ↑$0.26→$0.40/1M, completion ↑$2.60→$4.00/1M; developers using this model.
&lt;/li&gt;
&lt;li&gt;Qwen Qwen3 Next 80B A3B Thinking: prompt ↑$0.0975→$0.15/1M, completion ↑$0.78→$1.20/1M; developers using this model.
&lt;/li&gt;
&lt;li&gt;Qwen Qwen Plus 0728 (thinking): prompt ↑$0.26→$0.40/1M, completion ↑$0.78→$1.20/1M; developers using this model.
&lt;/li&gt;
&lt;li&gt;Qwen Qwen3 30B A3B Thinking 2507: prompt ↑$0.13→$0.20/1M, completion ↑$1.56→$2.40/1M; developers using this model.
&lt;/li&gt;
&lt;li&gt;OpenAI gpt-oss-20b: prompt unchanged $0.03/1M, completion ↓$0.14→$0.13/1M; developers using this model.
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;Cheapest models today:&lt;/em&gt; inclusionAI Ling-2.6-flash ($0.01/1M prompt, $0.03/1M completion), IBM Granite 4.0 Micro ($0.017/1M prompt, $0.112/1M completion), Mistral Nemo ($0.019/1M prompt, $0.03/1M completion).&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://4663437Mehdi.github.io/token-ledger/entry.html?d=2026-07-29" rel="noopener noreferrer"&gt;The Token Ledger&lt;/a&gt;. Subscribe for the daily digest.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>api</category>
      <category>news</category>
    </item>
    <item>
      <title>AI Model Pricing Digest – 2026-07-28</title>
      <dc:creator>4663437Mehdi</dc:creator>
      <pubDate>Tue, 28 Jul 2026 09:46:27 +0000</pubDate>
      <link>https://dev.to/4663437mehdi/ai-model-pricing-digest-2026-07-28-111</link>
      <guid>https://dev.to/4663437mehdi/ai-model-pricing-digest-2026-07-28-111</guid>
      <description>&lt;h1&gt;
  
  
  AI Model Pricing Digest – 2026-07-28
&lt;/h1&gt;

&lt;p&gt;&lt;strong&gt;Most cost‑impacting change:&lt;/strong&gt; OpenAI cut prices for the GPT‑5.6 Terra series, slashing both prompt and completion costs by ~40‑50%.&lt;/p&gt;

&lt;h2&gt;
  
  
  Added
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Qwen: Qwen3.7 Flash&lt;/strong&gt; – New model. Prompt $0.03/1M, completion $0.13/1M. &lt;em&gt;Who should care:&lt;/em&gt; Teams needing ultra‑long context (1M tokens) at low cost.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Removed
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;OpenAI: GPT-5 Chat&lt;/strong&gt; – Deleted. Was $1.25/1M prompt, $10/1M completion. &lt;em&gt;Who should care:&lt;/em&gt; Users of this model must migrate to alternatives.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OpenAI: GPT-4o Search Preview&lt;/strong&gt; – Deleted. Was $2.5/1M prompt, $10/1M completion. &lt;em&gt;Who should care:&lt;/em&gt; Anyone relying on search‑augmented preview should switch.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Price Changes
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;OpenAI: GPT-5.6 Luna Pro&lt;/strong&gt; – Prompt ↓ $1.00 → $0.50/1M (-$0.50); Completion ↓ $6.00 → $3.00/1M (-$3.00). &lt;em&gt;Who should care:&lt;/em&gt; Cost‑sensitive workloads using Luna Pro.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OpenAI: GPT-5.6 Luna&lt;/strong&gt; – Same cuts as Luna Pro. &lt;em&gt;Who should care:&lt;/em&gt; Luna users.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OpenAI: GPT-5.6 Terra Pro&lt;/strong&gt; – Prompt ↓ $2.50 → $1.25/1M (-$1.25); Completion ↓ $15.00 → $7.50/1M (-$7.50). &lt;em&gt;Who should care:&lt;/em&gt; High‑volume Terra Pro users see ~40% savings.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OpenAI: GPT-5.6 Terra&lt;/strong&gt; – Identical cuts to Terra Pro. &lt;em&gt;Who should care:&lt;/em&gt; Terra users.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Z.ai: GLM 5.2&lt;/strong&gt; – Prompt ↓ $0.8036 → $0.7686/1M (-$0.035); Completion ↓ $2.5256 → $2.4156/1M (-$0.11). &lt;em&gt;Who should care:&lt;/em&gt; Minor savings for GLM 5.2 adopters.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;NVIDIA: Nemotron 3 Ultra&lt;/strong&gt; – Prompt ↑ $0.50 → $0.60/1M (+$0.10); Completion ↑ $2.20 → $3.60/1M (+$1.40). &lt;em&gt;Who should care:&lt;/em&gt; Budget impact for Nemotron 3 Ultra users.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Google: Gemma 4 26B A4B&lt;/strong&gt; – Prompt ↑ $0.12 → $0.14/1M (+$0.02); Completion ↑ $0.35 → $0.42/1M (+$0.07). &lt;em&gt;Who should care:&lt;/em&gt; Slight cost rise for Gemma 4 users.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Qwen: Qwen3.5-35B-A3B&lt;/strong&gt; – Prompt ↓ $0.15 → $0.14/1M (-$0.01); Completion unchanged $1.00/1M. &lt;em&gt;Who should care:&lt;/em&gt; Negligible saving for this model.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;Total models tracked: 341.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://4663437Mehdi.github.io/token-ledger/entry.html?d=2026-07-28" rel="noopener noreferrer"&gt;The Token Ledger&lt;/a&gt;. Subscribe for the daily digest.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>api</category>
      <category>news</category>
    </item>
    <item>
      <title>Token Ledger Digest – 2026-07-27</title>
      <dc:creator>4663437Mehdi</dc:creator>
      <pubDate>Mon, 27 Jul 2026 10:50:17 +0000</pubDate>
      <link>https://dev.to/4663437mehdi/token-ledger-digest-2026-07-27-1hjo</link>
      <guid>https://dev.to/4663437mehdi/token-ledger-digest-2026-07-27-1hjo</guid>
      <description>&lt;h1&gt;
  
  
  Token Ledger Digest – 2026-07-27
&lt;/h1&gt;

&lt;p&gt;&lt;strong&gt;Most cost‑impacting change&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;NVIDIA: Nemotron 3 Ultra&lt;/strong&gt; – completion price dropped from &lt;strong&gt;$3.60&lt;/strong&gt; to &lt;strong&gt;$2.20&lt;/strong&gt; per 1M tokens; prompt price fell from &lt;strong&gt;$0.60&lt;/strong&gt; to &lt;strong&gt;$0.50&lt;/strong&gt; per 1M tokens.
&lt;em&gt;Who should care:&lt;/em&gt; Existing users see a ~39% reduction in generation cost; new adopters benefit from lower pricing.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Other price changes&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Z.ai: GLM 5.2&lt;/strong&gt; – prompt rose from &lt;strong&gt;$0.68&lt;/strong&gt; to &lt;strong&gt;$0.80&lt;/strong&gt; per 1M; completion rose from &lt;strong&gt;$2.15&lt;/strong&gt; to &lt;strong&gt;$2.53&lt;/strong&gt; per 1M.
&lt;em&gt;Who should care:&lt;/em&gt; Sizable cost increase for workloads using GLM 5.2.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MoonshotAI: Kimi K2.7 Code&lt;/strong&gt; – prompt fell from &lt;strong&gt;$0.75&lt;/strong&gt; to &lt;strong&gt;$0.73&lt;/strong&gt; per 1M; completion unchanged at &lt;strong&gt;$3.50&lt;/strong&gt; per 1M.
&lt;em&gt;Who should care:&lt;/em&gt; Minor prompt‑side savings.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Qwen: Qwen3.5-35B-A3B&lt;/strong&gt; – prompt rose from &lt;strong&gt;$0.14&lt;/strong&gt; to &lt;strong&gt;$0.15&lt;/strong&gt; per 1M; completion unchanged at &lt;strong&gt;$1.00&lt;/strong&gt; per 1M.
&lt;em&gt;Who should care:&lt;/em&gt; Negligible impact.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Removed models&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;OpenAI: GPT-4o-mini Search Preview&lt;/strong&gt; – prompt &lt;strong&gt;$0.15&lt;/strong&gt;/1M, completion &lt;strong&gt;$0.60&lt;/strong&gt;/1M.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Inflection: Inflection 3 Pi&lt;/strong&gt; – prompt &lt;strong&gt;$2.50&lt;/strong&gt;/1M, completion &lt;strong&gt;$10.00&lt;/strong&gt;/1M.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Inflection: Inflection 3 Productivity&lt;/strong&gt; – same pricing as Pi.
&lt;em&gt;Who should care:&lt;/em&gt; Anyone relying on these models must migrate to alternatives; they are no longer available for new calls.
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Total models tracked: &lt;strong&gt;342&lt;/strong&gt;.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://4663437Mehdi.github.io/token-ledger/entry.html?d=2026-07-27" rel="noopener noreferrer"&gt;The Token Ledger&lt;/a&gt;. Subscribe for the daily digest.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>api</category>
      <category>news</category>
    </item>
    <item>
      <title>The Token Ledger – 2026-07-26</title>
      <dc:creator>4663437Mehdi</dc:creator>
      <pubDate>Sun, 26 Jul 2026 09:22:29 +0000</pubDate>
      <link>https://dev.to/4663437mehdi/the-token-ledger-2026-07-26-bnd</link>
      <guid>https://dev.to/4663437mehdi/the-token-ledger-2026-07-26-bnd</guid>
      <description>&lt;h1&gt;
  
  
  The Token Ledger – 2026-07-26
&lt;/h1&gt;

&lt;p&gt;&lt;strong&gt;Most cost‑impacting change:&lt;/strong&gt; Qwen’s Qwen3.6 27B completion price fell from $2.40 to $2.00 per 1M tokens (‑$0.40/M), a 16.7% reduction.&lt;/p&gt;

&lt;h3&gt;
  
  
  Price changes
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Z.ai: GLM 5.2&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Prompt: $0.749 → $0.685 /M
&lt;/li&gt;
&lt;li&gt;Completion: $2.354 → $2.152 /M
&lt;/li&gt;
&lt;li&gt;
&lt;em&gt;Who should care:&lt;/em&gt; Teams running latency‑critical inference where both prompt and completion costs matter.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;MoonshotAI: Kimi K2.7 Code&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Prompt: $0.78 → $0.75 /M
&lt;/li&gt;
&lt;li&gt;Completion: unchanged at $3.50 /M
&lt;/li&gt;
&lt;li&gt;
&lt;em&gt;Who should care:&lt;/em&gt; Code‑generation workflows sensitive to prompt pricing.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Qwen: Qwen3.6 27B&lt;/strong&gt; &lt;em&gt;(lead change)&lt;/em&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Prompt: $0.289 → $0.300 /M (+$0.011/M)
&lt;/li&gt;
&lt;li&gt;Completion: $2.400 → $2.000 /M (‑$0.400/M)
&lt;/li&gt;
&lt;li&gt;
&lt;em&gt;Who should care:&lt;/em&gt; Applications generating long completions (e.g., chat, summarization) that benefit from lower output cost.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;DeepSeek: DeepSeek V4 Flash&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Prompt: $0.094 → $0.14 /M (+$0.046/M)
&lt;/li&gt;
&lt;li&gt;Completion: $0.188 → $0.28 /M (+$0.092/M)
&lt;/li&gt;
&lt;li&gt;
&lt;em&gt;Who should care:&lt;/em&gt; Cost‑conscious users of flash‑mode models; higher per‑token cost may shift usage to alternatives.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;OpenAI: gpt‑oss‑20b&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Prompt: unchanged at $0.030 /M
&lt;/li&gt;
&lt;li&gt;Completion: $0.130 → $0.140 /M (+$0.010/M)
&lt;/li&gt;
&lt;li&gt;
&lt;em&gt;Who should care:&lt;/em&gt; Light‑weight completion tasks where a small cost increase may affect budgeting.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Qwen: Qwen3 30B A3B Instruct 2507&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Prompt: $0.100 → $0.048 /M (‑$0.052/M)
&lt;/li&gt;
&lt;li&gt;Completion: $0.300 → $0.193 /M (‑$0.107/M)
&lt;/li&gt;
&lt;li&gt;
&lt;em&gt;Who should care:&lt;/em&gt; Instruction‑following pipelines seeing notable savings on both input and output.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Qwen: Qwen3 30B A3B&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Prompt: $0.130 → $0.120 /M (‑$0.010/M)
&lt;/li&gt;
&lt;li&gt;Completion: $0.520 → $0.500 /M (‑$0.020/M)
&lt;/li&gt;
&lt;li&gt;
&lt;em&gt;Who should care:&lt;/em&gt; General‑purpose deployments where modest per‑token reductions accumulate over high volume.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;No models were added or removed today.&lt;/em&gt;  &lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Three cheapest models (per‑token pricing):&lt;/strong&gt;  &lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;inclusionAI: Ling‑2.6‑flash – Prompt $0.00001, Completion $0.00003 /M
&lt;/li&gt;
&lt;li&gt;IBM: Granite 4.0 Micro – Prompt $0.000017, Completion $0.000112 /M
&lt;/li&gt;
&lt;li&gt;Mistral: Mistral Nemo – Prompt $0.000019, Completion $0.00003 /M&lt;/li&gt;
&lt;/ol&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://4663437Mehdi.github.io/token-ledger/entry.html?d=2026-07-26" rel="noopener noreferrer"&gt;The Token Ledger&lt;/a&gt;. Subscribe for the daily digest.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>api</category>
      <category>news</category>
    </item>
  </channel>
</rss>
