The Token Ledger Digest – 2026-07-30
Biggest cost impact: DeepSeek DeepSeek V3 – prompt price rose from $0.20 to $0.26 / 1M tokens (+$0.06), completion from $0.80 to $1.03 / 1M tokens (+$0.23). Total increase ≈ $0.29 / 1M tokens.
Who should care: Teams running high‑volume chat or code‑gen workloads on DeepSeek V3; budget forecasts need upward adjustment.
Other price changes:
Z.ai GLM 5.2 – prompt ↓ $0.74 → $0.70 / 1M (-$0.05); completion ↓ $2.34 → $2.20 / 1M (-$0.14). Total –$0.19 / 1M.
Who should care: Users of GLM 5.2 seeing modest savings.MoonshotAI Kimi Latest – prompt ↓ $3.00 → $2.90 / 1M (-$0.10); completion unchanged at $15.00 / 1M. Total –$0.10 / 1M.
Who should care: Cost‑sensitive prompt‑heavy applications.Google Gemma 4 31B – prompt ↓ $0.14 → $0.10 / 1M (-$0.04); completion ↓ $0.40 → $0.34 / 1M (-$0.06). Total –$0.10 / 1M.
Who should care: Developers using Gemma 4 for latency‑critical tasks.Qwen Qwen3 Coder Next – prompt ↓ $0.18 → $0.12 / 1M (-$0.06); completion ↓ $0.90 → $0.80 / 1M (-$0.10). Total –$0.16 / 1M.
Who should care: Coding assistants benefiting from lower token cost.Qwen Qwen3 VL 30B A3B Instruct – prompt ↑ $0.13 → $0.15 / 1M (+$0.02); completion ↑ $0.52 → $0.60 / 1M (+$0.08). Total +$0.10 / 1M.
Who should care: Vision‑language workloads seeing a slight cost rise.Qwen Qwen2.5 7B Instruct – prompt ↑ $0.04 → $0.10 / 1M (+$0.06); completion ↑ $0.10 → $0.20 / 1M (+$0.10). Total +$0.16 / 1M.
Who should care: Light‑weight inference services now noticeably more expensive.
No models were added or removed today.
Cheapest models available: inclusionAI Ling‑2.6‑flash ($0.01 / 1M prompt, $0.03 / 1M completion), IBM Granite 4.0 Micro ($0.017 / 1M prompt, $0.112 / 1M completion), Mistral Nemo ($0.019 / 1M prompt, $0.03 / 1M completion).
Total models tracked: 367.
Originally published at The Token Ledger. Subscribe for the daily digest.
Top comments (0)