Token Ledger Digest – 2026-08-08
Lead change – biggest cost impact:
Z.ai: GLM 5.2 – Prompt price fell from $0.679/1M to $0.266/1M (‑$0.413/1M); completion price fell from $2.134/1M to $0.836/1M (‑$1.298/1M).
Who should care: Teams running high‑volume inference or fine‑tuning on GLM 5.2 will see per‑token costs drop ~60 % for prompts and ~61 % for completions.
Other price changes
Thinking Machines: Inkling – Prompt: $1.00 → $0.95/1M (‑$0.05/1M). Completion unchanged at $4.05/1M.
Who should care: Users sensitive to prompt‑only costs (e.g., classification) save 5 % per million tokens.DeepSeek: DeepSeek V4 Flash 0423 – Prompt: $0.0882 → $0.14/1M (+$0.0518/1M). Completion: $0.1764 → $0.28/1M (+$0.1036/1M).
Who should care: Cost‑sensitive applications will see ~59 % higher prompt and ~59 % higher completion expenses.MoonshotAI: Kimi K2.6 – Prompt: $0.589 → $0.5795/1M (‑$0.0095/1M). Completion: $2.48 → $2.44/1M (‑$0.04/1M).
Who should care: Minor savings (~1.6 % prompt, ~1.6 % completion) for latency‑critical workloads.NVIDIA: Nemotron 3 Super – Prompt: $0.30 → $0.085/1M (‑$0.215/1M). Completion: $0.90 → $0.40/1M (‑$0.50/1M).
Who should care: Large‑scale generation jobs benefit from ~71 % lower prompt and ~56 % lower completion costs.DeepSeek: DeepSeek V3.2 – Prompt: $0.269 → $0.26/1M (‑$0.009/1M). Completion: $0.40 → $0.38/1M (‑$0.02/1M).
Who should care: Small reductions (~3 % prompt, ~5 % completion) for existing V3.2 users.
Cheapest models today (per‑token prices):
- inclusionAI: Ling-2.6-flash – Prompt $0.01/1M, Completion $0.03/1M
- IBM: Granite 4.0 Micro – Prompt $0.017/1M, Completion $0.112/1M
- Mistral: Mistral Nemo – Prompt $0.019/1M, Completion $0.03/1M
No models were added or removed today. Total model count remains 400.
Originally published at The Token Ledger. Subscribe for the daily digest.
Top comments (0)