Token Ledger Digest – 2026-07-29
The biggest cost impact today is a price cut for NVIDIA Nemotron 3 Ultra, dropping completion from $3.60 to $2.20 per 1M tokens (‑$1.40/1M).
Added (28 models)
- Google Gemini 3.6 Flash (batch): added; prompt $0.75/1M, completion $3.75/1M; teams needing batch inference.
- Google Gemini 3.5 Flash Lite (batch): added; prompt $0.15/1M, completion $1.25/1M; teams needing batch inference.
- Anthropic Claude Sonnet 5 (batch): added; prompt $1.00/1M, completion $5.00/1M; teams needing batch inference.
- Anthropic Claude Fable 5 (batch): added; prompt $5.00/1M, completion $25.00/1M; teams needing batch inference.
- MiniMax MiniMax M3 (batch): added; prompt $0.15/1M, completion $0.60/1M; teams needing batch inference.
- Anthropic Claude Opus 4.8 (batch): added; prompt $2.50/1M, completion $12.50/1M; teams needing batch inference.
- Google Gemini 3.5 Flash (batch): added; prompt $0.75/1M, completion $4.50/1M; teams needing batch inference.
- Google Gemini 3.1 Flash Lite (batch): added; prompt $0.125/1M, completion $0.75/1M; teams needing batch inference.
- OpenAI GPT-5.5 (batch): added; prompt $2.50/1M, completion $15.00/1M; teams needing batch inference.
- Anthropic Claude Opus 4.7 (batch): added; prompt $2.50/1M, completion $12.50/1M; teams needing batch inference.
- OpenAI GPT-5.4 Nano (batch): added; prompt $0.10/1M, completion $0.625/1M; teams needing batch inference.
- OpenAI GPT-5.4 Mini (batch): added; prompt $0.375/1M, completion $2.25/1M; teams needing batch inference.
- OpenAI GPT-5.4 (batch): added; prompt $1.25/1M, completion $7.50/1M; teams needing batch inference.
- Google Gemini 3.1 Pro Preview (batch): added; prompt $1.00/1M, completion $6.00/1M; teams needing batch inference.
- Anthropic Claude Opus 4.6 (batch): added; prompt $2.50/1M, completion $12.50/1M; teams needing batch inference.
Removed (2 models)
- Poolside Laguna M.1: removed; prompt $0.20/1M, completion $0.40/1M; users of this model.
- Poolside Laguna M.1 (free): removed; prompt $0.00/1M, completion $0.00/1M; users of the free tier.
Price Changes (13 models)
- Z.ai GLM 5.2: prompt ↓$0.7686→$0.7448/1M, completion ↓$2.4156→$2.3408/1M; developers using this model.
- NVIDIA Nemotron 3 Ultra: prompt ↓$0.60→$0.50/1M, completion ↓$3.60→$2.20/1M; developers using this model.
- Qwen Qwen3.6 Max Preview: prompt ↓$1.04→$1.027/1M, completion ↓$6.24→$6.162/1M; developers using this model.
- Google Gemma 4 26B A4B: prompt ↓$0.14→$0.07/1M, completion ↓$0.42→$0.34/1M; developers using this model.
- Qwen Qwen3 Coder Next: prompt ↑$0.11→$0.18/1M, completion ↑$0.80→$0.90/1M; developers using this model.
- Qwen Qwen3 VL 8B Thinking: prompt ↑$0.117→$0.18/1M, completion ↑$1.365→$2.10/1M; developers using this model.
- Qwen Qwen3 VL 30B A3B Thinking: prompt ↑$0.13→$0.20/1M, completion ↑$1.56→$2.40/1M; developers using this model.
- Qwen Qwen3 VL 30B A3B Instruct: prompt ↓$0.15→$0.13/1M, completion ↓$0.60→$0.52/1M; developers using this model.
- Qwen Qwen3 VL 235B A22B Thinking: prompt ↑$0.26→$0.40/1M, completion ↑$2.60→$4.00/1M; developers using this model.
- Qwen Qwen3 Next 80B A3B Thinking: prompt ↑$0.0975→$0.15/1M, completion ↑$0.78→$1.20/1M; developers using this model.
- Qwen Qwen Plus 0728 (thinking): prompt ↑$0.26→$0.40/1M, completion ↑$0.78→$1.20/1M; developers using this model.
- Qwen Qwen3 30B A3B Thinking 2507: prompt ↑$0.13→$0.20/1M, completion ↑$1.56→$2.40/1M; developers using this model.
- OpenAI gpt-oss-20b: prompt unchanged $0.03/1M, completion ↓$0.14→$0.13/1M; developers using this model.
Cheapest models today: inclusionAI Ling-2.6-flash ($0.01/1M prompt, $0.03/1M completion), IBM Granite 4.0 Micro ($0.017/1M prompt, $0.112/1M completion), Mistral Nemo ($0.019/1M prompt, $0.03/1M completion).
Originally published at The Token Ledger. Subscribe for the daily digest.
Top comments (0)