Token Ledger Digest – 2026-08-12
Price Changes (most cost‑impacting first)
Z.ai: GLM 5.2
- Prompt: $0.76 → $0.50 /1M tokens (‑$0.26)
- Completion: $2.42 → $3.15 /1M tokens (+$0.73) Who should care: Teams running high‑volume completion workloads will see per‑million costs rise ~30%; prompt‑heavy flows benefit from a 34% drop.
Qwen: Qwen3 VL 235B A22B Instruct
- Prompt: $0.21 → $0.26 /1M (+$0.05)
- Completion: $1.90 → $1.04 /1M (‑$0.86) Who should care: Vision‑language apps that generate long outputs will cut completion spend by ~45%; slight prompt increase is negligible.
DeepSeek: DeepSeek V3.1 Terminus
- Prompt: unchanged $0.27 /1M
- Completion: $1.00 → $0.95 /1M (‑$0.05) Who should care: Minor savings for completion‑heavy use cases.
OpenAI: gpt-oss-120b
- Prompt: $0.037 → $0.030 /1M (‑$0.007)
- Completion: unchanged $0.17 /1M Who should care: Developers using this model for prompt‑intensive tasks save <1% per million tokens.
Added Models
LiquidAI: LFM2.5-2.6B (free)
- Prompt/Completion: $0 /1M, context 128k Who should care: Prototyping or education projects needing zero‑cost inference.
NVIDIA: Nemotron 3.5 Lightning
- Prompt: $0.10 /1M, Completion: $0.25 /1M, context 262k Who should care: Users seeking a low‑cost, mid‑size model with decent throughput.
NVIDIA: Nemotron 3.5 Lightning (free)
- Prompt/Completion: $0 /1M, context 1M Who should care: Applications requiring very long context without cost.
ByteDance Seed: Seed-2.0-Code
- Prompt: $0.50 /1M, Completion: $3.00 /1M, context 262k Who should care: Code‑generation workloads where completion cost is acceptable for higher quality.
Total models tracked: 406.
Originally published at The Token Ledger. Subscribe for the daily digest.
Top comments (0)