The Token Ledger Digest – 2026-08-28
Lead: Google Gemini 3.7 Flash and Gemini Flash Latest doubled both prompt and completion rates.
Google: Gemini 3.7 Flash – prompt rose from $0.375/M to $0.75/M (+$0.375/M); completion from $1.875/M to $3.75/M (+$1.875/M).
Who should care: Teams using Gemini Flash for high‑volume inference; cost per million tokens jumps ~2.25×.Google Gemini Flash Latest – identical change (prompt $0.375→$0.75/M, completion $1.875→$3.75/M).
Who should care: Same as above; affects any workload pinned to the “latest” tag.
Other notable price shifts
DeepSeek V4 Pro 0813 – prompt fell $1.122/M → $0.66/M (−$0.462/M); completion $3.366/M → $1.98/M (−$1.386/M).
Who should care: Cost‑sensitive users of DeepSeek’s larger V4 model see ~45% lower prompt and ~41% lower completion cost.NVIDIA Nemotron 3 Ultra – prompt $0.60/M → $0.50/M (−$0.10/M); completion $3.60/M → $2.20/M (−$1.40/M).
Who should care: Users of Nemotron Ultra benefit from a 61% drop in completion cost.
Additions & removals
- 8 new models added, highlighted by Tencent Hy4 preview (prompt $0.834/M, completion $2.501/M) and several Mistral batch variants ranging from $0.15/M prompt to $0.75/M completion.
- 37 models removed, chiefly the entire OpenAI GPT‑5.6 Luna/Terra/Sol families and GPT‑5.5/5.4 series, eliminating high‑priced options (e.g., GPT‑5.5 Pro at $15/M prompt, $90/M completion).
Cheapest models today (per‑million): IBM Granite 4.0 Micro ($0.017 prompt, $0.112 completion), Mistral Nemo ($0.019 prompt, $0.03 completion), Ling‑3.0‑flash ($0.021 prompt, $0.063 completion).
Total models tracked: 388.
Originally published at The Token Ledger. Subscribe for the daily digest.
Top comments (0)