The Token Ledger Digest – 2026-08-27
Top cost impact: DeepSeek’s V4 Flash Vision Exp model saw its prompt price drop 50% and completion price halve, saving $0.88 per 1M tokens total.
Price changes
- DeepSeek V4 Flash Vision Exp – Prompt ↓$0.44→$0.22 /1M; Completion ↓$1.32→$0.66 /1M. Who: Vision‑heavy apps (OCR, image QA).
- DeepSeek V4 Flash Latest – Completion ↑$0.075→$0.10 /1M (prompt unchanged). Who: Low‑latency chat bots seeing a slight cost rise.
- DeepSeek V4 Flash 0731 – Prompt ↓$0.06→$0.05 /1M; Completion ↓$0.12→$0.10 /1M. Who: General‑purpose workloads on this snapshot.
- Tencent Hy3 – Prompt ↓$0.132→$0.0825 /1M; Completion ↓$0.528→$0.33 /1M. Who: Multimodal services using Hy3.
- MoonshotAI Kimi K2.7 Code – Prompt ↓$0.67→$0.66 /1M (completion flat). Who: Code‑generation pipelines.
- Qwen Qwen3.6 35B A3B – Prompt ↓$0.14→$0.10 /1M; Completion ↓$1.00→$0.90 /1M. Who: Mid‑size LLM users needing cheaper inference.
- DeepSeek V4 Flash 0423 – Prompt ↓$0.089→$0.078 /1M; Completion ↓$0.177→$0.156 /1M. Who: Users locked to the 0423 checkpoint.
- Z.ai GLM 4.6 – Prompt ↓$0.50→$0.43 /1M; Completion ↓$2.00→$1.75 /1M. Who: Enterprises running GLM‑4.6 at scale.
- Qwen Qwen3 235B A22B Instruct 2507 – Prompt ↓$0.09→$0.0875 /1M; Completion ↓$0.55→$0.35 /1M. Who: Large‑scale instruction‑following tasks.
Added models
- inclusionai/ling-3.0-flash-fin:free – $0 prompt, $0 completion (free). Who: Cost‑sensitive prototyping.
- qwen/qwen3.8-flash – Prompt $0.15 /1M, Completion $0.47 /1M, 1M‑token context. Who: Long‑context applications needing flash speed.
- z-ai/glm-5.3-flash – Prompt $0.075 /1M, Completion $0.25 /1M, 1.31M‑token context. Who: Users needing large context with low price.
Removed models
- stealth/ox-alpha – Free model removed. Who: Anyone relying on the zero‑cost Ox Alpha.
- z-ai/glm-5.2:batch – Prompt $1.40 /1M, Completion $4.40 /1M. Who: Batch‑job users must migrate.
- mistralai/ministral-8b – Prompt $0.11 /1M, Completion $0.11 /1M. Who: Low‑cost edge‑device developers.
No other meaningful shifts; total models now 417.
Three cheapest today: IBM Granite 4.0 Micro ($0.017 prompt / $0.112 completion per 1M), Mistral Nemo ($0.019 prompt / $0.030 completion per 1M), Ling‑3.0‑flash ($0.021 prompt / $0.063 completion per 1M).
Originally published at The Token Ledger. Subscribe for the daily digest.
Top comments (0)