Token Ledger Digest – 2026-09-09
Lead: DeepSeek V4 Pro 0813 (batch) delivers the biggest cost cut, lowering prompt price from $1.32 to $0.66 per 1M tokens and completion from $3.96 to $1.98 per 1M tokens – a total saving of $2.64 per 1M tokens. Teams running large‑scale batch workloads should revisit their cost forecasts.
Added Models
- Inception: Mercury 2.5 – new; prompt $0.04/1M, completion $0.15/1M, 260k context. Suitable for latency‑sensitive apps needing long context.
- Nex AGI: Nex‑N2.5‑Mini (free) – new; free. Ideal for prototyping or education.
- Nex AGI: Nex‑N2.5‑Pro (free) – new; free. Same use‑case as Mini.
- DeepSeek: DeepSeek V4 Flash Vision Exp (batch) – new; prompt $0.11/1M, completion $0.33/1M, 1M context. Vision‑focused batch pipelines.
- Z.ai: GLM 5.3 (batch) – new; prompt $0.70/1M, completion $2.20/1M, 1M context. High‑throughput batch jobs.
- Z.ai: GLM 5.2 (batch) – new; same pricing as GLM 5.3.
Removed Models
- Inception: Mercury 2.5 Preview – removed; was $0.04/1M prompt, $0.15/1M completion. Migrate to Mercury 2.5.
- Nex AGI: Nex‑N2‑Mini – removed; was $0.025/1M prompt, $0.10/1M completion. Replace with the free Mini or another low‑cost option.
- Nex AGI: Nex‑N2‑Pro – removed; was $0.25/1M prompt, $1.00/1M completion. Consider newer Pro variants or free alternatives.
Price Changes
- IBM: Granite 4.2 8B – prompt ↓ $0.10 → $0.06/1M; completion ↑ $0.15 → $0.25/1M. Cost‑sensitive inference workloads.
- Z.ai: GLM Flash Latest – prompt ↑ $0.071 → $0.075/1M; completion ↑ $0.238 → $0.250/1M. Minor uptick; monitor latency‑critical services.
- Z.ai: GLM 5.3 Flash (batch) – prompt ↓ $0.15 → $0.075/1M; completion ↓ $0.50 → $0.25/1M. Batch pipelines benefit from ~50% savings.
- Z.ai: GLM Latest – prompt ↓ $1.12 → $1.113/1M; completion ↓ $3.52 → $3.498/1M. Negligible impact; keep existing contracts.
- DeepSeek: DeepSeek V4 Pro 0813 (batch) – prompt ↓ $1.32 → $0.66/1M; completion ↓ $3.96 → $1.98/1M. Major saving for large batch jobs.
- Meta: Muse Glimmer 30B (batch) – prompt ↓ $0.35 → $0.175/1M; completion ↓ $1.50 → $0.75/1M. Substantial cut for creative‑generation batches.
- DeepSeek: DeepSeek V4 Flash 0731 (batch) – prompt ↓ $0.14 → $0.11/1M; completion ↑ $0.28 → $0.33/1M. Mixed effect; evaluate prompt‑heavy vs. completion‑heavy loads.
- MoonshotAI Kimi Latest – prompt ↓ $2.50 → $2.40/1M; completion ↓ $14.0 → $12.0/1M. Notable reduction for long‑form generation.
- Z.ai: GLM 4.6 – prompt ↓ $0.55 → $0.43/1M; completion ↓ $2.20 → $1.75/1M. Moderate saving for standard workloads.
- Qwen: Qwen3 Next 80B A3B Instruct – prompt ↓ $0.10 → $0.09/1M; completion unchanged $1.10/1M. Small prompt‑only gain.
- Qwen: Qwen3 235B A22B Instruct 2507 – prompt ↑ $0.09 → $0.22/1M; completion ↑ $0.55 → $0.88/1M. Cost increase; assess alternatives.
- MiniMax: MiniMax M1 – prompt ↑ $0.40 → $0.55/1M; completion unchanged $2.20/1M. Slight uplift; watch budget‑sensitive deployments.
If no meaningful changes had occurred, we would note that and list the three cheapest models today.
Originally published at The Token Ledger. Subscribe for the daily digest.
Top comments (0)