DEV Community

4663437Mehdi
4663437Mehdi

Posted on Originally published at 4663437mehdi.github.io

The Token Ledger Digest – 2026-09-02

The Token Ledger Digest – 2026-09-02

Most cost‑impacting change: DeepSeek: DeepSeek V4 Pro 0813 – prompt price rose from $0.66 to $1.12/1M tokens (+$0.46) and completion from $1.98 to $3.35/1M (+$1.37). Developers using this model for high‑volume generation will see inference costs jump ~92%.

Added models

  • Anthropic: Claude Fable 5.1 – prompt $10.00/1M, completion $50.00/1M. Teams needing long‑context (1M tokens) with strong reasoning should evaluate.
  • Anthropic: Claude Fable 5.1 (batch) – prompt $5.00/1M, completion $25.00/1M. Batch users can halve costs vs. synchronous version.
  • Inception: Mercury 2.5 Preview – prompt $0.04/1M, completion $0.15/1M. Ultra‑cheap for short‑form tasks; attractive for prototyping.
  • Z.ai: GLM Flash Latest – prompt $0.075/1M, completion $0.25/1M. Good balance of cost and 1.3M‑token context for retrieval‑augmented workflows.

Removed models

  • Claude Opus 5 (Fast) – prompt $10.00/1M, completion $50.00/1M. Users must migrate to Claude Fable 5.1 or alternatives.
  • Anthropic: Claude Opus 4.8 (Fast) – same pricing as Opus 5 Fast; same migration impact.

Price changes

  • Z.ai: GLM Latest – prompt $1.15/1M (down $0.02), completion $3.50/1M (down $0.46). Cost‑sensitive apps benefit slightly.
  • Qwen: Qwen3.8 2.4T A95B (batch) – prompt $2.00/1M (down $0.50), completion $6.00/1M (down $0.25). Batch users see ~12% lower spend.
  • DeepSeek: DeepSeek V4 Pro 0813 – prompt $1.12/1M (+$0.46), completion $3.35/1M (+$1.37). Notable cost increase; reconsider usage.
  • Z.ai: GLM 5.2 – prompt $0.97/1M (down $0.22), completion $3.04/1M (down $0.70). Moderate savings for existing users.
  • NVIDIA: Nemotron 3 Ultra – prompt $0.63/1M (+$0.13), completion $3.13/1M (+$0.93). Higher cost for those needing large model capacity.
  • DeepSeek: DeepSeek V4 Pro 0423 – prompt $1.04/1M (down $0.56), completion $2.08/1M (down $1.12). Significant reduction; attractive for cost‑conscious deployments.
  • DeepSeek: DeepSeek V4 Flash 0423 – prompt $0.089/1M (+$0.008), completion $0.177/1M (+$0.015). Minimal impact.
  • DeepSeek: DeepSeek V3.1 – prompt $0.25/1M (down $0.30), completion $0.95/1M (down $0.70). Large savings for chat‑oriented workloads.

Cheapest models today (per‑million)

  1. IBM: Granite 4.0 Micro – prompt $0.017/1M, completion $0.112/1M
  2. Mistral: Mistral Nemo – prompt $0.019/1M, completion $0.030/1M
  3. Ling-3.0-flash – prompt $0.021/1M, completion $0.063/1M

Total models tracked: 421.


Originally published at The Token Ledger. Subscribe for the daily digest.

Top comments (0)