DEV Community

4663437Mehdi
4663437Mehdi

Posted on • Originally published at 4663437mehdi.github.io

The Token Ledger Digest – 2026-08-09

The Token Ledger Digest – 2026-08-09

The most cost‑impacting change today is a steep price reduction for Z.ai: GLM 5.2, cutting both prompt and completion costs by over 80 %.

Z.ai: GLM 5.2

  • Prompt: fell from $0.266 / 1M to $0.07 / 1M (‑73.7 %).
  • Completion: fell from $0.836 / 1M to $0.22 / 1M (‑73.8 %). Who should care: Teams running high‑volume chat, summarization, or code‑gen workloads where token usage dominates cost.

DeepSeek V4 Flash Latest

  • Prompt: down slightly from $0.090 / 1M to $0.080 / 1M.
  • Completion: up from $0.180 / 1M to $0.252 / 1M (+40 %). Who should care: Applications that are completion‑heavy (e.g., creative writing) may see higher bills.

MoonshotAI Kimi Latest

  • Prompt: rose from $2.50 / 1M to $2.80 / 1M (+12 %).
  • Completion: unchanged at $14.00 / 1M. Who should care: Users sensitive to prompt cost, such as those feeding large context windows.

Qwen: Qwen3.6 35B A3B

  • Prompt: increased from $0.14 / 1M to $0.15 / 1M (+7 %).
  • Completion: steady at $1.00 / 1M. Who should care: Cost‑conscious developers using this model for moderate‑size generations.

MiniMax: MiniMax M2.7

  • Prompt: up from $0.27 / 1M to $0.30 / 1M (+11 %).
  • Completion: up from $1.08 / 1M to $1.20 / 1M (+11 %). Who should care: Balanced prompt/completion workloads will see a modest rise.

NVIDIA: Nemotron 3 Super

  • Prompt: jumped from $0.085 / 1M to $0.30 / 1M (+253 %).
  • Completion: rose from $0.40 / 1M to $0.90 / 1M (+125 %). Who should care: Anyone relying on this large model for inference should reassess budget or consider alternatives.

DeepSeek: DeepSeek V3.2

  • Prompt: slight increase from $0.26 / 1M to $0.269 / 1M (+3.5 %).
  • Completion: up from $0.38 / 1M to $0.40 / 1M (+5.3 %). Who should care: Minor impact; relevant for long‑running services tracking marginal cost drift.

No models were added or removed today.

Three cheapest models (per‑token):

  1. inclusionai/ling-2.6-flash – $0.01 / 1M prompt, $0.03 / 1M completion
  2. mistralai/mistral-nemo – $0.019 / 1M prompt, $0.03 / 1M completion
  3. ibm-granite/granite-4.0-h-micro – $0.017 / 1M prompt, $0.112 / 1M completion

Originally published at The Token Ledger. Subscribe for the daily digest.

Top comments (0)