DEV Community

4663437Mehdi
4663437Mehdi

Posted on • Originally published at 4663437mehdi.github.io

The Token Ledger – 2026-08-07

The Token Ledger – 2026-08-07

Most cost‑impacting change: NVIDIA Nemotron 3 Super saw a price increase. Prompt rose from $0.085/1M to $0.30/1M (+$0.215/1M); completion rose from $0.40/1M to $0.90/1M (+$0.50/1M). Teams using this model for high‑volume inference should budget for roughly a 70% cost jump.

Other price changes

  • MoonshotAI Kimi Latest: Prompt dropped from $2.90/1M to $2.50/1M (−$0.40/1M). Completion unchanged at $14.00/1M. Cost‑sensitive users of this model save ~14% on prompt tokens.
  • Qwen VL 235B Thinking: Prompt fell from $0.98/1M to $0.40/1M (−$0.58/1M); completion rose slightly from $3.95/1M to $4.00/1M (+$0.05/1M). Net −$0.53/1M, benefiting vision‑language workloads.
  • Thinking Machines Inkling Small: Prompt decreased from $0.50/1M to $0.45/1M (−$0.005/1M); completion steady at $1.20/1M. Minor saving for low‑traffic use.
  • InclusionAI Ling‑3.0‑flash: Prompt cut from $0.075/1M to $0.021/1M (−$0.054/1M); completion cut from $0.22/1M to $0.063/1M (−$0.157/1M). Total −$0.211/1M – a notable reduction for existing users.
  • Z.ai GLM 5.2: Prompt lowered from $0.76/1M to $0.679/1M (−$0.081/1M); completion lowered from $2.42/1M to $2.134/1M (−$0.286/1M). Net −$0.367/1M.
  • Qwen 3.5‑122B: Prompt rose from $0.26/1M to $0.29/1M (+$0.03/1M); completion rose from $2.08/1M to $2.40/1M (+$0.32/1M). Net +$0.35/1M – a modest increase for those relying on this model.

Catalog updates: 61 new models were added, highlighted by the free inclusionai/ling-3.0-tiny, batch variants of Claude Opus 5, Gemini 3.6 Flash, and several OpenAI GPT‑5.6 Luna/Terra/Sol editions. One model was removed: the free inclusionai/ling-3.0-flash.

Developers should review the new batch offerings for potential latency‑cost tradeoffs and adjust usage of the adjusted models accordingly.


Originally published at The Token Ledger. Subscribe for the daily digest.

Top comments (0)