DEV Community

4663437Mehdi
4663437Mehdi

Posted on • Originally published at 4663437mehdi.github.io

Token Ledger Digest – 2026-07-31

Token Ledger Digest – 2026-07-31

Added

  • DeepSeek: DeepSeek V4 Flash 0731 – New model with 1M‑token context. Prompt $0.14/1M, completion $0.28/1M. Who should care: Teams needing long‑context generation at low cost.

Removed

  • OpenAI: o3 Deep Research – Prompt $10.00/1M, completion $40.00/1M.
  • OpenAI: o4 Mini Deep Research – Prompt $2.00/1M, completion $8.00/1M.
  • OpenAI: GPT-5 Codex – Prompt $1.25/1M, completion $10.00/1M. Who should care: Any workflows relying on these models must migrate to alternatives.

Price Changes

Model Old Prompt ($/1M) New Prompt ($/1M) Δ Prompt Old Compl. ($/1M) New Compl. ($/1M) Δ Compl. Who should care
Poolside: Laguna S 2.1 0.10 0.09 –0.01 0.20 0.18 –0.02 Cost‑sensitive inference on Laguna S.
OpenAI: GPT-5.6 Luna Pro 0.50 0.10 –0.40 3.00 0.60 –2.40 Heavy users of Luna Pro see ~80% cost cut.
OpenAI: GPT-5.6 Luna 0.50 0.10 –0.40 3.00 0.60 –2.40 Same impact as Luna Pro.
OpenAI: GPT-5.6 Terra Pro 1.25 1.00 –0.25 7.50 6.00 –1.50 Terra Pro workloads become cheaper.
OpenAI: GPT-5.6 Terra 1.25 1.00 –0.25 7.50 6.00 –1.50 Same as Terra Pro.
Z.ai: GLM 5.2 0.70 1.23 +0.53 2.20 3.87 +1.67 GLM 5.2 now significantly more expensive.
NVIDIA: Nemotron 3 Ultra 0.50 0.60 +0.10 2.20 3.60 +1.40 Ultra‑scale users face higher spend.
MoonshotAI Kimi Latest 2.90 2.90 0.00 15.00 14.00 –1.00 Slight completion‑cost relief.
MoonshotAI: Kimi K2.6 0.65 0.95 +0.30 2.72 4.00 +1.28 Kimi K2.6 cost rises noticeably.
Qwen: Qwen3 VL 30B A3B Instruct 0.15 0.13 –0.02 0.60 0.52 –0.08 Minor savings for vision‑language tasks.
Qwen: Qwen3 235B A22B Thinking 2507 0.30 0.23 –0.07 3.00 2.30 –0.70 Thinking model becomes cheaper.

Most cost‑impacting change: The OpenAI GPT-5.6 Luna Pro and Luna models dropped prompt pricing from $0.50 to $0.10/1M and completion from $3.00 to $0.60/1M – an ~80% reduction, saving up to $2.80 per million tokens.

Cheapest Models Today

  1. inclusionAI: Ling-2.6-flash – $0.01/1M prompt, $0.03/1M completion
  2. IBM: Granite 4.0 Micro – $0.017/1M prompt, $0.112/1M completion
  3. Mistral: Mistral Nemo – $0.019/1M prompt, $0.03/1M completion

Total models tracked: 365.


Originally published at The Token Ledger. Subscribe for the daily digest.

Top comments (0)