DEV Community

4663437Mehdi
4663437Mehdi

Posted on • Originally published at 4663437mehdi.github.io

Token Ledger Digest – 2026-08-15

Token Ledger Digest – 2026-08-15

Most cost‑impacting change:

  • Z.ai: GLM 5.1 – Prompt price fell from $1.40 / 1M to $0.966 / 1M (−$0.434); completion price fell from $4.40 / 1M to $3.036 / 1M (−$1.364). Total cost ↓ $1.798 / 1M. Who should care: Teams running high‑volume inference on GLM 5.1 will see ~40% lower per‑million‑token spend.

Added Models

  • Qwen: Qwen3.8 27B – Prompt $0.45 / 1M, Completion $3.20 / 1M, context 262,144 tokens.
  • Dots Studio: Dots3‑Note Preview (free) – Prompt $0.00, Completion $0.00, context 512,000 tokens. Who should care: Developers needing a free, long‑context model for note‑taking or prototyping.

Price Changes (selected highlights)

Model Old Prompt ($/1M) New Prompt ($/1M) Δ Prompt Old Comp ($/1M) New Comp ($/1M) Δ Comp Δ Total ($/1M) Who should care
DeepSeek V4 Flash Latest 0.08 0.07 −0.01 0.252 0.14 −0.112 −0.122 Cost‑sensitive batch jobs
Z.ai: GLM 5.2 0.63 0.49 −0.14 1.98 1.54 −0.44 −0.58 Mid‑scale LLM users
MoonshotAI: Kimi K2.7 Code 0.67 0.71 +0.04 3.40 3.50 +0.10 +0.14 Teams needing slightly higher quality
DeepSeek V4 Flash 0423 0.14 0.074 −0.066 0.28 0.148 −0.132 −0.198 Flash inference workloads
Tencent: Hy3 preview 0.063 0.18 +0.117 0.21 0.60 +0.39 +0.507 Experimental preview users
MoonshotAI: Kimi K2.6 0.95 0.65 −0.30 4.00 3.41 −0.59 −0.89 Code‑gen pipelines
Qwen: Qwen3.5‑35B‑A3B 0.25 0.225 −0.025 1.25 1.80 +0.55 +0.525 Balanced prompt/completion shift
Qwen: Qwen3.5 397B A17B 0.50 0.39 −0.11 3.60 2.34 −1.26 −1.37 Large‑model cost optimizers
Z.ai: GLM 5 0.95 0.60 −0.35 2.55 1.92 −0.63 −0.98 General purpose GLM users
Z.ai: GLM 4.6 0.50 0.55 +0.05 2.00 2.20 +0.20 +0.25 Slight uptick for legacy GLM
Meta: Llama 4 Maverick 0.20 0.20 0.00 0.696 0.80 +0.104 +0.104 Maverick adopters

No models were removed today.

Total models tracked: 413.


Originally published at The Token Ledger. Subscribe for the daily digest.

Top comments (0)