The Token Ledger – 2026-08-16
Most cost‑impacting change: Qwen Qwen3.6 27B saw its completion price drop from $3.60 to $2.00 per 1M tokens (‑$1.60/1M) and its prompt price fall from $0.60 to $0.30 per 1M tokens (‑$0.30/1M). Developers running long‑form generation workloads should see noticeable savings.
DeepSeek V4 Flash Latest – Prompt: $0.070 → $0.067/1M (‑$0.003); Completion: $0.140 → $0.134/1M (‑$0.006).
Who should care: Users of latency‑sensitive flash variants.Z.ai GLM 5.2 – Prompt: $0.49 → $0.31/1M (‑$0.18); Completion: $1.54 → $0.97/1M (‑$0.57).
Who should care: Teams balancing quality and cost for mid‑size models.Qwen Qwen3.6 27B – Prompt: $0.60 → $0.30/1M (‑$0.30); Completion: $3.60 → $2.00/1M (‑$1.60).
Who should care: Heavy‑generation pipelines (chat, summarization).DeepSeek V4 Flash 0423 – Prompt: $0.074 → $0.061/1M (‑$0.013); Completion: $0.148 → $0.123/1M (‑$0.025).
Who should care: Applications needing ultra‑low prompt overhead.MoonshotAI Kimi K2.6 – Prompt: $0.65 → $0.54/1M (‑$0.11); Completion: $3.41 → $2.28/1M (‑$1.13).
Who should care: Users of long‑context reasoning models.Qwen Plus 0728 (thinking) – Prompt: $0.40 → $0.26/1M (‑$0.14); Completion: $1.20 → $0.78/1M (‑$0.42).
Who should care: Workloads that enable thinking mode.Qwen Qwen3 30B A3B – Prompt: $0.12 → $0.13/1M (+$0.01); Completion: $0.50 → $0.52/1M (+$0.02).
Who should care: Slight cost increase; monitor usage if budget‑tight.
Cheapest models today (per 1M tokens):
- inclusionAI Ling‑2.6‑flash – Prompt $0.01, Completion $0.03
- IBM Granite 4.0 Micro – Prompt $0.017, Completion $0.112
- Mistral Nemo – Prompt $0.019, Completion $0.03
Total models tracked: 413. No models were added or removed today.
Originally published at The Token Ledger. Subscribe for the daily digest.
Top comments (0)