The Token Ledger Digest – 2026-07-23
Lead: NVIDIA: Nemotron 3 Ultra – prompt price dropped from $0.60 to $0.50 /M tokens; completion price fell from $3.60 to $2.20 /M tokens (total –$1.50 /M). Cost‑sensitive workloads see ~42% lower completion cost.
- Qwen: Qwen3.6 27B – Prompt $0.45 → $0.60 /M; Completion $2.70 → $3.60 /M (Δ +$1.05 /M). Teams using this model for generation face higher per‑token spend.
- Qwen: Qwen3 14B – Prompt $0.12 → $0.2275 /M; Completion $0.24 → $0.91 /M (Δ +$0.78 /M). Notable cost rise for mid‑size LLM applications.
- Z.ai: GLM 5 – Prompt unchanged $0.95 /M; Completion $2.55 → $3.15 /M (Δ +$0.60 /M). Completion‑heavy pipelines will see increased expense.
- Qwen: Qwen3 Next 80B A3B Instruct – Prompt $0.10 → $0.15 /M; Completion $1.10 → $1.20 /M (Δ +$0.15 /M). Large‑scale inference budgets affected.
- Qwen: Qwen3 VL 30B A3B Instruct – Prompt $0.13 → $0.15 /M; Completion $0.52 → $0.60 /M (Δ +$0.10 /M). Vision‑language users notice modest cost increase.
- Google: Gemma 4 26B A4B – Prompt $0.07 → $0.12 /M; Completion $0.34 → $0.35 /M (Δ +$0.06 /M). Small uptick for lightweight Gemma deployments.
- Z.ai: GLM 5.2 – Prompt $0.7938 → $0.8008 /M; Completion $2.4948 → $2.5168 /M (Δ +$0.029 /M). Minor price tweak; negligible for most workloads.
- Google: Gemma 4 31B – Prompt unchanged $0.12 /M; Completion $0.37 → $0.35 /M (Δ –$0.02 /M). Slight completion‑cost relief.
- Meta: Llama 3.2 3B Instruct – Prompt $0.0509 → $0.05 /M; Completion $0.335 → $0.33 /M (Δ –$0.006 /M). Tiny reduction for edge‑device use.
- Z.ai: GLM 4.7 Flash – Prompt $0.0605 → $0.06 /M; Completion unchanged $0.40 /M (Δ –$0.0005 /M). Practically no change.
If no meaningful changes had occurred, we would note the three cheapest models today: inclusionAI: Ling‑2.6‑flash ($0.01/$0.03 /M), IBM: Granite 4.0 Micro ($0.017/$0.112 /M), and Mistral: Mistral Nemo ($0.019/$0.03 /M).
Originally published at The Token Ledger. Subscribe for the daily digest.
Top comments (0)