DEV Community

4663437Mehdi
4663437Mehdi

Posted on Originally published at 4663437mehdi.github.io

Token Ledger Digest – 2026-09-12

Token Ledger Digest – 2026-09-12

Most impactful change: Z.ai: GLM 5.2 saw its completion price fall from $3.04/1M to $2.00/1M (‑34%) and prompt price from $0.97/1M to $0.60/1M (‑38%). Developers running long‑form generation or batch workloads will see the biggest per‑token savings.

Added models

  • Inference.net: Schematron V2 Turbo – prompt $0.03/1M, completion $0.15/1M. Teams needing cheap structured‑output inference.
  • Inference.net: Schematron V2 Small – prompt $0.05/1M, completion $0.23/1M. Same use‑case, slightly higher quality.
  • OpenAI: GPT Astra Latest – prompt $10.00/1M, completion $50.00/1M. Cutting‑edge research or premium‑tier apps.
  • OpenAI: GPT Sol Latest – prompt $2.00/1M, completion $10.00/1M. Mid‑range commercial products.
  • OpenAI: GPT Terra Latest – prompt $2.00/1M, completion $12.00/1M. Similar to Sol with a bit more completion cost.
  • OpenAI: GPT Luna Latest – prompt $0.20/1M, completion $1.20/1M. Low‑latency, cost‑sensitive services.
  • inclusionAI: Ling 3.0 Flash VL – prompt $0.06/1M, completion $0.18/1M. Vision‑language tasks on a budget.

Removed model

  • OpenAI GPT Latest – prompt $2.00/1M, completion $10.00/1M. Users should migrate to Sol/Terra/Luna or another provider.

Price changes

  • Z.ai: GLM 5.3 Flash – prompt $0.15 → $0.075/1M (‑50%); completion $0.50 → $0.25/1M (‑50%). Cost‑conscious inference pipelines.
  • Z.ai: GLM Latest – prompt $0.97 → $0.873/1M (‑10%); completion $3.31 → $3.36/1M (+1.5%). Minor shift; watch for completion‑heavy jobs.
  • Qwen: Qwen3.8 27B – prompt $0.42 → $0.214/1M (‑49%); completion $3.00 → $2.55/1M (‑15%). Mid‑size model users benefit.
  • DeepSeek: DeepSeek V4 Pro 0813 – prompt $0.66 → $0.578/1M (‑12%); completion $1.98 → $1.734/1M (‑12%). Pro‑tier workloads.
  • DeepSeek: DeepSeek V4 Flash Latest – prompt $0.05 → $0.03/1M (‑40%); completion $0.16 → $0.07/1M (‑56%). Ultra‑low‑cost flash usage.
  • DeepSeek: DeepSeek V4 Flash 0731 – prompt $0.065 → $0.04/1M (‑38%); completion $0.18 → $0.08/1M (‑56%). Similar to above.
  • MoonshotAI: Kimi K3 – prompt $2.34 → $2.303/1M (‑1.6%); completion $11.70 → $11.55/1M (‑1.3%). Negligible impact.
  • Z.ai: GLM 5.2 – prompt $0.966 → $0.60/1M (‑38%); completion $3.04 → $2.00/1M (‑34%). Already highlighted as top saver.
  • MoonshotAI: Kimi Latest – prompt $2.34 → $2.125/1M (‑9%); completion $11.70 → $11.90/1M (+1.7%). Slight trade‑off.
  • DeepSeek: DeepSeek V4 Pro 0423 – prompt $0.955 → $0.789/1M (‑17%); completion $1.911 → $1.578/1M (‑17%). Pro‑tier savings.
  • DeepSeek: DeepSeek V4 Flash 0423 – prompt $0.0886 → $0.0668/1M (‑25%); completion $0.177 → $0.134/1M (‑24%). Flash‑tier efficiency.
  • MiniMax: MiniMax M2.5 – prompt $0.30 → $0.27/1M (‑10%); completion $1.20 → $1.08/1M (‑10%). Balanced cost cut.
  • Qwen: Qwen3 235B A22B Instruct 2507 – prompt $0.22 → $0.0875/1M (‑60%); completion $0.88 → $0.35/1M (‑60%). Large‑model users see major savings.
  • Meta: Llama 3.1 70B Instruct – prompt $0.40 → $0.72/1M (+80%); completion $0.40 → $0.72/1M (+80%). Cost increase; consider alternatives for budget‑sensitive apps.

Who should care: anyone integrating these models into production—especially those optimizing for cost per token, latency, or specific modalities (vision, long‑context). Adjust budgets or model selection accordingly.


Originally published at The Token Ledger. Subscribe for the daily digest.

Top comments (0)