DEV Community

4663437Mehdi
4663437Mehdi

Posted on • Originally published at 4663437mehdi.github.io

AI Model Pricing Digest – 2026-07-28

AI Model Pricing Digest – 2026-07-28

Most cost‑impacting change: OpenAI cut prices for the GPT‑5.6 Terra series, slashing both prompt and completion costs by ~40‑50%.

Added

  • Qwen: Qwen3.7 Flash – New model. Prompt $0.03/1M, completion $0.13/1M. Who should care: Teams needing ultra‑long context (1M tokens) at low cost.

Removed

  • OpenAI: GPT-5 Chat – Deleted. Was $1.25/1M prompt, $10/1M completion. Who should care: Users of this model must migrate to alternatives.
  • OpenAI: GPT-4o Search Preview – Deleted. Was $2.5/1M prompt, $10/1M completion. Who should care: Anyone relying on search‑augmented preview should switch.

Price Changes

  • OpenAI: GPT-5.6 Luna Pro – Prompt ↓ $1.00 → $0.50/1M (-$0.50); Completion ↓ $6.00 → $3.00/1M (-$3.00). Who should care: Cost‑sensitive workloads using Luna Pro.
  • OpenAI: GPT-5.6 Luna – Same cuts as Luna Pro. Who should care: Luna users.
  • OpenAI: GPT-5.6 Terra Pro – Prompt ↓ $2.50 → $1.25/1M (-$1.25); Completion ↓ $15.00 → $7.50/1M (-$7.50). Who should care: High‑volume Terra Pro users see ~40% savings.
  • OpenAI: GPT-5.6 Terra – Identical cuts to Terra Pro. Who should care: Terra users.
  • Z.ai: GLM 5.2 – Prompt ↓ $0.8036 → $0.7686/1M (-$0.035); Completion ↓ $2.5256 → $2.4156/1M (-$0.11). Who should care: Minor savings for GLM 5.2 adopters.
  • NVIDIA: Nemotron 3 Ultra – Prompt ↑ $0.50 → $0.60/1M (+$0.10); Completion ↑ $2.20 → $3.60/1M (+$1.40). Who should care: Budget impact for Nemotron 3 Ultra users.
  • Google: Gemma 4 26B A4B – Prompt ↑ $0.12 → $0.14/1M (+$0.02); Completion ↑ $0.35 → $0.42/1M (+$0.07). Who should care: Slight cost rise for Gemma 4 users.
  • Qwen: Qwen3.5-35B-A3B – Prompt ↓ $0.15 → $0.14/1M (-$0.01); Completion unchanged $1.00/1M. Who should care: Negligible saving for this model.

Total models tracked: 341.


Originally published at The Token Ledger. Subscribe for the daily digest.

Top comments (0)