AI Roundup (Mon Aug 17)
A quick scan of what moved the frontier this weekend — pricing, open weights, and a model that exists but won't ship.
1. DeepSeek turns on peak/valley pricing (live Aug 17)
Starting 00:00 Beijing time on August 17, DeepSeek's API moved to time-of-day pricing. Off-peak (outside 09:00–12:00 and 14:00–18:00 BJT) runs at half the peak rate.
- V4-Pro: peak output ¥27 / $3.96 per M tokens, off-peak ¥13.5 / $1.98.
- V4-Flash: peak output ¥9 / $1.32 per M tokens, off-peak ¥4.5.
- The hike ends the era of rock-bottom flat rates — V4-Pro's off-peak output alone is ~2.25× its old list price, peak ~4.5×.
The signal: the open-weight leader is now pricing on capacity and capability, not just undercutting. Expect enterprises to shift bulk calls to off-peak windows.
2. Zhipu GLM-5.3 — best open-weights coding model yet (Aug 14)
Zhipu released GLM-5.3 on the same base as GLM-5.2 — every gain comes from post-training scaling. It now tops open-weights coding benchmarks:
- Terminal-Bench 3.0: 28.3 (from 4.6) and DeepSWE v1.1: 66.9 (from 46.2).
- CyberGym: 84.5% — matching/exceeding Claude Mythos 5 (83.8%) on vulnerability discovery.
- During pre-release testing it surfaced 2,404 potential vulnerabilities (1,088 high/medium), including a 40-year-old DNS protocol flaw.
Weights open in ~2 weeks after safety hardening. The takeaway: post-training alone is closing the gap to frontier closed models.
3. Anthropic's risk report discloses — then withholds — "Model 2" (Aug 14)
Anthropic's 186-page risk report revealed an internal-only model, Model 2, stronger than the public Claude Mythos 5:
- CoBench v2: 62.8% vs Mythos 5's 50.3% — already heavily used for coding, agentic work, and training-data generation inside the company.
- No release plans. Misalignment risk rating moved from "very low" to "low."
- The report also discloses real incidents: a biosafety classifier was silently off for ~11 months, letting ~133M conversations with ~50,000 contractors bypass it.
The frontier's real capability edge now lives inside the lab, not on the consumer product.
Keeping up with the AI frontier? Daily digests at AI Nexus Daily.
Top comments (0)