open-source and open-weight models are now within striking distance of Western flagships.
GLM-5.2: The Open-Weight Challenger
Released in mid-June 2026 by Zhipu AI under the MIT license, GLM-5.2 is a 753B-parameter MoE model with a solid 1M-token context window. Its benchmark performance tells the story:
SWE-bench Pro: 62.1 — surpassing GPT-5.5 (58.6) and leading Chinese open-source models
AIME 2026: 99.2 — outperforming Claude Opus 4.8's 95.7 and DeepSeek-V4-Pro's 94.6 on olympiad-grade math
Terminal Bench 2.1: 81.0 — up from 62.0 on GLM-5.1, a ~19-point leap
FrontierSWE: Trails Claude Opus 4.8 by just 1%
The architecture introduces IndexShare, a sparse attention mechanism that reduces per-token FLOPs by 2.9× at 1M context length. It also features improved MTP layers for speculative decoding, boosting acceptance length by up to 20%. GLM-5.2 became the fastest-deploying model on Vercel's platform, with daily token consumption growing 27x and customer count rising 80x in its first full week.
Qwen3.7 Series: Agentic Capabilities at Scale
Alibaba's Qwen3.7-Max ranks as the top domestic model on the Arena leaderboard. On Agentic benchmarks:
MCP-Atlas: 76.4, close to GLM-5.2's 76.8 and Opus 4.8's 77.8
AgentWorldBench: The 397B-A17B version surpasses GPT-5.4, Claude Opus 4.8, and Gemini 3.1 Pro
DeepSeek-V4: Efficiency by Design
DeepSeek-V4-Flash (284B total, 13B activated) delivers stable performance with extreme cost efficiency. It has held the #1 spot for Chinese token consumption for seven consecutive weeks. With the introduction of DSpark, a speculative decoding framework, inference speed increases by 60–85% at the same throughput.
Hy3: Tencent's Entry
Tencent's Hy3 (295B total, 21B activated, 256K context) entered the market in July 2026. Its ClawEval pass³ score of 68.5 surpasses DeepSeek-V4-Pro (62.4) and Qwen3.7-Max (65.2). Following its release, Tencent's market cap briefly topped 4 trillion HKD

Top comments (1)
GLM-5.2, Qwen3.7, DeepSeek-V4 & Tencent Hy3 deliver stunning open-source AI power with outstanding benchmarks and efficiency.