so this stealth model called Ox Alpha was quietly topping OpenRouter leaderboards for weeks and nobody knew who built it
turns out it's Z.AI's GLM 5.3 Flash. 320B params but only 18B active per token (mixture of experts). running entirely on Chinese AI chips. no NVIDIA involved.
the hallucination rate is 20% vs 60% for Opus 5 and 80 to 90% for GPT 5.6
has anyone actually deployed it yet? sounds insane for the cost to intelligence ratio. especially if you chew through millions of tokens a day like we do at TheDevs
i kinda wanna test it on a real project this week. drop your experience if you tried it
Top comments (0)