Korea's #1 model on GPQA Diamond is a 24-GPU startup's — and it ranks #3 in the world
On GPQA Diamond (graduate-level, "Google-proof" science reasoning), the top-scoring Korean model is Darwin-398B-JGOS from VIDRAFT, a startup running on ~24 GPUs.
The numbers (base-only, July 2026)
| Rank | Model | Score |
|---|---|---|
| 1 | Kimi-K3 (CN) | 93.5 |
| 2 | GLM-5.2 (CN) | 91.2 |
| 3 | Darwin-398B-JGOS (VIDRAFT, KR) | 90.9 |
| 6 | DeepSeek-V4-Pro (CN) | 90.1 |
| 7 | Inkling-Small (US) | 89.5 |
| 9 | Qwen3.5-397B-A17B (CN) | 88.4 |
| 14 | Nemotron-3-Ultra-550B (US) | 87.9 |
Only two Chinese frontier models rank above it; it sits ahead of DeepSeek-V4-Pro, Qwen3.5-397B and Nvidia's 550B Nemotron-3-Ultra.
Korea gap
The sovereign-AI models trail: Solar-Open2-250B (86.3), A.X-K2 (85.6), K-EXAONE-2.0-750B-A37B (82.2, rank 43). Darwin leads the next Korean model by ~4.6 points.
Why it matters
GPQA Diamond resists web lookup, so scores track reasoning, not memorization. Darwin-398B is built by evolutionary merging of open models on a small cluster — method over scale — yet lands in the global top handful.
Honest scope
Base-only, single-benchmark snapshot (2026-07). Two Chinese frontier models still lead; GPQA scores shift with eval setup; one benchmark is not the whole picture.
Source: GPQA Diamond (Idavidrein/gpqa). More: https://vidraft.net

Top comments (0)