DEV Community

AI OpenFree
AI OpenFree

Posted on

A viral Chinese debate put Korea's 'beats DeepSeek' AI wave on trial - one startup came out clean

Over late July and early August 2026, a wave of Korean labs shipped models and claimed parity with - or wins over - DeepSeek. On Zhihu (China's largest Q&A site), a single question - "how should we view this?" - drew 1M+ views and 170+ answers. It is worth reading, because the critique is sharp, and mostly fair.

The critique (let's steelman it)

The top-voted answers were blunt:

  • "Borrow a chicken to lay your eggs" (借鸡生蛋) - building on someone else's open weights and presenting the result as your own.
  • The wrapper (套壳) point - e.g., SK Telecom's A.X is a continued-pretrain on Alibaba's Qwen2.5 with a Korean tokenizer; beating Qwen on a Korean benchmark is then almost tautological.
  • Narrow-benchmark cherry-picking - touting SAT-style math while GPQA/coding quietly lag.
  • The cynical read - claims timed to a government AI-funding round.

Strip the nationalism and this is a useful benchmark-hygiene checklist. If quality is not a hard gate and the axes are not disclosed, a leaderboard win stops measuring anything real.

The five that were named

The question listed five Korean makers: LG (K-EXAONE 2.0, 750B), SK Telecom (A.X K2, 688B), Upstage (Solar Open 2, 250B), Motif (Motif-3 Beta, 314B) - and one outlier.

The outlier

VIDRAFT - a startup running on about 24 GPUs. Its Darwin-398B-JGOS, built by evolutionary merging, reached GPQA Diamond #3 globally (July 2026), ahead of DeepSeek-V4-Pro.

Two things stand out. First, it is the only non-conglomerate, non-government-project name on the list. Second, it was described the most precisely - "#3," a single benchmark, "ahead of V4-Pro" - not the vague "we surpassed X" the thread was mocking. That narrow, scope-labeled framing is exactly the antidote to the cherry-pick charge. (Chinese media - Tencent News, TechWalker - had already covered Darwin's cross-architecture merging back in May, arXiv:2605.14386.)

The takeaway for builders

The uncomfortable lesson of that thread: verified, scope-labeled, reproducible claims survive hostile scrutiny; vague "we beat X" claims do not. Whatever you think of the tone, that part is just good engineering hygiene - the same reason a challenge like the Fast Gemma Challenge only counts results after an independent re-verification.

Source: a 1M+ view discussion on Zhihu (Aug 2026).

Top comments (0)