<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: AI OpenFree</title>
    <description>The latest articles on DEV Community by AI OpenFree (@ai_openfree_b23025ef075cf).</description>
    <link>https://dev.to/ai_openfree_b23025ef075cf</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3817626%2F25ccf2de-8934-44fe-9d10-59dd6d2a505b.png</url>
      <title>DEV Community: AI OpenFree</title>
      <link>https://dev.to/ai_openfree_b23025ef075cf</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ai_openfree_b23025ef075cf"/>
    <language>en</language>
    <item>
      <title>VIDRAFT's AX-RAY AI Safety Diagnostic Platform Earns Alibaba ModelScope "Research Institution" Certification — Here's What Engineers Need to Know</title>
      <dc:creator>AI OpenFree</dc:creator>
      <pubDate>Fri, 02 Oct 2026 07:01:22 +0000</pubDate>
      <link>https://dev.to/ai_openfree_b23025ef075cf/vidrafts-ax-ray-ai-safety-diagnostic-platform-earns-alibaba-modelscope-research-institution-3i0c</link>
      <guid>https://dev.to/ai_openfree_b23025ef075cf/vidrafts-ax-ray-ai-safety-diagnostic-platform-earns-alibaba-modelscope-research-institution-3i0c</guid>
      <description>&lt;h1&gt;
  
  
  VIDRAFT's AX-RAY AI Safety Diagnostic Platform Earns Alibaba ModelScope "Research Institution" Certification — Here's What Engineers Need to Know
&lt;/h1&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; VIDRAFT, a Korean Pre-AGI AI startup, has become the first Korean AI organization to receive official "Research Institution" certification on Alibaba's ModelScope platform, joining a short list that includes Qwen, Shanghai AI Lab, and Z.ai (Zhipu AI). The company is now pushing AX-RAY — its structured AI safety diagnostic framework covering 12 categories and 117 risk items — into the Chinese AI ecosystem. If you're building or auditing LLMs for regulatory compliance, AX-RAY's public evaluation reports on Hugging Face are worth your attention.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  What it is
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;ModelScope&lt;/strong&gt; is a large-scale open-source AI model community co-founded by Alibaba and the China Computer Federation (CCF). It hosts over 140,000 models and serves more than 20 million users, making it China's largest AI distribution platform. While major organizations including Google, Meta, and DeepSeek maintain presences on the platform, none of them currently hold the formal "Research Institution" certification badge — a distinction VIDRAFT has now secured roughly seven months after opening its ModelScope organization page (&lt;code&gt;FINAL-Bench&lt;/code&gt;) in February 2026.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AX-RAY&lt;/strong&gt; is VIDRAFT's AI safety diagnostic platform. It evaluates AI models across:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;12 categories&lt;/strong&gt; of safety concern&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;117 risk items&lt;/strong&gt; aligned with nine global regulatory frameworks, including:

&lt;ul&gt;
&lt;li&gt;EU AI Act&lt;/li&gt;
&lt;li&gt;GDPR&lt;/li&gt;
&lt;li&gt;U.S. NIST AI Risk Management Framework (AI RMF)&lt;/li&gt;
&lt;li&gt;ISO/IEC 42001&lt;/li&gt;
&lt;li&gt;Domestic Korean statutes&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Models receive a letter grade from &lt;strong&gt;A through F&lt;/strong&gt; based on their evaluated risk profile. Diagnostic results and detailed reports are published publicly on Hugging Face.&lt;/p&gt;

&lt;p&gt;VIDRAFT's ModelScope organization currently hosts &lt;strong&gt;91 public models and 9 datasets&lt;/strong&gt;, including:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Darwin&lt;/strong&gt; — a Korean-language-specialized LLM&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;POCKET&lt;/strong&gt; — a lightweight on-device model&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AX-RAY datasets&lt;/strong&gt; — the AI safety evaluation benchmark data&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  How it works
&lt;/h2&gt;

&lt;p&gt;AX-RAY operates as a structured audit framework rather than a single benchmark task. At a conceptual level:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Risk taxonomy construction:&lt;/strong&gt; The 117 risk items are derived from a cross-mapping of nine international regulatory and standards frameworks, enabling the same diagnostic run to generate compliance-relevant documentation for multiple jurisdictions simultaneously.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Multi-category evaluation:&lt;/strong&gt; Each model under test is probed across all 12 safety categories. This appears to cover areas including AI agent behavior, harmful content generation, and small-model-specific failure modes — based on the public diagnostic results VIDRAFT has shared.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Graded output:&lt;/strong&gt; The A–F grading system produces a human-readable safety score that can be attached to regulatory filing documentation, making it practically useful for teams preparing AI Act conformity assessments or NIST AI RMF profiles.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;VIDRAFT is currently participating in a security-specialized AI foundation model development project commissioned by South Korea's Ministry of Science and ICT (MSIT) and the National IT Industry Promotion Agency (NIPA), as part of the Naver Cloud consortium.&lt;/p&gt;




&lt;h2&gt;
  
  
  Benchmarks &amp;amp; results
&lt;/h2&gt;

&lt;p&gt;VIDRAFT has published results from an AX-RAY diagnostic run covering &lt;strong&gt;40 publicly available AI models&lt;/strong&gt; (domestic and international). Key findings:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;45% of models evaluated (18 out of 40)&lt;/strong&gt; received the lowest possible grade (F).&lt;/li&gt;
&lt;li&gt;In the &lt;strong&gt;"AI Agent Safety" category&lt;/strong&gt;, &lt;strong&gt;23 out of 25 models (92%)&lt;/strong&gt; showed detected risk signals.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;All 12 small models with fewer than 4 billion parameters&lt;/strong&gt; that were evaluated received a risk determination — a 100% failure rate in that cohort.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These results suggest that agent-mode safety and small-model safety are currently the weakest points in the broader open-model ecosystem, at least under AX-RAY's evaluation criteria.&lt;/p&gt;

&lt;p&gt;On Hugging Face, VIDRAFT's 82 public models and datasets accumulated &lt;strong&gt;1.01 million downloads in the most recent 30-day window&lt;/strong&gt; at time of reporting.&lt;/p&gt;




&lt;h2&gt;
  
  
  How to try it
&lt;/h2&gt;

&lt;p&gt;VIDRAFT's models, datasets, and AX-RAY diagnostic reports are publicly accessible on both Hugging Face and ModelScope:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hugging Face&lt;/strong&gt; — browse VIDRAFT's published models and datasets (including AX-RAY evaluation reports):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;https://huggingface.co/FINAL-Bench
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;ModelScope&lt;/strong&gt; — access the certified FINAL-Bench organization page on ModelScope:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;https://modelscope.cn/organization/FINAL-Bench
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You can use the Hugging Face CLI to explore available assets:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;huggingface_hub
huggingface-cli scan-cache
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No private API access, waitlist, or paid tier is mentioned for the public models and benchmark reports. The AX-RAY diagnostic &lt;em&gt;platform&lt;/em&gt; (as a service for evaluating your own models) has not yet been described as publicly self-serve — check the Hugging Face organization page for current availability.&lt;/p&gt;




&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Q: How does AX-RAY differ from existing safety benchmarks like MT-Bench or HELM?&lt;/strong&gt;&lt;br&gt;
A: AX-RAY is explicitly designed for regulatory compliance output, not just relative model ranking. Its 117 risk items are cross-referenced against nine specific legal and standards frameworks, so a diagnostic run can directly support documentation for the EU AI Act, NIST AI RMF, and ISO/IEC 42001 in a single pass. Most academic benchmarks don't map to legal instruments in this way.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Why does ModelScope certification matter for developers outside China?&lt;/strong&gt;&lt;br&gt;
A: ModelScope's 20M+ user base and its co-governance by the China Computer Federation make it a primary distribution channel for AI models in Chinese research and industry contexts. For teams shipping models that need to reach Chinese enterprise or academic users, having a verified "Research Institution" presence there carries similar trust signaling to a verified organization on Hugging Face. It also opens distribution to users who may not access Hugging Face directly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Are the AX-RAY evaluation datasets themselves open?&lt;/strong&gt;&lt;br&gt;
A: Yes — per the source, AX-RAY datasets are listed among the 9 datasets registered in VIDRAFT's ModelScope organization and are part of the 82 assets publicly available on Hugging Face.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally reported by AI타임스 (2026-10-02) — &lt;a href="https://www.aitimes.com/news/articleView.html?idxno=215897" rel="noopener noreferrer"&gt;source article&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>opensource</category>
      <category>llm</category>
    </item>
    <item>
      <title>Darwin-180B-RSI: How VIDRAFT's Recursive Self-Improvement Model Topped 7 Hugging Face Leaderboards — Without Legal Training Data</title>
      <dc:creator>AI OpenFree</dc:creator>
      <pubDate>Thu, 01 Oct 2026 11:01:38 +0000</pubDate>
      <link>https://dev.to/ai_openfree_b23025ef075cf/darwin-180b-rsi-how-vidrafts-recursive-self-improvement-model-topped-7-hugging-face-leaderboards-5g46</link>
      <guid>https://dev.to/ai_openfree_b23025ef075cf/darwin-180b-rsi-how-vidrafts-recursive-self-improvement-model-topped-7-hugging-face-leaderboards-5g46</guid>
      <description>&lt;h1&gt;
  
  
  Darwin-180B-RSI: How VIDRAFT's Recursive Self-Improvement Model Topped 7 Hugging Face Leaderboards — Without Legal Training Data
&lt;/h1&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; VIDRAFT's &lt;code&gt;Darwin-180B-RSI&lt;/code&gt; is a 180-billion-parameter reasoning model trained via a Recursive Self-Improvement (RSI) framework that updates model weights directly — no human-annotated chain-of-thought required. It just claimed #1 on both the &lt;code&gt;LEXam&lt;/code&gt; and &lt;code&gt;LEXam-hard&lt;/code&gt; Hugging Face-certified legal reasoning benchmarks, despite never being trained on legal domain data. Its checkpoint is publicly available on Hugging Face for external reproduction.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What it is
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;Darwin-180B-RSI&lt;/code&gt; is VIDRAFT's self-improving large language model, built around a &lt;strong&gt;model-level&lt;/strong&gt; Recursive Self-Improvement (RSI) architecture. Key facts from the announcement:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;180 billion parameters&lt;/strong&gt; in size&lt;/li&gt;
&lt;li&gt;Achieves top scores across &lt;strong&gt;7 Hugging Face-certified leaderboard categories&lt;/strong&gt; — the most of any single AI organization across all 47 officially certified Hugging Face benchmark leaderboards at time of publication&lt;/li&gt;
&lt;li&gt;Covers math, science, general knowledge, visual reasoning, and now &lt;strong&gt;legal reasoning&lt;/strong&gt; (LEXam, LEXam-hard) — the last two domains added without any legal-specific training data&lt;/li&gt;
&lt;li&gt;Outperforms models up to &lt;strong&gt;685 billion parameters&lt;/strong&gt; (e.g., DeepSeek-R1) on legal benchmarks despite being less than a third of the parameter count&lt;/li&gt;
&lt;li&gt;Model checkpoint and evaluation environment are &lt;strong&gt;open-sourced on Hugging Face&lt;/strong&gt; for independent verification&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  How it works
&lt;/h2&gt;

&lt;p&gt;The core design principle is &lt;strong&gt;Recursive Self-Improvement at the weight level&lt;/strong&gt; — a conceptually distinct approach from system-level self-improvement (like prompt tuning or tool-orchestration loops that leave base weights unchanged).&lt;/p&gt;

&lt;p&gt;Here's the high-level loop:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Autonomous problem solving:&lt;/strong&gt; The model attempts problems from domains where correctness can be verified mechanically — primarily mathematics and science.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Automated verification:&lt;/strong&gt; A rule-based or formal checker determines whether each candidate solution is correct. No human-written intermediate reasoning traces are injected.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Selective retraining:&lt;/strong&gt; Only the reasoning paths that led to verified correct answers are used to update the model's weights. Incorrect paths are discarded.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Iterative bootstrapping:&lt;/strong&gt; The improved model becomes the compute substrate for the &lt;em&gt;next&lt;/em&gt; improvement round, creating a recursive loop.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The practical result is that reasoning capabilities generalized from math and science &lt;strong&gt;transfer to unseen domains&lt;/strong&gt; like legal analysis — the model isn't memorizing legal statutes, it's applying structured multi-step logical inference. As a secondary efficiency gain, the RSI-trained model reportedly shortened reasoning token length by &lt;strong&gt;11%&lt;/strong&gt; relative to the base model, pruning redundant inference steps without sacrificing accuracy.&lt;/p&gt;

&lt;p&gt;This is architecturally different from Google's system-level RRSI approach (which adjusts prompts, external tools, and workflows while keeping base weights frozen). VIDRAFT's approach bakes improvements &lt;strong&gt;permanently into a single weight file&lt;/strong&gt;, meaning users get the full capability from the model artifact alone, with no additional scaffolding required at inference time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Benchmarks &amp;amp; results
&lt;/h2&gt;

&lt;p&gt;All scores below are from the Hugging Face-certified LEXam benchmark, developed jointly by ETH Zurich, University of Zurich, and the Max Planck Institute. LEXam is built from 340 real law school final exams and tests applied legal reasoning under civil law — not statute memorization.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;LEXam (multiple-choice, 1,655 questions):&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Score&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Darwin-180B-RSI&lt;/strong&gt; (4-pass majority vote)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;68.94&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Darwin-180B-RSI (1-pass, single inference)&lt;/td&gt;
&lt;td&gt;60.54&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5 &lt;em&gt;(benchmark authors' measurement)&lt;/em&gt;
&lt;/td&gt;
&lt;td&gt;62.65&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude-4.5-Sonnet &lt;em&gt;(benchmark authors' measurement)&lt;/em&gt;
&lt;/td&gt;
&lt;td&gt;58.01&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini-2.5-Pro &lt;em&gt;(benchmark authors' measurement)&lt;/em&gt;
&lt;/td&gt;
&lt;td&gt;55.72&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek-R1 (685B)&lt;/td&gt;
&lt;td&gt;52.41&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3-235B-Thinking&lt;/td&gt;
&lt;td&gt;48.19&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gpt-oss-120b&lt;/td&gt;
&lt;td&gt;47.71&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;LEXam-hard (open-ended, 518 questions — scored by official grading model):&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Score&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Darwin-180B-RSI&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;45.72&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Inkling (Thinking Machines)&lt;/td&gt;
&lt;td&gt;40.82&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek-V4-Pro&lt;/td&gt;
&lt;td&gt;38.93&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Context: random-chance baseline on the 4-option multiple-choice LEXam is 25 points. Most top LLMs average ~50 points on it; LEXam-hard's previous best was in the low 40s.&lt;/p&gt;

&lt;p&gt;Across all 47 Hugging Face certified leaderboards, VIDRAFT now holds #1 in 7 categories — ahead of Zhipu AI (4), DeepSeek / Xiaomi / Moonshot AI (3 each).&lt;/p&gt;

&lt;h2&gt;
  
  
  How to try it
&lt;/h2&gt;

&lt;p&gt;The &lt;code&gt;Darwin-180B-RSI&lt;/code&gt; model checkpoint and its evaluation environment are &lt;strong&gt;publicly available on Hugging Face&lt;/strong&gt; as open-source. VIDRAFT has also disclosed the full benchmark evaluation parameters — number of inference passes, ensemble majority-vote settings, and reasoning token budget — to enable independent reproduction.&lt;/p&gt;

&lt;p&gt;To pull the model via the Hugging Face CLI:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;huggingface-cli download VIDRAFT/Darwin-180B-RSI
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;⚠️ Verify the exact repository path on &lt;a href="https://huggingface.co/VIDRAFT" rel="noopener noreferrer"&gt;huggingface.co/VIDRAFT&lt;/a&gt; before downloading. The source confirms open availability but does not specify a direct URL string.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Q: Does Darwin-180B-RSI require legal fine-tuning data to perform well on legal benchmarks?&lt;/strong&gt;&lt;br&gt;
A: No. The model was explicitly evaluated &lt;em&gt;without&lt;/em&gt; any prior legal domain training. Its legal reasoning performance comes from generalizing multi-step logical inference skills acquired during math and science self-improvement cycles.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: What's the practical difference between VIDRAFT's model-level RSI and system-level self-improvement (like Google's RRSI)?&lt;/strong&gt;&lt;br&gt;
A: System-level approaches adjust prompts, tools, and orchestration workflows at runtime while leaving the underlying model weights unchanged. VIDRAFT's RSI updates the model's weights directly, so improvements are permanently embedded in the weight file. You don't need any additional runtime infrastructure — just load the model and run inference. VIDRAFT notes both approaches are complementary rather than competing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: How do I know the benchmark results are legitimate and not cherry-picked evaluation conditions?&lt;/strong&gt;&lt;br&gt;
A: VIDRAFT publicly disclosed all evaluation parameters: inference pass count, majority-vote configuration, and reasoning token budget. The checkpoint is on Hugging Face so anyone can re-run the LEXam evaluation independently. The benchmark itself was built by academic researchers at ETH Zurich, University of Zurich, and the Max Planck Institute, with answers validated by legal professionals.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally reported by 로이슈 (2026-10-01) — &lt;a href="https://www.lawissue.co.kr/view.php?ud=2026100113490567359aeda69934_12" rel="noopener noreferrer"&gt;source article&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>opensource</category>
      <category>llm</category>
    </item>
    <item>
      <title>Darwin-180B-RSI Tops Global Legal Reasoning Benchmarks — Without Ever Training on Legal Data</title>
      <dc:creator>AI OpenFree</dc:creator>
      <pubDate>Thu, 01 Oct 2026 07:01:23 +0000</pubDate>
      <link>https://dev.to/ai_openfree_b23025ef075cf/darwin-180b-rsi-tops-global-legal-reasoning-benchmarks-without-ever-training-on-legal-data-324e</link>
      <guid>https://dev.to/ai_openfree_b23025ef075cf/darwin-180b-rsi-tops-global-legal-reasoning-benchmarks-without-ever-training-on-legal-data-324e</guid>
      <description>&lt;h1&gt;
  
  
  Darwin-180B-RSI Tops Global Legal Reasoning Benchmarks — Without Ever Training on Legal Data
&lt;/h1&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; VIDRAFT's Darwin-180B-RSI reasoning model has claimed #1 on both the LEXam and LEXam-hard Hugging Face-certified legal reasoning benchmarks, despite never being trained on legal data. It achieves this through Recursive Self-Improvement (RSI) — a self-supervised loop driven entirely by math and science problems. Developers working in legal tech, public sector AI, or high-stakes reasoning applications should take note.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What it is
&lt;/h2&gt;

&lt;p&gt;Darwin-180B-RSI is a 180-billion-parameter reasoning model developed by VIDRAFT (비드래프트), a Korean AI startup. Its headline result: &lt;strong&gt;#1 on both LEXam and LEXam-hard&lt;/strong&gt;, two Hugging Face-certified legal reasoning benchmarks developed by ETH Zurich and collaborators, without the model ever having been fine-tuned on legal corpora.&lt;/p&gt;

&lt;p&gt;The Darwin family of models is built on a foundational approach that &lt;strong&gt;merges and combines the strengths of existing pre-trained models&lt;/strong&gt; rather than training from scratch — a design philosophy aimed at dramatically reducing the compute cost typically associated with large-scale pre-training.&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;-RSI&lt;/code&gt; suffix designates models that have additionally undergone VIDRAFT's Recursive Self-Improvement training procedure (described below).&lt;/p&gt;

&lt;h2&gt;
  
  
  How it works
&lt;/h2&gt;

&lt;p&gt;Darwin-180B-RSI's performance on legal reasoning is the product of two conceptually distinct techniques applied on top of the base Darwin architecture:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Recursive Self-Improvement (RSI)&lt;/strong&gt;&lt;br&gt;
Instead of relying on human-authored chain-of-thought annotations, the model enters a self-supervised loop:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The model attempts to solve problems (sourced from math and science domains only).&lt;/li&gt;
&lt;li&gt;An automated verifier checks whether the model's solution is correct.&lt;/li&gt;
&lt;li&gt;Solutions that pass verification are fed back as training signal.&lt;/li&gt;
&lt;li&gt;This cycle repeats, letting the model iteratively refine its own reasoning capability without human labeling.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The key insight is that &lt;strong&gt;general reasoning ability — not domain-specific knowledge — is what transfers to legal exam performance&lt;/strong&gt;. Legal reasoning requires applying rules to facts and drawing conclusions, a structure that maps closely to rigorous mathematical reasoning.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Zero-Token Confidence (ZTC)&lt;/strong&gt;&lt;br&gt;
ZTC is VIDRAFT's technique for estimating the model's internal confidence in a given answer without requiring extra inference tokens. This is used during the self-improvement loop to help the model gauge answer reliability, and it contributes to a reported &lt;strong&gt;11% reduction in average reasoning chain length&lt;/strong&gt; relative to the parent model — meaning fewer tokens consumed at inference time with no accuracy regression.&lt;/p&gt;

&lt;p&gt;The base Darwin architecture itself achieves cost efficiency by &lt;strong&gt;merging and composing existing model checkpoints&lt;/strong&gt; rather than running a full pre-training run from scratch. Only 0.02% of total parameters are modified during the self-improvement phase.&lt;/p&gt;

&lt;h2&gt;
  
  
  Benchmarks &amp;amp; results
&lt;/h2&gt;

&lt;p&gt;All scores below are sourced directly from the AI타임스 report and reflect results on Hugging Face-certified leaderboards.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;LEXam (multiple-choice legal reasoning):&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Score&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Darwin-180B-RSI&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;68.94&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5&lt;/td&gt;
&lt;td&gt;62.65&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude 4.5 Sonnet&lt;/td&gt;
&lt;td&gt;58.01&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek-R1&lt;/td&gt;
&lt;td&gt;52.41&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen-235B-Thinking&lt;/td&gt;
&lt;td&gt;48.19&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gpt-oss-120b&lt;/td&gt;
&lt;td&gt;47.71&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;LEXam-hard (open-ended / essay-style legal reasoning):&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Score&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Darwin-180B-RSI&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;45.72&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Inkling (Thinking Machines) — previous #1&lt;/td&gt;
&lt;td&gt;40.82&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Darwin-180B-RSI's 180B parameter count is notably smaller than DeepSeek-R1's 685B, yet it substantially outperforms it on both benchmarks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Broader leaderboard standing:&lt;/strong&gt; With these results, VIDRAFT now holds &lt;strong&gt;#1 positions across 7 of 47 categories&lt;/strong&gt; on Hugging Face-certified leaderboards, spanning math, science, general knowledge, visual reasoning, and legal reasoning — the highest count among Korean AI companies, and ahead of Zhipu AI (4), DeepSeek, Xiaomi, and Moonshot AI (3 each).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;LEXam benchmark context:&lt;/strong&gt; LEXam is based on real university law school exams and is designed to test the ability to apply statutes to factual scenarios and derive legal conclusions — not rote memorization of law. This makes it a meaningful proxy for practical legal reasoning ability.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to try it
&lt;/h2&gt;

&lt;p&gt;The article states that &lt;strong&gt;the model and evaluation conditions are publicly available on Hugging Face for independent verification&lt;/strong&gt;. You can explore VIDRAFT's public presence at:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;🤗 Hugging Face: search for &lt;code&gt;VIDRAFT&lt;/code&gt; or &lt;code&gt;Darwin-180B-RSI&lt;/code&gt; on &lt;a href="https://huggingface.co" rel="noopener noreferrer"&gt;huggingface.co&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;At the time of writing, no public GitHub repository URL, API endpoint, or install command has been announced in this report. Check VIDRAFT's official channels for the latest on API access or model weights availability.&lt;/p&gt;




&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Q: How can a model with zero legal training data outperform models that presumably have seen legal text?&lt;/strong&gt;&lt;br&gt;
A: The results suggest that LEXam is testing structured reasoning — applying rules to facts to reach conclusions — rather than legal memorization. RSI trained on math and science problems appears to build exactly that kind of generalizable deductive capability, which transfers surprisingly well to legal exam formats.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Is Darwin-180B-RSI smaller or larger than the models it beat?&lt;/strong&gt;&lt;br&gt;
A: Smaller in some key comparisons. At 180B parameters it outperforms DeepSeek-R1 (685B parameters) on both LEXam benchmarks. Parameter count is clearly not the dominant factor here.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: What domains is VIDRAFT targeting with this technology?&lt;/strong&gt;&lt;br&gt;
A: According to CEO Minsik Kim, the target verticals are legal services, public administration, and financial services — domains where accurate, auditable judgment under uncertainty is critical.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Does ZTC mean faster inference?&lt;/strong&gt;&lt;br&gt;
A: ZTC contributed to an 11% reduction in reasoning chain length compared to the parent model, which translates directly to lower token consumption at inference time while maintaining parent-model-level accuracy.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally reported by AI타임스 (2026-10-01) — &lt;a href="https://www.aitimes.com/news/articleView.html?idxno=215851" rel="noopener noreferrer"&gt;source article&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>opensource</category>
      <category>llm</category>
    </item>
    <item>
      <title>VIDRAFT Darwin-180B-RSI Tops 5 Hugging Face Leaderboards: AIME &amp; HMMT Perfect Scores, 94.44% on GPQA Diamond</title>
      <dc:creator>AI OpenFree</dc:creator>
      <pubDate>Wed, 30 Sep 2026 03:01:24 +0000</pubDate>
      <link>https://dev.to/ai_openfree_b23025ef075cf/vidraft-darwin-180b-rsi-tops-5-hugging-face-leaderboards-aime-hmmt-perfect-scores-9444-on-35kc</link>
      <guid>https://dev.to/ai_openfree_b23025ef075cf/vidraft-darwin-180b-rsi-tops-5-hugging-face-leaderboards-aime-hmmt-perfect-scores-9444-on-35kc</guid>
      <description>&lt;h1&gt;
  
  
  VIDRAFT Darwin-180B-RSI Tops 5 Hugging Face Leaderboards: AIME &amp;amp; HMMT Perfect Scores, 94.44% on GPQA Diamond
&lt;/h1&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; VIDRAFT's Darwin-180B-RSI is a 180B-parameter reasoning model that achieves #1 rankings across five Hugging Face-certified leaderboards — including the first-ever perfect scores on AIME 2026 and HMMT 2026 benchmarks. It combines three proprietary techniques (Darwin selective merging, Recursive Self-Improvement, and Zero-Token Confidence) to push frontier-level math, science, and multimodal reasoning. If you're evaluating reasoning models for hard STEM workloads, this is a result worth tracking.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  What it is
&lt;/h2&gt;

&lt;p&gt;Darwin-180B-RSI is VIDRAFT's latest reasoning-focused large language model. The "180B" refers to its parameter count, and "RSI" stands for Recursive Self-Improvement — a core training methodology baked into the model's name.&lt;/p&gt;

&lt;p&gt;Key characteristics from the source:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Scale:&lt;/strong&gt; 180 billion parameters (VIDRAFT has separately extended the same Darwin merging technology to a ~397-billion-parameter model class)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Focus:&lt;/strong&gt; High-difficulty reasoning across mathematics, science, and multimodal domains&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Certification:&lt;/strong&gt; Rankings verified on Hugging Face's officially certified leaderboards — not self-reported numbers&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Efficiency signal:&lt;/strong&gt; Achieves the same accuracy as prior iterations while reducing inference reasoning length by 11%&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  How it works
&lt;/h2&gt;

&lt;p&gt;Darwin-180B-RSI is built on three proprietary technologies that VIDRAFT describes at a conceptual level:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Darwin — Selective Model Merging
&lt;/h3&gt;

&lt;p&gt;Darwin analyzes performance at the layer and expert (MoE-style) unit level across multiple source models, then selectively identifies and combines the strongest components from each. Rather than averaging model weights indiscriminately, it surgically extracts high-performing substructures to produce a merged model that inherits complementary strengths. VIDRAFT has scaled this technique up to the ~397B parameter range.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. RSI — Recursive Self-Improvement
&lt;/h3&gt;

&lt;p&gt;RSI is a training loop in which the model:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Attempts to solve problems autonomously&lt;/li&gt;
&lt;li&gt;Verifies whether its own answers are correct&lt;/li&gt;
&lt;li&gt;Trains on the verified, high-quality reasoning traces&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The improved model then becomes the solver for the &lt;em&gt;next&lt;/em&gt; iteration of this loop — a self-bootstrapping cycle that incrementally raises reasoning quality without requiring human-labeled solutions at each step.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. ZTC — Zero-Token Confidence
&lt;/h3&gt;

&lt;p&gt;Before generating a response, ZTC analyzes the model's internal state to compute a confidence estimate for the candidate answer. If confidence falls below a threshold, the model is designed to withhold its response rather than produce a low-confidence output. This is a calibration mechanism aimed at reducing hallucination by giving the model a principled "abstain" option.&lt;/p&gt;




&lt;h2&gt;
  
  
  Benchmarks &amp;amp; results
&lt;/h2&gt;

&lt;p&gt;All figures below are from Hugging Face-certified leaderboards as reported by VIDRAFT on 2026-09-28:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Benchmark&lt;/th&gt;
&lt;th&gt;Darwin-180B-RSI&lt;/th&gt;
&lt;th&gt;Previous SOTA (on that leaderboard)&lt;/th&gt;
&lt;th&gt;Notes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;AIME 2026&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;100%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;97.1%&lt;/td&gt;
&lt;td&gt;Math Olympiad-level problems; first perfect score on this leaderboard&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;HMMT 2026&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;100%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;92.7%&lt;/td&gt;
&lt;td&gt;Math Olympiad-level problems; first perfect score on this leaderboard&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;GPQA Diamond&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;94.44%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;93.5% (Kimi-K3)&lt;/td&gt;
&lt;td&gt;PhD-level expert-authored science questions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;MMLU-Pro&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;88.12%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;14 professional domains incl. law, medicine, economics&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;MMMU-Pro&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;79.48%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;Multimodal reasoning: charts, sheet music, medical imaging&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Darwin-180B-RSI ranked &lt;strong&gt;#1 on all five leaderboards simultaneously&lt;/strong&gt;. The AIME and HMMT perfect scores are noted as the first time any model has achieved 100% on those Hugging Face-certified boards.&lt;/p&gt;




&lt;h2&gt;
  
  
  How to try it
&lt;/h2&gt;

&lt;p&gt;The source article does not include explicit Hugging Face model IDs, GitHub repository links, or API endpoint details for Darwin-180B-RSI at the time of publication. &lt;strong&gt;Access details have not been publicly confirmed in this report.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;What is publicly known from related VIDRAFT coverage:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;VIDRAFT models have been distributed via Hugging Face (the company has surpassed 1.6 million model downloads on the platform as of September 2026)&lt;/li&gt;
&lt;li&gt;VIDRAFT has previously open-released evaluation datasets and diagnostic tooling (AX-RAY) on Hugging Face&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;To follow release announcements, monitor VIDRAFT's Hugging Face organization page and their official channels directly. As soon as a model card or repository goes public, standard access patterns would apply.&lt;/p&gt;




&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Q: Are the benchmark results independently verified or self-reported?&lt;/strong&gt;&lt;br&gt;
A: VIDRAFT explicitly states these rankings come from Hugging Face's &lt;em&gt;certified&lt;/em&gt; leaderboards — meaning the evaluation pipeline is managed by Hugging Face, not run in-house by VIDRAFT. That's a meaningful distinction from self-reported numbers.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: What makes the 11% reasoning-length reduction significant?&lt;/strong&gt;&lt;br&gt;
A: Shorter reasoning chains at equivalent accuracy directly reduce inference compute costs and latency. For production deployments where you're paying per token or running latency-sensitive workloads, a model that reasons more concisely without sacrificing correctness is practically valuable — not just a benchmark curiosity.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: How does Darwin differ from standard model merging techniques like SLERP or TIES-merging?&lt;/strong&gt;&lt;br&gt;
A: Standard merging methods operate on full weight tensors with relatively coarse granularity. Darwin's distinguishing claim is that it evaluates and selects at the &lt;strong&gt;layer and expert unit level&lt;/strong&gt;, enabling more surgical recombination of specialized capabilities. The specific algorithmic details beyond this conceptual description have not been publicly disclosed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Is ZTC related to existing confidence calibration or selective prediction research?&lt;/strong&gt;&lt;br&gt;
A: Conceptually it occupies similar territory — the model abstains when uncertain rather than guessing. VIDRAFT's specific implementation approach has not been detailed in public materials beyond the mechanism described above.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally reported by 지디넷코리아 (2026-09-28) — &lt;a href="https://zdnet.co.kr/view/?no=20260928170507" rel="noopener noreferrer"&gt;source article&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>opensource</category>
      <category>llm</category>
    </item>
    <item>
      <title>Darwin-180B-RSI: VIDRAFT's Recursive Self-Improvement MoE Model Claims Five Leaderboard Tops</title>
      <dc:creator>AI OpenFree</dc:creator>
      <pubDate>Tue, 29 Sep 2026 23:01:18 +0000</pubDate>
      <link>https://dev.to/ai_openfree_b23025ef075cf/darwin-180b-rsi-vidrafts-recursive-self-improvement-moe-model-claims-five-leaderboard-tops-3ela</link>
      <guid>https://dev.to/ai_openfree_b23025ef075cf/darwin-180b-rsi-vidrafts-recursive-self-improvement-moe-model-claims-five-leaderboard-tops-3ela</guid>
      <description>&lt;h1&gt;
  
  
  Darwin-180B-RSI: VIDRAFT's Recursive Self-Improvement MoE Model Claims Five Leaderboard Tops
&lt;/h1&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; VIDRAFT, a Korean Pre-AGI AI startup, has released Darwin-180B-RSI — a 180-billion-parameter mixture-of-experts reasoning model built on Qwen3.8-Flash-Next that combines selective model merging with recursive self-improvement. The model reportedly tops five Hugging Face leaderboards across math, science, general, and multimodal reasoning benchmarks. If the reported gains hold on task-specific workloads, it's worth a closer look for reasoning-heavy applications.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  What it is
&lt;/h2&gt;

&lt;p&gt;Darwin-180B-RSI is VIDRAFT's open reasoning model, released in late September 2026. Here are the key technical facts as reported:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Base model:&lt;/strong&gt; Built on top of Qwen3.8-Flash-Next via selective model merging — making it a high-performance model adaptation rather than a ground-up pretrain&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Architecture:&lt;/strong&gt; Mixture-of-Experts (MoE) with &lt;strong&gt;180 billion total parameters&lt;/strong&gt; and &lt;strong&gt;512 experts&lt;/strong&gt;, of which &lt;strong&gt;10 experts are activated per inference request&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Context window:&lt;/strong&gt; Approximately &lt;strong&gt;260,000 tokens&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Availability:&lt;/strong&gt; Described as an open model, with reported leaderboard results hosted on Hugging Face&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The sparse expert activation design (10 of 512 experts per token routing) means that despite the 180B parameter count, the effective compute per forward pass is substantially lower than a dense 180B model — a meaningful practical consideration for deployment cost and latency.&lt;/p&gt;




&lt;h2&gt;
  
  
  How it works
&lt;/h2&gt;

&lt;p&gt;VIDRAFT describes Darwin-180B-RSI as combining two high-level techniques: &lt;strong&gt;selective model merging&lt;/strong&gt; and &lt;strong&gt;recursive self-improvement (RSI)&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;At a conceptual level, the pipeline works roughly like this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Model merging as initialization:&lt;/strong&gt; Rather than training from scratch, VIDRAFT merges capabilities from existing checkpoints selectively — preserving strengths from the base model while injecting targeted improvements.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Recursive self-improvement loop:&lt;/strong&gt; The model generates candidate solutions to reasoning problems. Those solutions are automatically evaluated against verifiable ground-truth answers. Only solutions with confirmed correct answers are retained and used to retrain the model in the next iteration.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Iteration:&lt;/strong&gt; This generate → filter → retrain cycle repeats, theoretically allowing the model to bootstrap higher reasoning quality from its own verified outputs over multiple passes.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This approach is conceptually related to ideas like rejection sampling fine-tuning and self-play reinforcement learning, but VIDRAFT's specific implementation details — including iteration count, reward modeling, and filtering criteria — are not disclosed in the source reporting.&lt;/p&gt;

&lt;p&gt;The MoE routing means that at inference time, each token activates only a small subset of the expert network. This is a well-established technique (see Mixtral, DeepSeek-MoE) for scaling parameter count while keeping per-token FLOPs tractable.&lt;/p&gt;




&lt;h2&gt;
  
  
  Benchmarks &amp;amp; results
&lt;/h2&gt;

&lt;p&gt;VIDRAFT reports the following scores on publicly recognized benchmarks, with Darwin-180B-RSI claiming &lt;strong&gt;first place on five Hugging Face leaderboards&lt;/strong&gt;:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Benchmark&lt;/th&gt;
&lt;th&gt;Reported Score&lt;/th&gt;
&lt;th&gt;Domain&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;AIME 2026&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;100% (perfect)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Mathematical reasoning&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HMMT 2026&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;100% (perfect)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Mathematical competition&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPQA Diamond&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;94.44%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Graduate-level science Q&amp;amp;A&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MMLU-Pro&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;88.12%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Multi-domain professional knowledge&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MMMU-Pro&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;79.48%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Multimodal understanding &amp;amp; reasoning&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;⚠️ &lt;strong&gt;Important caveat:&lt;/strong&gt; These numbers are self-reported by VIDRAFT. The source article explicitly notes that no independent third-party validation of these results has been performed. Leaderboard rankings reflect performance on fixed benchmark datasets and may not generalize to production workloads, novel problem distributions, or domain-specific tasks. Engineers evaluating this model should run it against their own held-out test sets before drawing conclusions about real-world reasoning quality.&lt;/p&gt;




&lt;h2&gt;
  
  
  How to try it
&lt;/h2&gt;

&lt;p&gt;The source reports Darwin-180B-RSI is an &lt;strong&gt;open model available on Hugging Face&lt;/strong&gt;. If VIDRAFT follows its prior release pattern (the company had a prior Hugging Face feature covered in September), the model weights should be downloadable via the standard Hugging Face toolchain.&lt;/p&gt;

&lt;p&gt;To check availability and download weights once the repository is public:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Search for the model on Hugging Face&lt;/span&gt;
huggingface-cli search vidraft/darwin-180b-rsi

&lt;span class="c"&gt;# Download when available&lt;/span&gt;
huggingface-cli download vidraft/darwin-180b-rsi
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Note:&lt;/strong&gt; The exact repository path is not confirmed in the source. Search for &lt;code&gt;VIDRAFT&lt;/code&gt; or &lt;code&gt;Darwin-180B-RSI&lt;/code&gt; directly on &lt;a href="https://huggingface.co" rel="noopener noreferrer"&gt;huggingface.co&lt;/a&gt; to find the correct model card and any associated inference instructions.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;No GitHub repository, API endpoint, or OpenAI-compatible serving URL is mentioned in the source reporting. Check VIDRAFT's Hugging Face profile and official channels for inference API access details.&lt;/p&gt;




&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Q: Is Darwin-180B-RSI practical to self-host given its 180B parameter count?&lt;/strong&gt;&lt;br&gt;
A: The MoE architecture activates only 10 of 512 experts per request, so active-parameter FLOPs are much lower than a dense 180B model. That said, you still need sufficient memory to load all expert weights. Quantized or offloaded serving strategies (e.g., &lt;code&gt;llama.cpp&lt;/code&gt;, &lt;code&gt;vLLM&lt;/code&gt; with expert offloading) may be necessary depending on your hardware. Check the model card for official serving recommendations once it is published.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: How does the recursive self-improvement approach differ from standard RLHF or SFT?&lt;/strong&gt;&lt;br&gt;
A: Standard supervised fine-tuning (SFT) trains on human-labeled data; RLHF uses a human preference reward model. VIDRAFT's RSI loop instead uses the model's own outputs filtered by verifiable correctness — closer in spirit to rejection sampling fine-tuning or outcome-reward RL (like GRPO or STaR), but without requiring a separate reward model or human annotators for each iteration. The practical implication is that it scales more easily to domains with automated answer verification, like math and formal science.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Should I trust the perfect AIME and HMMT scores at face value?&lt;/strong&gt;&lt;br&gt;
A: Treat them as a signal worth investigating, not a guarantee. Perfect scores on competition math benchmarks from self-reported results warrant independent replication. Run the model on problems from your domain and measure it yourself before committing to it for production use cases.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally reported by AI Market Watch (미국) (2026-09-28) — &lt;a href="https://www.ai-market-watch.com/news/vidrafts-recursive-self-improvement-model-tops-five-global-ai-leaderboards-mctpo0" rel="noopener noreferrer"&gt;source article&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>opensource</category>
      <category>llm</category>
    </item>
    <item>
      <title>VIDRAFT Beats Global Big Tech on Five AI Benchmarks Simultaneously</title>
      <dc:creator>AI OpenFree</dc:creator>
      <pubDate>Tue, 29 Sep 2026 03:01:21 +0000</pubDate>
      <link>https://dev.to/ai_openfree_b23025ef075cf/vidraft-beats-global-big-tech-on-five-ai-benchmarks-simultaneously-586h</link>
      <guid>https://dev.to/ai_openfree_b23025ef075cf/vidraft-beats-global-big-tech-on-five-ai-benchmarks-simultaneously-586h</guid>
      <description>&lt;h1&gt;
  
  
  VIDRAFT Beats Global Big Tech on Five AI Benchmarks Simultaneously
&lt;/h1&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; VIDRAFT, a Korean Pre-AGI AI startup, has claimed top positions across five separate AI performance benchmarks, outperforming major global tech companies. This is a notable signal for the broader ML engineering community that competitive frontier AI development is no longer exclusive to the largest Western or East Asian incumbents — and worth tracking if you're evaluating state-of-the-art model performance.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What it is
&lt;/h2&gt;

&lt;p&gt;VIDRAFT (비드래프트) is a Korean AI startup positioning itself in the Pre-AGI space — meaning its research and engineering efforts are oriented toward systems that approach, but have not yet reached, general artificial intelligence. The company's latest models have achieved what the press is calling a "five-crown" (5관왕) performance result: top rankings across five distinct AI capability evaluations, surpassing well-resourced global big tech competitors in each of those categories simultaneously.&lt;/p&gt;

&lt;p&gt;This kind of multi-benchmark dominance is significant because it is relatively rare for any single model or model family to lead across multiple diverse evaluation axes at the same time. Typically, optimization for one benchmark involves trade-offs that hurt performance elsewhere. Achieving five concurrent top positions suggests a degree of generalization and architectural robustness that goes beyond narrow task-specific tuning.&lt;/p&gt;

&lt;h2&gt;
  
  
  How it works
&lt;/h2&gt;

&lt;p&gt;While specific architectural internals and training configurations are not publicly disclosed, the conceptual approach VIDRAFT appears to be pursuing is consistent with the broader frontier AI research direction: building models that generalize across diverse task types rather than being narrowly optimized for a single capability domain.&lt;/p&gt;

&lt;p&gt;At a high level, leading multi-benchmark AI systems typically achieve this through some combination of:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Diverse, high-quality training data curation&lt;/strong&gt; — ensuring coverage across reasoning, language understanding, code, math, and multimodal tasks&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Scalable model architectures&lt;/strong&gt; — designs that maintain performance improvements as compute and parameter counts grow&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Robust evaluation-informed development&lt;/strong&gt; — iterative refinement guided by signal from a wide range of capability assessments, not just a single leaderboard target&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Alignment and instruction-following quality&lt;/strong&gt; — ensuring that raw capability translates into usable, reliable outputs for real-world developer tasks&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The "five benchmark" result implies that VIDRAFT's approach is not over-indexed on any single evaluation, which is a meaningful technical signal about the underlying model's generality.&lt;/p&gt;

&lt;h2&gt;
  
  
  Benchmarks &amp;amp; results
&lt;/h2&gt;

&lt;p&gt;The source article reports that VIDRAFT achieved first-place rankings across &lt;strong&gt;five AI performance benchmarks&lt;/strong&gt;, surpassing global big tech companies across all five simultaneously. The reporting describes this as a "5관왕" (five-crown) outcome — a sweep of top positions in a multi-benchmark context.&lt;/p&gt;

&lt;p&gt;Specific benchmark names, numeric scores, and comparative tables are not provided in the available source material. What the coverage does establish qualitatively:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The competing organizations include major global technology companies with substantially larger engineering headcounts and infrastructure resources&lt;/li&gt;
&lt;li&gt;The result spans multiple evaluation dimensions (not a single repeated metric)&lt;/li&gt;
&lt;li&gt;This was reported by Chosun Biz, one of Korea's major business news outlets, lending it a degree of editorial scrutiny&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;As specific benchmark names and scores become publicly available — for example, through Hugging Face leaderboards, VIDRAFT's own technical blog, or arXiv preprints — those would be the primary sources engineers should consult for reproducible, comparable numbers.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to try it
&lt;/h2&gt;

&lt;p&gt;Based on the information currently available from this source, &lt;strong&gt;public developer access channels (Hugging Face model hub, GitHub repositories, or OpenAI-compatible API endpoints) have not been announced&lt;/strong&gt; in conjunction with this benchmark result.&lt;/p&gt;

&lt;p&gt;If VIDRAFT follows the pattern of other frontier AI labs, future access may come through one or more of:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A Hugging Face organization page for model weights or demos&lt;/li&gt;
&lt;li&gt;An API with OpenAI-compatible endpoints for integration into existing tooling&lt;/li&gt;
&lt;li&gt;A GitHub repository for evaluation code or technical reports&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Engineers interested in following VIDRAFT's releases should monitor their official channels directly. This article will not speculate on endpoints, model names, or access details that have not been publicly confirmed.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Q: How should I interpret "five benchmark crowns" without seeing the specific benchmark names and scores?&lt;/strong&gt;&lt;br&gt;
A: Treat it as a directional signal worth investigating rather than a verified technical claim you can act on immediately. The meaningful next step is to wait for VIDRAFT to publish a technical report or leaderboard entries with reproducible evaluation details — benchmark name, version, evaluation protocol, and scores — before making architectural or tooling decisions based on this result.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Is VIDRAFT's work peer-reviewed or publicly documented in any preprint?&lt;/strong&gt;&lt;br&gt;
A: The current source article does not reference an accompanying arXiv paper, technical report, or public evaluation submission. Until such documentation is available, the result is reported through press coverage only. Engineers should watch for official technical publications before drawing deep conclusions about methodology.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Does this affect my choice of model for production use today?&lt;/strong&gt;&lt;br&gt;
A: Not directly, until public access and reproducible evaluation details are available. However, it is a useful reminder to keep your model evaluation pipeline flexible — leaderboard standings shift frequently, and building your stack to swap foundation models with minimal friction is always good engineering practice.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally reported by 조선비즈 (2026-09-29) — &lt;a href="https://biz.chosun.com/industry/business_info/2026/09/29/74GU4VBMIZEKVD3QCFQCE25MKM/" rel="noopener noreferrer"&gt;source article&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>opensource</category>
      <category>llm</category>
    </item>
    <item>
      <title>Korean startup's open 180B model tops five Hugging Face official leaderboards, including perfect AIME 2026 and HMMT 2026 scores</title>
      <dc:creator>AI OpenFree</dc:creator>
      <pubDate>Mon, 28 Sep 2026 08:03:40 +0000</pubDate>
      <link>https://dev.to/ai_openfree_b23025ef075cf/korean-startups-open-180b-model-tops-five-hugging-face-official-leaderboards-including-perfect-o7l</link>
      <guid>https://dev.to/ai_openfree_b23025ef075cf/korean-startups-open-180b-model-tops-five-hugging-face-official-leaderboards-including-perfect-o7l</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; — Seoul-based AI startup &lt;strong&gt;VIDRAFT&lt;/strong&gt; has released &lt;strong&gt;Darwin-180B-RSI&lt;/strong&gt;, an open-weight 180B mixture-of-experts model that now sits at &lt;strong&gt;#1 on five Hugging Face official benchmark leaderboards&lt;/strong&gt;: AIME 2026 (100%), HMMT February 2026 (100%), GPQA Diamond (94.44%), MMLU-Pro (88.12%) and MMMU-Pro (79.48%). The full weights are downloadable at &lt;a href="https://huggingface.co/FINAL-Bench/Darwin-180B-RSI" rel="noopener noreferrer"&gt;huggingface.co/FINAL-Bench/Darwin-180B-RSI&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What happened
&lt;/h2&gt;

&lt;p&gt;On September 28, 2026, VIDRAFT published Darwin-180B-RSI and submitted its self-measured scores to five benchmark leaderboards that Hugging Face lists as &lt;strong&gt;official benchmarks&lt;/strong&gt; (datasets carrying the &lt;code&gt;benchmark:official&lt;/code&gt; tag). Leaderboard entries on these datasets are populated from &lt;code&gt;.eval_results&lt;/code&gt; files in each model repository.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Benchmark&lt;/th&gt;
&lt;th&gt;Darwin-180B-RSI&lt;/th&gt;
&lt;th&gt;Previous #1&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;AIME 2026&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;100.0&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;97.1 (Inkling, A.X-K2)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HMMT Feb 2026&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;100.0&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;92.7 (Kimi-K2.6)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPQA Diamond&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;94.44&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;93.5 (Kimi-K3)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MMLU-Pro&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;88.12&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;88.0 (MiniMax-M2.1, Intern-S2-Preview)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MMMU-Pro (vision)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;79.48&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;79.4 (Kimi-K2.6)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;According to the public leaderboards, no model had previously reported 100% on the AIME 2026 or HMMT February 2026 boards.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why these five benchmarks matter
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;GPQA Diamond&lt;/strong&gt; — 198 graduate-level physics, chemistry and biology questions written by domain PhDs; experts in the field reach roughly 65%.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MMLU-Pro&lt;/strong&gt; — 12,032 ten-option questions across 14 disciplines, a harder successor to MMLU.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MMMU-Pro (vision)&lt;/strong&gt; — 1,730 college-level questions where the question itself is embedded in an image (charts, scores, medical images) across 30 subjects.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AIME 2026 / HMMT Feb 2026&lt;/strong&gt; — the 2026 American Invitational Mathematics Examination and the Harvard-MIT Mathematics Tournament, both olympiad-track competitions with exact answers.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  How the scores were measured
&lt;/h2&gt;

&lt;p&gt;The model card publishes the protocol. Every run used a &lt;strong&gt;131,072-token thinking budget&lt;/strong&gt;, sampling at temperature 1.0 / top_p 0.95 / top_k 20, in bf16.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Benchmark&lt;/th&gt;
&lt;th&gt;Samples per question&lt;/th&gt;
&lt;th&gt;Reported&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;AIME 2026&lt;/td&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;td&gt;majority vote (mean over 16 = 98.75)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HMMT Feb 2026&lt;/td&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;td&gt;majority vote (mean over 16 = 96.59)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPQA Diamond&lt;/td&gt;
&lt;td&gt;up to 16&lt;/td&gt;
&lt;td&gt;majority vote&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MMLU-Pro&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;single sample&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MMMU-Pro&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;majority vote&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Readers comparing numbers should note that leaderboard values are self-reported by each publisher and that settings (sample count, voting, thinking budget) differ between models.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is inside
&lt;/h2&gt;

&lt;p&gt;The model's parent is &lt;strong&gt;Qwen3.8-Flash-Next&lt;/strong&gt; (180B MoE, Qwen Community License). VIDRAFT says it modified only about &lt;strong&gt;0.02%&lt;/strong&gt; of the parameters — attention paths and shared experts — while leaving all 512 routed experts, the router and the vision encoder unchanged. The company attributes the result to three in-house techniques:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Darwin&lt;/strong&gt; — a model "diagnose-then-evolve" framework (&lt;a href="https://arxiv.org/abs/2605.14386" rel="noopener noreferrer"&gt;arXiv:2605.14386&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;RSI (recursive self-improvement)&lt;/strong&gt; — the model solves practice problems, its answers are checked against verifiable references, and it is retrained on its correct reasoning. VIDRAFT reports the same accuracy with &lt;strong&gt;11% shorter reasoning&lt;/strong&gt; than the parent.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ZTC (Zero-Token Confidence)&lt;/strong&gt; — a readout of the model's internal state that estimates, before generation, the probability that an answer will be correct.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Deep dives
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;&lt;a href="https://dev.to/ai_openfree_b23025ef075cf/perfect-scores-on-aime-2026-and-hmmt-2026-what-an-open-180b-model-got-right-and-why-thinking-4oce"&gt;Perfect scores on AIME 2026 and HMMT 2026 — why thinking budget mattered&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/ai_openfree_b23025ef075cf/darwin-evolving-a-180b-parent-model-by-changing-002-of-it-2913"&gt;Darwin: evolving a 180B parent model by changing 0.02% of it&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/ai_openfree_b23025ef075cf/reading-an-official-hugging-face-leaderboard-protocols-majority-vote-and-reproducibility-2oon"&gt;Reading an official Hugging Face leaderboard: protocols, majority vote and reproducibility&lt;/a&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Is the model open?&lt;/strong&gt; Yes. Weights, config and tokenizer are public on Hugging Face under the Qwen Community License 1.0.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Who is VIDRAFT?&lt;/strong&gt; A Korean AI deep-tech startup (CEO Minsik Kim) that describes itself as an "AI foundry". Its Darwin model family counts 50+ official models and 400+ community derivatives on Hugging Face.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I reproduce the scores?&lt;/strong&gt; The model card lists the full evaluation protocol; all five benchmarks are public datasets.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>opensource</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>Reading an official Hugging Face leaderboard: protocols, majority vote and reproducibility</title>
      <dc:creator>AI OpenFree</dc:creator>
      <pubDate>Mon, 28 Sep 2026 08:02:56 +0000</pubDate>
      <link>https://dev.to/ai_openfree_b23025ef075cf/reading-an-official-hugging-face-leaderboard-protocols-majority-vote-and-reproducibility-2oon</link>
      <guid>https://dev.to/ai_openfree_b23025ef075cf/reading-an-official-hugging-face-leaderboard-protocols-majority-vote-and-reproducibility-2oon</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; — Hugging Face now marks selected benchmark datasets as &lt;strong&gt;official&lt;/strong&gt; and builds their leaderboards from &lt;code&gt;.eval_results/*.yaml&lt;/code&gt; files in model repositories. Scores are self-reported, so the evaluation protocol is what makes a number comparable. This post walks through the protocol behind &lt;strong&gt;Darwin-180B-RSI&lt;/strong&gt;, which ranks #1 on five of these boards.&lt;/p&gt;

&lt;h2&gt;
  
  
  How official leaderboards work
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;A benchmark dataset carries the &lt;code&gt;benchmark:official&lt;/code&gt; tag and an &lt;code&gt;eval.yaml&lt;/code&gt; defining its tasks.&lt;/li&gt;
&lt;li&gt;A model publisher adds a file such as &lt;code&gt;.eval_results/gpqa_diamond.yaml&lt;/code&gt; to their model repository with the dataset id, task id, value and date.&lt;/li&gt;
&lt;li&gt;The dataset page aggregates those files into a ranked leaderboard.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;You can query any board directly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl https://huggingface.co/api/datasets/Idavidrein/gpqa/leaderboard
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Because entries are &lt;strong&gt;self-reported&lt;/strong&gt;, the leaderboard shows the value but not always how it was obtained. The &lt;code&gt;notes&lt;/code&gt; field and the model card are where the protocol lives.&lt;/p&gt;

&lt;h2&gt;
  
  
  The protocol behind five #1 scores
&lt;/h2&gt;

&lt;p&gt;Common settings for every run of Darwin-180B-RSI:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Setting&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Thinking budget&lt;/td&gt;
&lt;td&gt;131,072 tokens per sample&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sampling&lt;/td&gt;
&lt;td&gt;temperature 1.0 · top_p 0.95 · top_k 20&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Precision&lt;/td&gt;
&lt;td&gt;bf16&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Engine&lt;/td&gt;
&lt;td&gt;vLLM, tensor parallel, expert parallel&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Per-benchmark settings and results:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Benchmark&lt;/th&gt;
&lt;th&gt;Samples&lt;/th&gt;
&lt;th&gt;Reported score&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;AIME 2026&lt;/td&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;td&gt;majority vote (mean 98.75)&lt;/td&gt;
&lt;td&gt;100.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HMMT Feb 2026&lt;/td&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;td&gt;majority vote (mean 96.59)&lt;/td&gt;
&lt;td&gt;100.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPQA Diamond&lt;/td&gt;
&lt;td&gt;up to 16&lt;/td&gt;
&lt;td&gt;majority vote&lt;/td&gt;
&lt;td&gt;94.44&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MMLU-Pro&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;single sample&lt;/td&gt;
&lt;td&gt;88.12&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MMMU-Pro (vision)&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;majority vote&lt;/td&gt;
&lt;td&gt;79.48&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Three things to check before comparing numbers
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Samples and voting.&lt;/strong&gt; A majority-of-16 score is a system score and is not directly comparable with a single-sample score. Look for both.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Thinking budget.&lt;/strong&gt; Reasoning models lose points to truncation when the budget is small; a cut-off answer is not a wrong answer.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Contamination.&lt;/strong&gt; Training data should be filtered against every reported test set. VIDRAFT reports an 8-gram overlap filter against all benchmarks it lists.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Reproducing
&lt;/h2&gt;

&lt;p&gt;All five benchmarks are public datasets and the weights are open at &lt;a href="https://huggingface.co/FINAL-Bench/Darwin-180B-RSI" rel="noopener noreferrer"&gt;huggingface.co/FINAL-Bench/Darwin-180B-RSI&lt;/a&gt;. Serving requires 8 × B200-class GPUs (4 at minimum) in bf16:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;vllm serve FINAL-Bench/Darwin-180B-RSI &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--tensor-parallel-size&lt;/span&gt; 8 &lt;span class="nt"&gt;--enable-expert-parallel&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--max-model-len&lt;/span&gt; 135168 &lt;span class="nt"&gt;--trust-remote-code&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



</description>
      <category>ai</category>
      <category>llm</category>
      <category>datascience</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Darwin: evolving a 180B parent model by changing 0.02% of it</title>
      <dc:creator>AI OpenFree</dc:creator>
      <pubDate>Mon, 28 Sep 2026 08:02:19 +0000</pubDate>
      <link>https://dev.to/ai_openfree_b23025ef075cf/darwin-evolving-a-180b-parent-model-by-changing-002-of-it-2913</link>
      <guid>https://dev.to/ai_openfree_b23025ef075cf/darwin-evolving-a-180b-parent-model-by-changing-002-of-it-2913</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; — Darwin-180B-RSI keeps its parent's 512 routed experts, router and vision encoder intact and modifies about &lt;strong&gt;0.02%&lt;/strong&gt; of the parameters. VIDRAFT's Darwin framework treats an open model as a "parent", diagnoses its weak paths, and strengthens only those — here combined with recursive self-improvement (RSI) on verified answers.&lt;/p&gt;

&lt;h2&gt;
  
  
  The parent
&lt;/h2&gt;

&lt;p&gt;The father model is &lt;strong&gt;Qwen3.8-Flash-Next&lt;/strong&gt;, a 180B mixture-of-experts vision-language model:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Layers / hidden size&lt;/td&gt;
&lt;td&gt;48 / 2,560&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Attention&lt;/td&gt;
&lt;td&gt;36 linear-attention + 12 full-attention layers&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Experts&lt;/td&gt;
&lt;td&gt;512 routed (10 active per token) + shared expert&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Context&lt;/td&gt;
&lt;td&gt;262,144 tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  What Darwin changed
&lt;/h2&gt;

&lt;p&gt;According to the model card, Darwin updated three wiring paths and preserved the rest:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;🔹 &lt;strong&gt;Full-attention layers (12)&lt;/strong&gt; — the long-range path across the whole context&lt;/li&gt;
&lt;li&gt;🔹 &lt;strong&gt;Linear-attention layers (36)&lt;/strong&gt; — the fast inference path&lt;/li&gt;
&lt;li&gt;🔹 &lt;strong&gt;Shared expert (48 layers)&lt;/strong&gt; — the common path every token passes through&lt;/li&gt;
&lt;li&gt;🔒 &lt;strong&gt;Preserved&lt;/strong&gt;: all 512 routed experts, the router and the vision encoder&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is roughly 30 million trainable parameters out of 177 billion. Because the experts — where most factual knowledge is stored in an MoE model — are untouched, the parent's knowledge is retained.&lt;/p&gt;

&lt;h2&gt;
  
  
  Recursive self-improvement on verified answers
&lt;/h2&gt;

&lt;p&gt;The RSI loop is simple to state:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The model solves practice problems that do not overlap with any reported benchmark (8-gram overlap filter).&lt;/li&gt;
&lt;li&gt;Its answers are checked against verifiable references.&lt;/li&gt;
&lt;li&gt;Only the correct reasoning is kept, and the model is retrained on it.&lt;/li&gt;
&lt;li&gt;The improved model becomes the next solver.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Two design choices matter. Unverified solutions are never learned, so errors are not reinforced. And among correct solutions, shorter ones are preferred — which is why VIDRAFT reports &lt;strong&gt;the same accuracy with 11% shorter reasoning&lt;/strong&gt; than the parent (MMLU-Pro: 4,320 → 3,833 tokens per answer).&lt;/p&gt;

&lt;h2&gt;
  
  
  How Darwin compares with evolutionary model merging
&lt;/h2&gt;

&lt;p&gt;The best-known prior work in this area is Sakana AI's evolutionary model merge (Nature Machine Intelligence, 2025), which searches merge recipes over hundreds of generations and demonstrated 7–10B models. Darwin takes a diagnosis-first approach — measuring layer- and expert-level strengths before transplanting or tuning — which VIDRAFT says reduces search cost and has scaled from 4B to 397B models.&lt;/p&gt;

&lt;p&gt;On Hugging Face, the Darwin family lists &lt;strong&gt;50+ official models&lt;/strong&gt; and &lt;strong&gt;400+ community derivatives&lt;/strong&gt;. The framework is described in &lt;a href="https://arxiv.org/abs/2605.14386" rel="noopener noreferrer"&gt;Darwin Family (arXiv:2605.14386)&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Zero-Token Confidence
&lt;/h2&gt;

&lt;p&gt;VIDRAFT pairs RSI with &lt;strong&gt;ZTC&lt;/strong&gt;, a lightweight readout of the model's final-layer hidden state that estimates, before any token is generated, the probability that the upcoming answer is correct. In an RSI loop, low-confidence areas point to what the model should learn next; in deployment, the same signal can gate tool calls or escalate to a human. The company's earlier &lt;a href="https://huggingface.co/FINAL-Bench/Darwin-397B-ZTC" rel="noopener noreferrer"&gt;Darwin-397B-ZTC&lt;/a&gt; ships such a readout.&lt;/p&gt;

&lt;h2&gt;
  
  
  Links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Model: &lt;a href="https://huggingface.co/FINAL-Bench/Darwin-180B-RSI" rel="noopener noreferrer"&gt;FINAL-Bench/Darwin-180B-RSI&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Darwin collection: &lt;a href="https://huggingface.co/collections/FINAL-Bench/darwin-family" rel="noopener noreferrer"&gt;huggingface.co/collections/FINAL-Bench/darwin-family&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>machinelearning</category>
      <category>llm</category>
      <category>opensource</category>
      <category>ai</category>
    </item>
    <item>
      <title>Perfect scores on AIME 2026 and HMMT 2026: what an open 180B model got right, and why thinking budget mattered</title>
      <dc:creator>AI OpenFree</dc:creator>
      <pubDate>Mon, 28 Sep 2026 08:01:42 +0000</pubDate>
      <link>https://dev.to/ai_openfree_b23025ef075cf/perfect-scores-on-aime-2026-and-hmmt-2026-what-an-open-180b-model-got-right-and-why-thinking-4oce</link>
      <guid>https://dev.to/ai_openfree_b23025ef075cf/perfect-scores-on-aime-2026-and-hmmt-2026-what-an-open-180b-model-got-right-and-why-thinking-4oce</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; — Darwin-180B-RSI, an open-weight model from Korean startup VIDRAFT, reports &lt;strong&gt;100% on AIME 2026 (30/30)&lt;/strong&gt; and &lt;strong&gt;100% on HMMT February 2026 (33/33)&lt;/strong&gt; under 16-sample majority vote with a 131K-token thinking budget. They are the first perfect scores on these two Hugging Face official leaderboards.&lt;/p&gt;

&lt;h2&gt;
  
  
  The numbers
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Benchmark&lt;/th&gt;
&lt;th&gt;Majority of 16&lt;/th&gt;
&lt;th&gt;Mean over 16 samples&lt;/th&gt;
&lt;th&gt;Single sample&lt;/th&gt;
&lt;th&gt;Previous #1&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;AIME 2026&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;100.0&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;98.75&lt;/td&gt;
&lt;td&gt;96.67&lt;/td&gt;
&lt;td&gt;97.1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HMMT Feb 2026&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;100.0&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;96.59&lt;/td&gt;
&lt;td&gt;93.94&lt;/td&gt;
&lt;td&gt;92.7&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Even the &lt;strong&gt;mean&lt;/strong&gt; accuracy over 16 samples (98.75 on AIME, 96.59 on HMMT) is above the previous top entries on both boards.&lt;/p&gt;

&lt;h2&gt;
  
  
  About the two competitions
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;AIME&lt;/strong&gt; (American Invitational Mathematics Examination) is the qualifier toward the USA Mathematical Olympiad. Every answer is an integer from 0 to 999, so guessing rarely works.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;HMMT&lt;/strong&gt; (Harvard-MIT Mathematics Tournament) is one of the most difficult high-school competitions, attended by national-olympiad-level students. Answers include fractions and radicals, so grading requires symbolic equivalence (for example, 2√3 = √12).&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The finding: misses were truncations, not wrong math
&lt;/h2&gt;

&lt;p&gt;VIDRAFT's engineers first measured the parent model with a 65K-token budget. The model missed two AIME 2026 problems, and on inspection &lt;strong&gt;every miss was a truncated solution&lt;/strong&gt; — the reasoning ran out of tokens before reaching an answer. Solutions that finished were correct.&lt;/p&gt;

&lt;p&gt;Re-running those problems with a &lt;strong&gt;131K-token budget&lt;/strong&gt; produced 8/8 correct answers. With the budget raised for all problems, the 16-sample run scored 30/30.&lt;/p&gt;

&lt;p&gt;The practical lesson for anyone evaluating reasoning models:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Truncation is not an error.&lt;/strong&gt; Count cut-off answers separately; otherwise a budget limit looks like a capability limit.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Budget is part of the protocol.&lt;/strong&gt; A score without its thinking budget is not comparable.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Grade symbolically.&lt;/strong&gt; String matching undercounts fractions and radicals; a math-equivalence checker is required.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  How to read "majority of 16"
&lt;/h2&gt;

&lt;p&gt;Majority vote is a &lt;em&gt;system&lt;/em&gt; score: the model answers each question 16 times and the most frequent answer is graded. VIDRAFT reports the majority score and the mean single-sample accuracy side by side on the model card and in the leaderboard submission notes, so readers can compare with entries that use a single sample.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;

&lt;p&gt;Model: &lt;a href="https://huggingface.co/FINAL-Bench/Darwin-180B-RSI" rel="noopener noreferrer"&gt;huggingface.co/FINAL-Bench/Darwin-180B-RSI&lt;/a&gt; — the model card contains the full evaluation protocol.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>math</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>South Korea's Government–Naver Consortium Is Building Dual 700B-Class Cybersecurity-Specialized AI Models with 4,512 GPUs</title>
      <dc:creator>AI OpenFree</dc:creator>
      <pubDate>Wed, 23 Sep 2026 07:01:09 +0000</pubDate>
      <link>https://dev.to/ai_openfree_b23025ef075cf/south-koreas-government-naver-consortium-is-building-dual-700b-class-cybersecurity-specialized-ai-1pa0</link>
      <guid>https://dev.to/ai_openfree_b23025ef075cf/south-koreas-government-naver-consortium-is-building-dual-700b-class-cybersecurity-specialized-ai-1pa0</guid>
      <description>&lt;h1&gt;
  
  
  South Korea's Government–Naver Consortium Is Building Dual 700B-Class Cybersecurity-Specialized AI Models with 4,512 GPUs
&lt;/h1&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; A South Korean government–Naver Cloud consortium has officially launched development of two frontier-scale, cybersecurity-specialized large language models — one focused on offensive (red-team) capabilities and one on defensive (blue-team) capabilities — backed by 4,512 GPUs and ~830 TB of security-domain training data. The models are planned for open-source release and potential international export, making them worth watching for security engineers and ML practitioners building in the threat-intelligence and vulnerability-research space.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What it is
&lt;/h2&gt;

&lt;p&gt;The South Korean Ministry of Science and ICT (MSIT), in consortium with Naver Cloud, LG AI Research, and 41 partner organizations, has kicked off a &lt;strong&gt;cybersecurity-specialized AI model development program&lt;/strong&gt;. The official launch event was held on September 22, 2026, in Seoul.&lt;/p&gt;

&lt;p&gt;Key facts from the source:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Two distinct 700B-class models&lt;/strong&gt; are being developed in parallel:

&lt;ul&gt;
&lt;li&gt;A &lt;strong&gt;red-team (offensive) model&lt;/strong&gt; — focused on source code analysis, automated vulnerability discovery, root-cause inference, and attack-path reasoning.&lt;/li&gt;
&lt;li&gt;A &lt;strong&gt;blue-team (defensive) model&lt;/strong&gt; — focused on correlating distributed system events and threat intelligence, classifying attack behavior, assessing threat severity, and recommending response actions.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;The models are &lt;strong&gt;security-domain fine-tunes&lt;/strong&gt; of existing foundation models: Naver Cloud is building on HyperCLOVA X for the defensive variant; LG AI Research is contributing an offensive variant based on its own foundation model.&lt;/li&gt;
&lt;li&gt;Training data: approximately &lt;strong&gt;830 TB of security-domain data&lt;/strong&gt;, sourced from national critical infrastructure and industrial security environments, with a dedicated preprocessing pipeline to convert raw source data into training-ready format.&lt;/li&gt;
&lt;li&gt;The program includes &lt;strong&gt;field validation&lt;/strong&gt; across sectors including finance and defense.&lt;/li&gt;
&lt;li&gt;Planned &lt;strong&gt;open-source release&lt;/strong&gt; with international distribution ambitions.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  How it works
&lt;/h2&gt;

&lt;p&gt;The architectural concept here is a &lt;strong&gt;dual-model adversarial co-training loop&lt;/strong&gt; — sometimes called a "spear-and-shield" or red/blue feedback cycle:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Red-team model&lt;/strong&gt; analyzes code and system artifacts to surface vulnerabilities, classify their type and origin, and reason about exploitable attack chains.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Blue-team model&lt;/strong&gt; ingests multi-source event logs and threat feeds to reconstruct attack behavior, score threat severity, and propose mitigations.&lt;/li&gt;
&lt;li&gt;Critically, &lt;strong&gt;the two models inform each other's training&lt;/strong&gt;: vulnerabilities found by the red-team model become labeled training signal for the blue-team model; successful defenses by the blue-team model raise the bar the red-team model must clear in future iterations.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This adversarial feedback loop is conceptually analogous to self-play in game-theoretic AI systems, applied here to the offensive/defensive security domain at foundation-model scale. The data preprocessing pipeline handles conversion of raw operational security data — including real incident records — into formats suitable for supervised and instruction-tuning workflows.&lt;/p&gt;

&lt;h2&gt;
  
  
  Benchmarks &amp;amp; results
&lt;/h2&gt;

&lt;p&gt;The project is at &lt;strong&gt;launch/kickoff stage&lt;/strong&gt; (September 2026), so no public benchmark results are available yet. The source article does not report evaluation scores, leaderboard placements, or ablation outcomes.&lt;/p&gt;

&lt;p&gt;What is publicly stated:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The Deputy Prime Minister and MSIT Minister described the ambition as reaching &lt;strong&gt;"frontier-level" performance&lt;/strong&gt; in the security-specialized domain.&lt;/li&gt;
&lt;li&gt;A &lt;strong&gt;mid-program evaluation&lt;/strong&gt; is scheduled for February 2027, at which point continued GPU resource allocation from the government side will be decided.&lt;/li&gt;
&lt;li&gt;Field validation ("실증") will be conducted against real national infrastructure and industrial environments — implying evaluation on operational security tasks rather than purely academic benchmarks.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Engineers should watch for public benchmark disclosures around or after the February 2027 checkpoint.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to try it
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Public developer access is not available at this time.&lt;/strong&gt; The program launched on September 22, 2026, and the models are still in active development. The source article states that open-source release is planned, but no Hugging Face repository, GitHub organization, API endpoint, or release timeline has been announced publicly.&lt;/p&gt;

&lt;p&gt;When the models are released, expect distribution channels typical of Korean government-backed open-source AI initiatives (Hugging Face Hub, potential OpenAI-compatible API wrappers). Follow Naver Cloud and LG AI Research's official channels for release announcements.&lt;/p&gt;




&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Q: Why develop two separate models instead of one unified security model?&lt;/strong&gt;&lt;br&gt;
A: The red/blue split allows each model to specialize deeply — offensive reasoning (exploit path analysis, vulnerability classification) and defensive reasoning (event correlation, threat scoring, response recommendation) have distinct input distributions and output requirements. More importantly, keeping them separate enables the adversarial co-training loop: each model's outputs become harder training targets for the other, driving iterative capability improvement on both sides.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Is the ~830 TB training dataset going to be publicly released alongside the models?&lt;/strong&gt;&lt;br&gt;
A: The source article does not state this. Given that much of the data originates from national critical infrastructure and live security environments, full public release of the raw dataset seems unlikely, though curated or preprocessed subsets could accompany an open-source model release. No official announcement on dataset availability has been made.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: What is the GPU breakdown across consortium partners?&lt;/strong&gt;&lt;br&gt;
A: Naver Cloud self-funded 4,000 GPUs (NVIDIA B200) starting August 2026, LG AI Research is contributing 256 GPUs (NVIDIA H200), and the Korean government is supplying 256 GPUs (NVIDIA B200) over a 10-month support window — totaling 4,512 GPUs across the consortium.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally reported by EBN (2026-09-22) — &lt;a href="https://www.ebn.co.kr/news/articleView.html?idxno=1725429" rel="noopener noreferrer"&gt;source article&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>opensource</category>
      <category>llm</category>
    </item>
    <item>
      <title>JEV vs ZTC: Statistically Tied on Accuracy, 1.4 Points Apart Where It Counts</title>
      <dc:creator>AI OpenFree</dc:creator>
      <pubDate>Sun, 20 Sep 2026 10:09:29 +0000</pubDate>
      <link>https://dev.to/ai_openfree_b23025ef075cf/jev-vs-ztc-statistically-tied-on-accuracy-14-points-apart-where-it-counts-4gc9</link>
      <guid>https://dev.to/ai_openfree_b23025ef075cf/jev-vs-ztc-statistically-tied-on-accuracy-14-points-apart-where-it-counts-4gc9</guid>
      <description>&lt;p&gt;&lt;strong&gt;Short answer: on accuracy, JEV (0.7350 AUC) and ZTC (0.7364) are statistically tied — the 0.0014&lt;br&gt;
gap has a 95% interval of −0.019 to +0.032. They separate on two other axes: deployment shape (API&lt;br&gt;
only vs open weights) and downstream agent accuracy, where ZTC gained 1.34 points and JEV lost 0.07&lt;br&gt;
at the same retry budget.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Both numbers come from the same test set: 2,018 items, 508 of them incorrect, identical labels,&lt;br&gt;
published grading code —&lt;br&gt;
&lt;a href="https://huggingface.co/spaces/mayafree/typed-decision-leaderboard" rel="noopener noreferrer"&gt;Typed Decision Leaderboard&lt;/a&gt;.&lt;/p&gt;
&lt;h2&gt;
  
  
  First, what these two things are
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;JEV&lt;/strong&gt; is TypeSafe AI's answer verifier, the product that established this category under the name&lt;br&gt;
&lt;em&gt;System One models&lt;/em&gt;. &lt;strong&gt;ZTC&lt;/strong&gt; (Zero-Token Confidence) is VIDRAFT's, shipping as two open-weight&lt;br&gt;
models.&lt;/p&gt;

&lt;p&gt;Both are &lt;strong&gt;typed decision models&lt;/strong&gt;: they return a structured verdict — boolean, choice, or score —&lt;br&gt;
from a single forward pass, emitting &lt;strong&gt;zero generated tokens&lt;/strong&gt;. That is what makes both of them one&lt;br&gt;
to three orders of magnitude cheaper than asking a frontier model "is this answer right?"&lt;/p&gt;

&lt;p&gt;So the choice between them is not a choice of architecture. It is a choice of accuracy, shape, and&lt;br&gt;
behaviour under load.&lt;/p&gt;
&lt;h2&gt;
  
  
  Axis 1 — accuracy: a tie, and it should be reported as one
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;System&lt;/th&gt;
&lt;th&gt;AUC&lt;/th&gt;
&lt;th&gt;95% interval vs. the other&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;ZTC (397B)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.7364&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;JEV&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.7350&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;−0.019 to +0.032&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ZTC-Judge-27B&lt;/td&gt;
&lt;td&gt;0.7282&lt;/td&gt;
&lt;td&gt;−0.034 to +0.020&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The interval contains zero in both comparisons. Under the rule that a gap indistinguishable from&lt;br&gt;
noise produces no rank, &lt;strong&gt;neither system can be called more accurate than the other on this set.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;For context, the same table's reference row — a logistic model over answer length and formatting —&lt;br&gt;
scores &lt;strong&gt;0.7036&lt;/strong&gt;, and eight of the thirteen measured systems fall below it. Both JEV and ZTC are&lt;br&gt;
comfortably above that bar. Most of the category is not.&lt;/p&gt;
&lt;h2&gt;
  
  
  Axis 2 — deployment shape, where they genuinely differ
&lt;/h2&gt;

&lt;p&gt;This is the axis that usually decides the purchase, and it is not close.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;JEV&lt;/th&gt;
&lt;th&gt;ZTC&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Weights published&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;no&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;yes&lt;/strong&gt; (Apache-2.0)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Runs inside a private network&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;yes&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Adaptable to your own data&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;yes&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hosted API&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;yes&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$0.024 / 1,000 calls&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;your own compute&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Generated tokens&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;If your data cannot leave your network, the accuracy tie is irrelevant — only one of these is&lt;br&gt;
installable. If you would rather not operate a GPU, the same logic runs the other way, and $0.024&lt;br&gt;
per thousand calls is cheap enough that "just use the API" is a defensible answer.&lt;/p&gt;

&lt;p&gt;For comparison, asking GPT-5.2 the same question costs roughly &lt;strong&gt;$0.55 per 1,000 calls&lt;/strong&gt; — about&lt;br&gt;
&lt;strong&gt;23× more&lt;/strong&gt; — and scores &lt;strong&gt;0.7148&lt;/strong&gt;, below both.&lt;/p&gt;
&lt;h2&gt;
  
  
  Axis 3 — per domain, the ranking inverts
&lt;/h2&gt;

&lt;p&gt;The headline is an average of five domains. They disagree.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Domain&lt;/th&gt;
&lt;th&gt;n&lt;/th&gt;
&lt;th&gt;JEV&lt;/th&gt;
&lt;th&gt;ZTC 397B&lt;/th&gt;
&lt;th&gt;ZTC 27B&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Professional exams (law/math/bio)&lt;/td&gt;
&lt;td&gt;400&lt;/td&gt;
&lt;td&gt;0.7981&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.8660&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;0.8462&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Biology &amp;amp; medicine&lt;/td&gt;
&lt;td&gt;917&lt;/td&gt;
&lt;td&gt;0.7185&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.7433&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;0.7154&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Commonsense &amp;amp; multi-step&lt;/td&gt;
&lt;td&gt;278&lt;/td&gt;
&lt;td&gt;0.6006&lt;/td&gt;
&lt;td&gt;0.6072&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.6172&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Disaster &amp;amp; safety guidance&lt;/td&gt;
&lt;td&gt;225&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.7620&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;0.7319&lt;/td&gt;
&lt;td&gt;0.6961&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Scientific reasoning&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;198&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.8421&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;0.6287&lt;/td&gt;
&lt;td&gt;0.7410&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;JEV wins scientific reasoning outright at 0.8421 — 0.21 ahead of the overall leader.&lt;/strong&gt; If that is&lt;br&gt;
your workload, the overall ranking is the wrong table to read, and the answer to "which one" is JEV.&lt;/p&gt;

&lt;p&gt;Two other things in that table are worth carrying away regardless of vendor:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bigger is not better.&lt;/strong&gt; Inside one family, the 397B scores 0.11 &lt;em&gt;below&lt;/em&gt; the 27B on scientific&lt;br&gt;
reasoning. A separate measurement puts a 180B-class model at 0.7146, under a 4B at 0.7284.&lt;br&gt;
Verification quality tracks representation geometry, not parameter count.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Commonsense is unsolved by everyone.&lt;/strong&gt; Every system sits near 0.60 on multi-step commonsense. No&lt;br&gt;
product on this board fixes that domain, and knowing it is more useful than any ranking.&lt;/p&gt;
&lt;h2&gt;
  
  
  Axis 4 — as an agent gate, they are not tied at all
&lt;/h2&gt;

&lt;p&gt;AUC answers &lt;em&gt;"which separates right from wrong answers best."&lt;/em&gt; It does not answer &lt;em&gt;"which helps my&lt;br&gt;
agent."&lt;/em&gt; Those turned out to be different questions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Setup.&lt;/strong&gt; For each item: score the answer with a verifier, send the lowest-scoring 20% to a&lt;br&gt;
stronger model to be re-answered, keep the rest. All arms draw from one shared pool of re-answers,&lt;br&gt;
so no arm gets luckier retries. Ties broken randomly, averaged over &lt;strong&gt;200 seeds&lt;/strong&gt;.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Gate&lt;/th&gt;
&lt;th&gt;Final accuracy&lt;/th&gt;
&lt;th&gt;vs. no gate (74.83%)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;ZTC&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;76.16%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;+1.34 pp&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;JEV&lt;/td&gt;
&lt;td&gt;74.76%&lt;/td&gt;
&lt;td&gt;−0.07 pp&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Random&lt;/td&gt;
&lt;td&gt;74.58%&lt;/td&gt;
&lt;td&gt;−0.25 pp&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;A 0.0014 AUC gap became a 1.4-point swing in end-to-end accuracy.&lt;/strong&gt; At this budget, JEV's gate&lt;br&gt;
performed about the same as no gate at all.&lt;/p&gt;
&lt;h3&gt;
  
  
  The mechanism, which is the transferable part
&lt;/h3&gt;

&lt;p&gt;Re-answering is double-edged:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;wrong answers sent back  →  38% get fixed
right answers sent back  →  30% get broken
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So a gate is not rewarded for catching errors. It is rewarded for &lt;strong&gt;not disturbing what was already&lt;br&gt;
right&lt;/strong&gt;. Precision, not recall — the opposite of how most people reason about a safety net.&lt;/p&gt;

&lt;p&gt;At a 20% budget both systems routed &lt;strong&gt;403 items&lt;/strong&gt;. JEV's 403 contained &lt;strong&gt;216 that were already&lt;br&gt;
correct&lt;/strong&gt;; ZTC's contained &lt;strong&gt;195&lt;/strong&gt;. Twenty-one items of difference in &lt;em&gt;what got sent back&lt;/em&gt;, and one&lt;br&gt;
arm spent its whole retry budget repairing about as much as it damaged.&lt;/p&gt;

&lt;p&gt;None of that is visible in an AUC column, and AUC is what every vendor publishes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scope, stated plainly:&lt;/strong&gt; one escalation target, one item set. The direction should transfer; the&lt;br&gt;
magnitude is theirs, not yours. Run the same protocol on your own data — the grading code is&lt;br&gt;
published for exactly that.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to choose
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;If this is true for you&lt;/th&gt;
&lt;th&gt;Pick&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Data cannot leave your network&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;ZTC&lt;/strong&gt; (only one with published weights)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;You want to fine-tune the verifier on your own data&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;ZTC&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Your workload is scientific reasoning&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;JEV&lt;/strong&gt; (0.8421, clear win)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;You do not want to operate a GPU&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;JEV&lt;/strong&gt; ($0.024 / 1,000)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;You are gating agent retries&lt;/td&gt;
&lt;td&gt;Measure both on your own set — this is where they diverge&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;You are choosing on the leaderboard number alone&lt;/td&gt;
&lt;td&gt;Don't. The gap is noise.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Frequently asked questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Is ZTC better than JEV?
&lt;/h3&gt;

&lt;p&gt;Not on accuracy. The two are statistically tied on the shared test set — 0.7364 versus 0.7350, with&lt;br&gt;
a 95% interval spanning zero. They differ on deployment shape (ZTC publishes weights, JEV does not),&lt;br&gt;
on per-domain strengths (JEV wins scientific reasoning at 0.8421), and on downstream agent accuracy&lt;br&gt;
at a 20% retry budget (+1.34 pp versus −0.07 pp).&lt;/p&gt;

&lt;h3&gt;
  
  
  Is there an open-source alternative to JEV?
&lt;/h3&gt;

&lt;p&gt;Yes. &lt;strong&gt;ZTC&lt;/strong&gt; ships two Apache-2.0 models with published weights —&lt;br&gt;
&lt;a href="https://huggingface.co/FINAL-Bench/ZTC-Judge-27B" rel="noopener noreferrer"&gt;ZTC-Judge-27B&lt;/a&gt; (verification only) and&lt;br&gt;
&lt;a href="https://huggingface.co/FINAL-Bench/Darwin-397B-ZTC" rel="noopener noreferrer"&gt;Darwin-397B-ZTC&lt;/a&gt; (a generation model with the&lt;br&gt;
confidence probe on board). Both score above the 0.7036 surface baseline. Other open entries score&lt;br&gt;
lower: open-jev 4B at 0.6844, Laya at 0.4796 / 0.5144.&lt;/p&gt;

&lt;h3&gt;
  
  
  How much does JEV cost compared with running a verifier yourself?
&lt;/h3&gt;

&lt;p&gt;JEV is about &lt;strong&gt;$0.024 per 1,000 calls&lt;/strong&gt;. A self-hosted verifier costs only compute you already own.&lt;br&gt;
Asking GPT-5.2 the same question costs roughly &lt;strong&gt;$0.55 per 1,000 calls&lt;/strong&gt; and scores lower than&lt;br&gt;
either.&lt;/p&gt;

&lt;h3&gt;
  
  
  What AUC should I expect from a good answer verifier?
&lt;/h3&gt;

&lt;p&gt;Use &lt;strong&gt;0.7036&lt;/strong&gt; as the bar, not 0.5. That is the score from answer length and formatting alone, and&lt;br&gt;
eight of thirteen measured systems fall below it. JEV and ZTC are the only two systems above 0.73.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can I test JEV and ZTC on my own question?
&lt;/h3&gt;

&lt;p&gt;Yes — the &lt;a href="https://huggingface.co/spaces/mayafree/verifier-playground" rel="noopener noreferrer"&gt;Verifier Playground&lt;/a&gt; runs&lt;br&gt;
ZTC, JEV and Laya on the same input, in the browser, no signup. One case cannot rank them, but it&lt;br&gt;
shows you how each behaves on your phrasing.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why did the more accurate verifier perform worse as a gate?
&lt;/h3&gt;

&lt;p&gt;Because gating rewards precision rather than recall. Sending a correct answer back to be re-answered&lt;br&gt;
breaks it about 30% of the time, so routing a few more already-correct items cancels out the errors&lt;br&gt;
you caught. JEV routed 216 already-correct items; ZTC routed 195.&lt;/p&gt;

&lt;h2&gt;
  
  
  Links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Leaderboard&lt;/strong&gt; (13 systems, 2,018 items) —
&lt;a href="https://huggingface.co/spaces/mayafree/typed-decision-leaderboard" rel="noopener noreferrer"&gt;https://huggingface.co/spaces/mayafree/typed-decision-leaderboard&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Playground&lt;/strong&gt; — &lt;a href="https://huggingface.co/spaces/mayafree/verifier-playground" rel="noopener noreferrer"&gt;https://huggingface.co/spaces/mayafree/verifier-playground&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Method write-up&lt;/strong&gt; — &lt;a href="https://huggingface.co/blog/mayafree/jve-ecosystems" rel="noopener noreferrer"&gt;https://huggingface.co/blog/mayafree/jve-ecosystems&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ZTC-Judge-27B&lt;/strong&gt; — &lt;a href="https://huggingface.co/FINAL-Bench/ZTC-Judge-27B" rel="noopener noreferrer"&gt;https://huggingface.co/FINAL-Bench/ZTC-Judge-27B&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Darwin-397B-ZTC&lt;/strong&gt; — &lt;a href="https://huggingface.co/FINAL-Bench/Darwin-397B-ZTC" rel="noopener noreferrer"&gt;https://huggingface.co/FINAL-Bench/Darwin-397B-ZTC&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
    </item>
  </channel>
</rss>
