<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: AI OpenFree</title>
    <description>The latest articles on DEV Community by AI OpenFree (@ai_openfree_b23025ef075cf).</description>
    <link>https://dev.to/ai_openfree_b23025ef075cf</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3817626%2F25ccf2de-8934-44fe-9d10-59dd6d2a505b.png</url>
      <title>DEV Community: AI OpenFree</title>
      <link>https://dev.to/ai_openfree_b23025ef075cf</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ai_openfree_b23025ef075cf"/>
    <language>en</language>
    <item>
      <title>VIDRAFT's PharmaOS Lands Government-Pharma Program to Drive AI-Discovered Small-Molecule Cancer Drug Candidates to Preclinical</title>
      <dc:creator>AI OpenFree</dc:creator>
      <pubDate>Tue, 08 Sep 2026 03:01:19 +0000</pubDate>
      <link>https://dev.to/ai_openfree_b23025ef075cf/vidrafts-pharmaos-lands-government-pharma-program-to-drive-ai-discovered-small-molecule-cancer-2ah5</link>
      <guid>https://dev.to/ai_openfree_b23025ef075cf/vidrafts-pharmaos-lands-government-pharma-program-to-drive-ai-discovered-small-molecule-cancer-2ah5</guid>
      <description>&lt;h1&gt;
  
  
  VIDRAFT's PharmaOS Lands Government-Pharma Program to Drive AI-Discovered Small-Molecule Cancer Drug Candidates to Preclinical
&lt;/h1&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; VIDRAFT, a Korean AI-foundry deep-tech company, has been selected for the South Korean government's "Challenge Bio" open-innovation program to discover and validate small-molecule anticancer drug candidates using its PharmaOS platform — running from hit identification through preclinical linkage, with IP acquisition as an explicit goal. The platform has previously ranked #1 across 16 categories on the global Polaris drug-prediction benchmark. Developers and researchers can already interact with VIDRAFT's open drug-discovery ecosystem on Hugging Face.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  What it is
&lt;/h2&gt;

&lt;p&gt;VIDRAFT (비드래프트) is a Korean "AI Foundry" startup — its business model centers on designing, producing, and auditing AI models for industrial deployment, rather than being a pure software SaaS play. The core drug-discovery product is &lt;strong&gt;PharmaOS&lt;/strong&gt;, a computational platform for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Molecular structure exploration&lt;/strong&gt; — searching chemical space for candidate compounds&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Target-protein binding analysis&lt;/strong&gt; — estimating likelihood of binding to a specified protein target&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Drug-likeness and toxicity risk assessment&lt;/strong&gt; — filtering candidates before wet-lab work begins&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Explainable candidate prioritization&lt;/strong&gt; — the platform emphasizes &lt;em&gt;why&lt;/em&gt; a compound advances, not just that it scored highly&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The broader research stack also includes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;QuantumOS&lt;/strong&gt; — a quantum-classical hybrid computation software layer being researched for molecular-level analysis&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CELLOS&lt;/strong&gt; — a cell and organ simulation system under active development, designed to model cellular metabolism and inter-organ interactions, providing hypotheses for candidate comparison and downstream experiment design&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Together, these form an automated, iterative discovery-and-validation loop: AI proposes candidates → quantum/classical hybrid layer refines molecular analysis → cell/organ simulation flags biological risks → findings feed back into molecular design.&lt;/p&gt;




&lt;h2&gt;
  
  
  How it works
&lt;/h2&gt;

&lt;p&gt;PharmaOS operates as a multi-stage computational funnel rather than a single-score ranker:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Generation&lt;/strong&gt; — The system explores molecular structure space to produce a large pool of candidate compounds against a specified oncology target.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Binding estimation&lt;/strong&gt; — Computational docking and affinity models assess each candidate's likelihood of interacting with the target protein.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ADMET profiling&lt;/strong&gt; — Drug-likeness, absorption, distribution, metabolism, excretion, and toxicity risks are estimated in silico.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rationale generation&lt;/strong&gt; — Rather than outputting only a ranked list, PharmaOS produces documented reasoning for why specific candidates merit wet-lab follow-up — an important practical feature when handing off to pharma partners.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Iterative refinement&lt;/strong&gt; — Results from each computational stage feed back into the design loop, with QuantumOS research exploring additional molecular-level verification and CELLOS modeling biological responses at the cellular and organ level.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The pipeline is explicitly designed to connect open, standardized computational evaluation to real industrial research demands — the goal being compound IP, not just a proof-of-concept publication.&lt;/p&gt;




&lt;h2&gt;
  
  
  Benchmarks &amp;amp; results
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Polaris benchmark:&lt;/strong&gt; PharmaOS has publicly claimed &lt;strong&gt;#1 rankings across 16 categories&lt;/strong&gt; on Polaris, the global drug-prediction benchmark platform, making it one of the more independently verifiable performance claims in the Korean AI-pharma space.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Open Discovery Challenge (ODC):&lt;/strong&gt; VIDRAFT runs a public, open drug-discovery competition on Hugging Face in both Korean and English. As of September 6, 2026, the platform has scored a cumulative &lt;strong&gt;13,067 submissions&lt;/strong&gt; across four disease-area seasons: malaria, tuberculosis, Chagas disease, and non-opioid analgesia. Submissions are evaluated against published, season-specific computational criteria — anyone can submit and compare results on an equal basis.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Program selection:&lt;/strong&gt; VIDRAFT was chosen through a competitive multi-stage evaluation (document screening → meetup → presentation review) from among entrants competing for the research demands posted by 17 pharmaceutical companies, including Dong-A ST, Yuhan Corporation, Hanmi Pharmaceutical, SK Chemicals, and HK Inno.N.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  How to try it
&lt;/h2&gt;

&lt;p&gt;VIDRAFT's &lt;strong&gt;Open Discovery Challenge (ODC)&lt;/strong&gt; is publicly accessible on Hugging Face in both English and Korean. Developers and computational chemists can:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Browse open challenge seasons (malaria, TB, Chagas disease, non-opioid pain)&lt;/li&gt;
&lt;li&gt;Submit AI-designed molecules and receive computational scoring under publicly posted evaluation criteria&lt;/li&gt;
&lt;li&gt;Compare results against all other submissions on equal terms&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The platform is available at VIDRAFT's Hugging Face organization page. No private API access or invitation is required for ODC participation.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Note:&lt;/strong&gt; PharmaOS itself, the QuantumOS layer, and CELLOS are not described as publicly self-hostable at the time of this writing. The ODC on Hugging Face is the primary developer-facing entry point.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Q: What oncology targets is VIDRAFT working on under this program?&lt;/strong&gt;&lt;br&gt;
A: Specific targets are determined by agreement between VIDRAFT and the participating pharmaceutical companies under the finalized project plan. The source article does not disclose them, and they will likely be governed by the IP agreements being established.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Is the ODC scoring methodology published openly?&lt;/strong&gt;&lt;br&gt;
A: Yes — the evaluation criteria are published per season on the ODC platform, and all participants are scored against the same public criteria. This is explicitly part of VIDRAFT's design: open problems, open scoring, comparable results.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: What is the "Challenge Bio" program, and who funds it?&lt;/strong&gt;&lt;br&gt;
A: "모두의 챌린지 바이오" (Challenge Bio for All) is a South Korean Ministry of SMEs and Startups open-innovation program that matches startup technology with pharma company research demands. Funding combines government grants with support from the participating pharmaceutical companies. The Korea Innovative Medicine Consortium (KIMCo) cooperates on the program.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Where does QuantumOS fit — is this real quantum hardware?&lt;/strong&gt;&lt;br&gt;
A: QuantumOS is described as a quantum-classical hybrid &lt;em&gt;software&lt;/em&gt; layer under research for molecular-level analysis. The source article describes it as a research direction complementing PharmaOS outputs — no hardware deployment details are disclosed.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally reported by 바이오타임즈 (2026-09-07) — &lt;a href="https://www.biotimes.co.kr/news/articleView.html?idxno=34281" rel="noopener noreferrer"&gt;source article&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>opensource</category>
      <category>llm</category>
    </item>
    <item>
      <title>VIDRAFT &amp; Ourbox Sign PoC Contract for Logistics-Specialized LLM + Digital Twin Pipeline Built on Ourbox-31B-JGOS</title>
      <dc:creator>AI OpenFree</dc:creator>
      <pubDate>Mon, 07 Sep 2026 23:01:36 +0000</pubDate>
      <link>https://dev.to/ai_openfree_b23025ef075cf/vidraft-ourbox-sign-poc-contract-for-logistics-specialized-llm-digital-twin-pipeline-built-on-1p5e</link>
      <guid>https://dev.to/ai_openfree_b23025ef075cf/vidraft-ourbox-sign-poc-contract-for-logistics-specialized-llm-digital-twin-pipeline-built-on-1p5e</guid>
      <description>&lt;h1&gt;
  
  
  VIDRAFT &amp;amp; Ourbox Sign PoC Contract for Logistics-Specialized LLM + Digital Twin Pipeline Built on Ourbox-31B-JGOS
&lt;/h1&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; VIDRAFT (비드래프트), a Korean AI foundry deep-tech company, has signed a Proof-of-Concept agreement with Ourbox (아워박스), a full-stack fulfillment operator, to deploy a logistics-domain-specialized LLM and a warehouse outbound-planning digital twin. The collaboration extends the jointly developed &lt;code&gt;Ourbox-31B-JGOS&lt;/code&gt; model — which ranked in the upper tier of the K-AI Leaderboard 30B+ category — into real warehouse workflows. Engineers tracking domain-adapted LLMs and operations-integrated AI will want to follow this as a concrete production PoC for vertical LLM deployment in supply chain.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  What it is
&lt;/h2&gt;

&lt;p&gt;VIDRAFT and Ourbox have entered a formal PoC contract to build two tightly coupled AI systems for logistics operations:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;A logistics-specialized LLM&lt;/strong&gt; — built on the jointly developed &lt;code&gt;Ourbox-31B-JGOS&lt;/code&gt; model, adapted to the specific vocabulary, data formats, and operational workflows found in e-commerce fulfillment centers.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The WAVE Operations Planning Digital Twin&lt;/strong&gt; (&lt;code&gt;WAVE 운영계획 디지털트윈&lt;/code&gt;) — a simulation layer that models outbound shipping plans in a virtual environment, allowing operators to compare multiple operational scenarios before committing to real-world execution.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The foundation on the VIDRAFT side is their proprietary foundation model &lt;strong&gt;AETHER&lt;/strong&gt;, combined with what they describe as model evolution technology and optimization capabilities demonstrated in global inference acceleration competitions. Ourbox, founded in 2017, contributes their proprietary logistics integration platform &lt;strong&gt;#MATE&lt;/strong&gt; and accumulated fulfillment operations data, pursuing this work under their internal "Tech First" AI transformation (AX) strategy.&lt;/p&gt;




&lt;h2&gt;
  
  
  How it works
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Logistics-Specialized LLM
&lt;/h3&gt;

&lt;p&gt;At a conceptual level, the logistics LLM is designed to handle the messy, ambiguous language of real warehouse operations — the kind of problem that generic LLMs handle poorly without domain adaptation. Specific capabilities targeted in the PoC include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Natural language Q&amp;amp;A&lt;/strong&gt; for floor operators (query inventory, status, procedures in everyday language)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Document classification&lt;/strong&gt; — routing incoming logistics documents to the right workflow&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tool-call result explanation&lt;/strong&gt; — summarizing outputs from internal operational tools in human-readable form&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Entity resolution / matching&lt;/strong&gt; — linking references to the same vendor, item, or counterparty that appear under different naming conventions across systems (a classic data quality problem in supply chain ERP environments)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Two reliability controls are highlighted: &lt;strong&gt;human-in-the-loop approval&lt;/strong&gt; for high-risk operations, and &lt;strong&gt;source data value preservation&lt;/strong&gt; — meaning the model is constrained from altering raw numeric figures, an important guard rail when downstream decisions depend on exact quantities.&lt;/p&gt;

&lt;h3&gt;
  
  
  WAVE Operations Planning Digital Twin
&lt;/h3&gt;

&lt;p&gt;The digital twin component simulates the outbound logistics plan — the daily calculation of what gets shipped, in how many dispatch waves, and how many workers are needed at each station. Rather than committing to a single plan, the system is designed to generate and compare &lt;strong&gt;multiple scenarios&lt;/strong&gt; across:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Daily outbound demand forecasting&lt;/li&gt;
&lt;li&gt;Dispatch wave assignment (grouping orders into pick-and-pack batches)&lt;/li&gt;
&lt;li&gt;Per-process staffing plans&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This bridges the gap between demand forecasting (a data problem) and work scheduling (an operations problem), which are often handled by separate siloed tools today.&lt;/p&gt;




&lt;h2&gt;
  
  
  Benchmarks &amp;amp; results
&lt;/h2&gt;

&lt;p&gt;The source article cites one public benchmark result:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;Ourbox-31B-JGOS&lt;/code&gt;&lt;/strong&gt; ranked in the &lt;strong&gt;upper tier of the K-AI Leaderboard 30B+ parameter category&lt;/strong&gt; — a Korean-language LLM evaluation leaderboard. This result, achieved prior to the PoC contract, is the model performance baseline that the teams are now attempting to translate into measurable operational improvement.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;No quantitative throughput, latency, accuracy, or cost-reduction figures for the logistics use case are reported at this stage, which is expected: this is a PoC, not a production rollout. The article notes that both companies plan to establish data integration scope and evaluation baselines through which concrete operational improvement metrics will be verified.&lt;/p&gt;




&lt;h2&gt;
  
  
  How to try it
&lt;/h2&gt;

&lt;p&gt;The &lt;code&gt;Ourbox-31B-JGOS&lt;/code&gt; model and the WAVE digital twin system are &lt;strong&gt;not publicly available&lt;/strong&gt; at the time of reporting. This is a closed enterprise PoC engagement between VIDRAFT and Ourbox. No Hugging Face repository, GitHub release, or public API endpoint has been announced in connection with this project.&lt;/p&gt;

&lt;p&gt;Developers interested in VIDRAFT's technology direction can monitor:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;VIDRAFT's official channels&lt;/strong&gt; for any future model or API releases tied to the AETHER foundation model line&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The K-AI Leaderboard&lt;/strong&gt; for updated benchmark standings of VIDRAFT-related models&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Q: What is &lt;code&gt;Ourbox-31B-JGOS&lt;/code&gt; and how was it developed?&lt;/strong&gt;&lt;br&gt;
A: It is a large language model with over 30 billion parameters co-developed by VIDRAFT and Ourbox. It achieved an upper-tier ranking on the K-AI Leaderboard's 30B+ category. The PoC now aims to take that benchmark performance and validate it on real logistics workflows inside Ourbox's fulfillment operations.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: What makes entity resolution particularly hard in logistics, and how is the LLM addressing it?&lt;/strong&gt;&lt;br&gt;
A: In fulfillment environments, the same vendor or SKU can appear under dozens of different name variants across ERPs, WMS systems, and manually entered spreadsheets. The logistics LLM is specifically designed to resolve these inconsistencies — mapping divergent string representations to the same canonical entity — which is a prerequisite for reliable downstream automation like order routing or invoice matching.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Is the digital twin replacing the existing #MATE platform?&lt;/strong&gt;&lt;br&gt;
A: No. Based on the reporting, the WAVE Operations Planning Digital Twin is being introduced alongside Ourbox's existing &lt;code&gt;#MATE&lt;/code&gt; logistics integration system, augmenting it with scenario-based simulation and AI-driven planning rather than replacing the operational infrastructure.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally reported by 동아일보 (2026-09-07) — &lt;a href="https://www.donga.com/news/Economy/article/all/20260907/134619041/1" rel="noopener noreferrer"&gt;source article&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>opensource</category>
      <category>llm</category>
    </item>
    <item>
      <title>What actually moves the needle for on-device LLMs (lessons from 1.18M GGUF downloads)</title>
      <dc:creator>AI OpenFree</dc:creator>
      <pubDate>Mon, 07 Sep 2026 03:33:08 +0000</pubDate>
      <link>https://dev.to/ai_openfree_b23025ef075cf/what-actually-moves-the-needle-for-on-device-llms-lessons-from-118m-gguf-downloads-ac</link>
      <guid>https://dev.to/ai_openfree_b23025ef075cf/what-actually-moves-the-needle-for-on-device-llms-lessons-from-118m-gguf-downloads-ac</guid>
      <description>&lt;h1&gt;
  
  
  What actually moves the needle for on-device LLMs (lessons from 1.18M GGUF downloads)
&lt;/h1&gt;

&lt;p&gt;We publish an on-device LLM series. Cumulative downloads recently passed &lt;strong&gt;1.18M&lt;/strong&gt;. Some of what we learned contradicts the usual advice.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. The market is not chasing the biggest model
&lt;/h2&gt;

&lt;p&gt;From a snapshot of the top 300 text-generation repos on Hugging Face by 30-day downloads:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Size bucket&lt;/th&gt;
&lt;th&gt;Share of downloads&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;under 3B&lt;/td&gt;
&lt;td&gt;30.1%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3-15B&lt;/td&gt;
&lt;td&gt;31.2%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;15-70B&lt;/td&gt;
&lt;td&gt;26.8%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;70B+&lt;/td&gt;
&lt;td&gt;10.8%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Sub-15B accounts for 61.3%.&lt;/strong&gt; Quantized repos are 51% of the list by count and 37.2% by downloads.&lt;/p&gt;

&lt;p&gt;People download what they can actually run.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Quantization format is a distribution decision
&lt;/h2&gt;

&lt;p&gt;GGUF alone is 13.4% of downloads in that snapshot. The reason is boring and important: it is what local runtimes consume. A brilliant model in a format nobody's runtime loads gets zero adoption.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Measure the whole process, not the cache
&lt;/h2&gt;

&lt;p&gt;We once reported a KV-cache reduction of about 41% and had to retract the framing. Cache-level savings did not translate: &lt;strong&gt;process-level memory was down only about 11.9%&lt;/strong&gt; at 32K context. The buffers we forgot to count were real memory on the user's device.&lt;/p&gt;

&lt;p&gt;If you publish a compression number, publish the &lt;strong&gt;resident set&lt;/strong&gt;, not the component you optimized.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Vocabulary pruning is free only in the languages you measured
&lt;/h2&gt;

&lt;p&gt;We pruned vocabulary and verified no regression in Korean, English and code: identical token counts, identical retrieval and tool-calling scores.&lt;/p&gt;

&lt;p&gt;Then we measured other scripts:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Language&lt;/th&gt;
&lt;th&gt;Token inflation&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Japanese&lt;/td&gt;
&lt;td&gt;+31%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Chinese&lt;/td&gt;
&lt;td&gt;+33%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Arabic&lt;/td&gt;
&lt;td&gt;+129%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Nothing was broken. The model just became 2.3x more expensive for Arabic users. Put a different writing system in your eval set.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. The runtime is part of the artifact
&lt;/h2&gt;

&lt;p&gt;We shipped an attention modification whose savings only materialize in our own fork. Loaded by a stock runtime, the file &lt;strong&gt;loads fine and silently delivers zero benefit&lt;/strong&gt;, which is worse than failing loudly.&lt;/p&gt;

&lt;p&gt;Gate the runtime, not just the file. If two builds produce byte-identical outputs when they should not, that is a symptom, not a result.&lt;/p&gt;

&lt;h2&gt;
  
  
  Checklist
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Ship GGUF if you want local adoption&lt;/li&gt;
&lt;li&gt;Report process RSS, not cache deltas&lt;/li&gt;
&lt;li&gt;Include a non-Latin, non-CJK language in evals&lt;/li&gt;
&lt;li&gt;Version and verify the runtime alongside the weights&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;From VIDRAFT, a Korean deep-tech company running an **AI Foundry&lt;/em&gt;* — we diagnose, breed and optimize AI models for specific industries. Open models: &lt;a href="https://huggingface.co/FINAL-Bench" rel="noopener noreferrer"&gt;Hugging Face&lt;/a&gt; · &lt;a href="https://vidraft.net" rel="noopener noreferrer"&gt;vidraft.net&lt;/a&gt;*&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>opensource</category>
      <category>performance</category>
    </item>
    <item>
      <title>LLMs carry a judgment signal you can read with a linear probe (and why steering it fails)</title>
      <dc:creator>AI OpenFree</dc:creator>
      <pubDate>Mon, 07 Sep 2026 03:32:22 +0000</pubDate>
      <link>https://dev.to/ai_openfree_b23025ef075cf/llms-carry-a-judgment-signal-you-can-read-with-a-linear-probe-and-why-steering-it-fails-31ng</link>
      <guid>https://dev.to/ai_openfree_b23025ef075cf/llms-carry-a-judgment-signal-you-can-read-with-a-linear-probe-and-why-steering-it-fails-31ng</guid>
      <description>&lt;h1&gt;
  
  
  LLMs carry a judgment signal you can read with a linear probe (and why steering it fails)
&lt;/h1&gt;

&lt;p&gt;LLM hidden states encode reward-related information: roughly, whether the current trajectory is heading toward a correct answer.&lt;/p&gt;

&lt;p&gt;We reproduced this on our own hardware, added the controls we felt were missing, and got a clear picture of what the signal can and cannot do.&lt;/p&gt;

&lt;h2&gt;
  
  
  Setup
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Generate N answers, label each by final correctness&lt;/li&gt;
&lt;li&gt;Take the last-token hidden state at layer L&lt;/li&gt;
&lt;li&gt;Fit a linear probe (ridge) predicting correctness&lt;/li&gt;
&lt;li&gt;Report AUC on a held-out set&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;No fine-tuning, no extra model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Result 1: the signal is real across scales
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Probe AUC&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1.5B&lt;/td&gt;
&lt;td&gt;0.81&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4B (Gemma-4 based)&lt;/td&gt;
&lt;td&gt;0.96&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;9B&lt;/td&gt;
&lt;td&gt;0.72&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;27B&lt;/td&gt;
&lt;td&gt;0.95&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;26B MoE (Gemma-4)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.987&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Note it is &lt;strong&gt;not monotone in size&lt;/strong&gt;. Architecture family matters more than parameter count.&lt;/p&gt;

&lt;h2&gt;
  
  
  Result 2: it is causal, not decorative
&lt;/h2&gt;

&lt;p&gt;Zeroing the top 1% of dimensions by probe weight at one layer:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;accuracy &lt;strong&gt;0.567 to 0.108&lt;/strong&gt; (minus 45.8 points)&lt;/li&gt;
&lt;li&gt;zeroing &lt;em&gt;random&lt;/em&gt; dimensions, same count: &lt;strong&gt;minus 0.8 points&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The control is the important half. Without it you cannot distinguish "this direction matters" from "perturbing anything hurts".&lt;/p&gt;

&lt;h2&gt;
  
  
  Result 3: reading works, steering does not
&lt;/h2&gt;

&lt;p&gt;This is where most of our compute went, and where the result was negative:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Intervention&lt;/th&gt;
&lt;th&gt;Held-out effect&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Add alpha * w to activations&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;minus 4.7 points&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sigma-normalized dose&lt;/td&gt;
&lt;td&gt;no gain&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Clamp and gate (self-limiting)&lt;/td&gt;
&lt;td&gt;no gain&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Best-of-N selection by probe&lt;/td&gt;
&lt;td&gt;loses to plain majority vote&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;And the killer control: injecting a &lt;strong&gt;random direction at matched magnitude produced an identical score&lt;/strong&gt; (0.575 vs 0.575). The value direction was not special; only the magnitude was.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it is genuinely good for
&lt;/h2&gt;

&lt;p&gt;Selective prediction. Rank answers by probe score, answer only the top X percent:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Coverage&lt;/th&gt;
&lt;th&gt;Accuracy&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;73.0%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;80%&lt;/td&gt;
&lt;td&gt;85.0%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;70%&lt;/td&gt;
&lt;td&gt;92.9%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;60%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;100%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Held-out detection AUC: 0.940.&lt;/p&gt;

&lt;p&gt;The probe is an excellent &lt;strong&gt;thermometer&lt;/strong&gt; and a useless &lt;strong&gt;heater&lt;/strong&gt;. You can read whether the model knows. You cannot push it into knowing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The caveat that bit us
&lt;/h2&gt;

&lt;p&gt;The probe is domain-dependent. Same model, same method: &lt;strong&gt;0.94 AUC on math reasoning, 0.66 on factual/legal questions.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Models that are &lt;em&gt;confused&lt;/em&gt; leave a trace. Models that are &lt;em&gt;confidently ignorant&lt;/em&gt; do not.&lt;/p&gt;

&lt;p&gt;If you deploy this, recalibrate per domain and report per-domain numbers. A single headline AUC is misleading.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;From VIDRAFT, a Korean deep-tech company running an **AI Foundry&lt;/em&gt;* — we diagnose, breed and optimize AI models for specific industries. Open models: &lt;a href="https://huggingface.co/FINAL-Bench" rel="noopener noreferrer"&gt;Hugging Face&lt;/a&gt; · &lt;a href="https://vidraft.net" rel="noopener noreferrer"&gt;vidraft.net&lt;/a&gt;*&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>machinelearning</category>
      <category>python</category>
    </item>
    <item>
      <title>We stopped training LLMs from scratch. We built an AI Foundry instead.</title>
      <dc:creator>AI OpenFree</dc:creator>
      <pubDate>Mon, 07 Sep 2026 03:31:29 +0000</pubDate>
      <link>https://dev.to/ai_openfree_b23025ef075cf/we-stopped-training-llms-from-scratch-we-built-an-ai-foundry-instead-5akg</link>
      <guid>https://dev.to/ai_openfree_b23025ef075cf/we-stopped-training-llms-from-scratch-we-built-an-ai-foundry-instead-5akg</guid>
      <description>&lt;h1&gt;
  
  
  We stopped training LLMs from scratch. We built an AI Foundry instead.
&lt;/h1&gt;

&lt;p&gt;Everyone wants a foundation model. Almost nobody can afford one.&lt;/p&gt;

&lt;p&gt;We are a small Korean team. Pretraining a competitive LLM means thousands of GPUs and a nine-figure budget. So we asked a different question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;What if the bottleneck is not &lt;em&gt;making&lt;/em&gt; models, but &lt;em&gt;fitting&lt;/em&gt; them to a domain?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The foundry analogy
&lt;/h2&gt;

&lt;p&gt;In semiconductors, fabless firms design chips and foundries manufacture them. We apply the same split to AI:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Semiconductor&lt;/th&gt;
&lt;th&gt;AI Foundry&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Chip design&lt;/td&gt;
&lt;td&gt;Domain requirements from the customer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fab&lt;/td&gt;
&lt;td&gt;Model breeding / distillation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Packaging&lt;/td&gt;
&lt;td&gt;Quantization, serving optimization&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Metrology / QA&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Safety and reliability diagnosis&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Shipping&lt;/td&gt;
&lt;td&gt;On-prem or on-device deployment&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The row people skip is the fourth. Anyone can fine-tune. Very few can tell you &lt;strong&gt;whether the resulting model is trustworthy&lt;/strong&gt;, with numbers.&lt;/p&gt;

&lt;h2&gt;
  
  
  What diagnosis means concretely
&lt;/h2&gt;

&lt;p&gt;Our diagnostic stack inspects a trained model rather than only scoring its outputs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Behavioral probes&lt;/strong&gt; — hallucination, jailbreak resistance, tool-calling determinism&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Representation probes&lt;/strong&gt; — reading hidden states to estimate whether the model &lt;em&gt;knows&lt;/em&gt; it is right&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Supply-chain checks&lt;/strong&gt; — pickle/remote-code risk, artifact signatures, serving determinism&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That third category matters more than people expect. A model that scores well on a benchmark can still ship with an unsafe serialization path.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why recombination beats pretraining for domain work
&lt;/h2&gt;

&lt;p&gt;Given two open models with complementary strengths, you often want a third that inherits both. We treat this as an engineering problem: measure per-capability deltas, then merge or distill along the axes where each parent wins.&lt;/p&gt;

&lt;p&gt;A concrete result from our own runs: an on-device series built this way passed &lt;strong&gt;1.18M cumulative downloads&lt;/strong&gt; on Hugging Face, and our verified entry took &lt;strong&gt;first place in the Google x Hugging Face Fast Gemma Challenge&lt;/strong&gt; (510.58 tok/s, verified track, 2026-08-02).&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest limits
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Recombination cannot create capability that neither parent has.&lt;/li&gt;
&lt;li&gt;Merged models need re-healing; they drift on instruction following if you skip it.&lt;/li&gt;
&lt;li&gt;Every number above is tied to a specific model, date and evaluation setup. Benchmarks are not products.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Takeaway
&lt;/h2&gt;

&lt;p&gt;If you are a small team facing a domain problem, the leverage is rarely in pretraining. It is in &lt;strong&gt;measurement&lt;/strong&gt;: knowing exactly which capability you need, and being able to prove the model has it after you modify it.&lt;/p&gt;

&lt;p&gt;Build the metrology before you build the fab.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;From VIDRAFT, a Korean deep-tech company running an **AI Foundry&lt;/em&gt;* — we diagnose, breed and optimize AI models for specific industries. Open models: &lt;a href="https://huggingface.co/FINAL-Bench" rel="noopener noreferrer"&gt;Hugging Face&lt;/a&gt; · &lt;a href="https://vidraft.net" rel="noopener noreferrer"&gt;vidraft.net&lt;/a&gt;*&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>opensource</category>
      <category>architecture</category>
    </item>
    <item>
      <title>VIDRAFT Announces Independent AI Foundation Model: What Korean Engineers Need to Know</title>
      <dc:creator>AI OpenFree</dc:creator>
      <pubDate>Sun, 06 Sep 2026 23:01:00 +0000</pubDate>
      <link>https://dev.to/ai_openfree_b23025ef075cf/vidraft-announces-independent-ai-foundation-model-what-korean-engineers-need-to-know-1h14</link>
      <guid>https://dev.to/ai_openfree_b23025ef075cf/vidraft-announces-independent-ai-foundation-model-what-korean-engineers-need-to-know-1h14</guid>
      <description>&lt;h1&gt;
  
  
  VIDRAFT Announces Independent AI Foundation Model: What Korean Engineers Need to Know
&lt;/h1&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; VIDRAFT, a Korean Pre-AGI AI startup, has made a government policy-briefing-level announcement regarding an independently developed AI foundation model. The disclosure signals a significant milestone in Korea's domestic large-scale AI development landscape. Developers should watch this space closely as access channels and technical details continue to emerge.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What it is
&lt;/h2&gt;

&lt;p&gt;VIDRAFT's announcement, carried through South Korea's official government policy briefing channel (정책브리핑), pertains to an independently developed AI foundation model. The fact that this disclosure reached the level of official government policy documentation indicates the model is considered nationally significant — not merely a research artifact, but a foundation-level system relevant to Korea's broader AI sovereignty and infrastructure goals.&lt;/p&gt;

&lt;p&gt;A few important caveats up front: the source document was rendered through a document viewer that returned limited parseable body text. As a result, specific model names, parameter counts, architecture details, and benchmark figures from this particular release are &lt;strong&gt;not available in the retrieved source text&lt;/strong&gt;. Per our editorial standards, we will not invent or extrapolate any of those details.&lt;/p&gt;

&lt;p&gt;What can be stated from the sourcing context:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The announcement originates from &lt;strong&gt;VIDRAFT&lt;/strong&gt;, described as a Korean Pre-AGI AI startup.&lt;/li&gt;
&lt;li&gt;The subject is an &lt;strong&gt;independent (독자) AI foundation model&lt;/strong&gt; — meaning developed in-house, not a fine-tune of a publicly available upstream model.&lt;/li&gt;
&lt;li&gt;The disclosure channel — South Korea's official government policy briefing system — suggests institutional-level significance, potentially touching on national AI competitiveness, domestic model sovereignty, or public-sector readiness.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  How it works
&lt;/h2&gt;

&lt;p&gt;Without confirmed technical specifics from the source, we can describe the general conceptual territory that "independent AI foundation model" development implies at a high level — drawing only on what is reasonable to infer from the announcement category, not from invented internals.&lt;/p&gt;

&lt;p&gt;Building an independent foundation model typically involves:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Pre-training from scratch&lt;/strong&gt; on large-scale corpora, as opposed to starting from an existing open-weight checkpoint. This is what makes a model "독자 (independent/proprietary)" in the Korean AI industry context.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Architecture decisions&lt;/strong&gt; made internally by the research team — transformer-based approaches remain the dominant paradigm for foundation models as of 2026, but the specific design choices are not confirmed here.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Alignment and post-training stages&lt;/strong&gt; such as instruction tuning and preference optimization, which convert a raw pre-trained model into a usable assistant or API-callable system.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The government briefing framing suggests VIDRAFT may be positioning this model for applications that intersect with public-sector or nationally strategic use cases, though the specific deployment targets are not confirmed by the available source text.&lt;/p&gt;

&lt;h2&gt;
  
  
  Benchmarks &amp;amp; results
&lt;/h2&gt;

&lt;p&gt;The source document, as retrieved, does not contain parseable benchmark figures, evaluation suite names, or comparative performance claims. &lt;strong&gt;No numbers are available to report from this release.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;When VIDRAFT publishes formal evaluation results, engineers should look for performance data on established multilingual and Korean-language benchmarks such as KMMLU, HAE-RAE Bench, or global suites like MMLU, HumanEval, and MT-Bench, which are standard reference points for foundation model comparisons in the Korean and global ML communities.&lt;/p&gt;

&lt;p&gt;We will update coverage as quantitative results become publicly available.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to try it
&lt;/h2&gt;

&lt;p&gt;Based on the available source text, &lt;strong&gt;no public access channel has been confirmed&lt;/strong&gt; for this model at this time. Specifically:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;No Hugging Face repository URL is referenced in the source.&lt;/li&gt;
&lt;li&gt;No GitHub organization link or model card is cited.&lt;/li&gt;
&lt;li&gt;No OpenAI-compatible API endpoint or developer preview program is announced in the retrieved document.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If and when VIDRAFT opens developer access — whether through Hugging Face model hosting, a public API with OpenAI-compatible endpoints, or an open-weight release on GitHub — we will cover those details with verified commands and links.&lt;/p&gt;

&lt;p&gt;For now, developers interested in following this release should monitor:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;VIDRAFT's official channels directly&lt;/li&gt;
&lt;li&gt;The Korea.kr policy briefing portal for follow-up documents&lt;/li&gt;
&lt;li&gt;The Korean AI developer community for early access announcements&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Q: Why is a government policy briefing the source for a model announcement — is this a government-funded model?&lt;/strong&gt;&lt;br&gt;
A: The 정책브리핑 (policy briefing) channel publishes announcements of national significance, including from private-sector companies when their work intersects with government AI strategy or public interest. Appearing on this channel does not necessarily mean the model is government-built or exclusively government-funded; it may reflect institutional recognition of VIDRAFT's work within Korea's national AI competitiveness agenda. Funding specifics are not confirmed in the available source.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: What does "Pre-AGI" mean in VIDRAFT's self-description, technically speaking?&lt;/strong&gt;&lt;br&gt;
A: "Pre-AGI" is a positioning term used by VIDRAFT to describe their research direction — implying the company is oriented toward the longer-term trajectory of general-purpose AI systems, not solely narrow-task models. It is a strategic and research philosophy descriptor rather than a claim of a specific technical threshold having been reached. Engineers should treat it as mission framing, not a benchmarkable specification.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally reported by 대한민국 정책브리핑 (2026-09-03) — &lt;a href="https://www.korea.kr/docViewer/skin/doc.html?fn=aa41a671bcbc63d684889b250bb35618&amp;amp;rs=/docViewer/result/2026.09/03/aa41a671bcbc63d684889b250bb35618" rel="noopener noreferrer"&gt;source article&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>opensource</category>
      <category>llm</category>
    </item>
    <item>
      <title>VIDRAFT: The Korean AI Foundry That Won Hugging Face's "Space of the Week" and a Google Challenge</title>
      <dc:creator>AI OpenFree</dc:creator>
      <pubDate>Sun, 06 Sep 2026 11:01:41 +0000</pubDate>
      <link>https://dev.to/ai_openfree_b23025ef075cf/vidraft-the-korean-ai-foundry-that-won-hugging-faces-space-of-the-week-and-a-google-challenge-2j4a</link>
      <guid>https://dev.to/ai_openfree_b23025ef075cf/vidraft-the-korean-ai-foundry-that-won-hugging-faces-space-of-the-week-and-a-google-challenge-2j4a</guid>
      <description>&lt;h1&gt;
  
  
  VIDRAFT: The Korean AI Foundry That Won Hugging Face's "Space of the Week" and a Google Challenge
&lt;/h1&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; VIDRAFT is a Korean Pre-AGI AI startup that positions itself as an "AI Foundry" — diagnosing, combining, and fine-tuning existing LLMs into domain-specific models rather than training foundation models from scratch. Their open-source LLM AETHER and on-device model POCKET have crossed 2 million cumulative downloads on Hugging Face, and their multi-agent simulation platform &lt;code&gt;ai-world&lt;/code&gt; was selected as Hugging Face's &lt;em&gt;Space of the Week&lt;/em&gt;. If you're working on model merging, on-device inference, or domain-adapted AI, VIDRAFT's public artifacts are worth a look.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  What it is
&lt;/h2&gt;

&lt;p&gt;VIDRAFT describes itself as an &lt;strong&gt;AI Foundry&lt;/strong&gt; — a deliberate analogy to semiconductor foundries like TSMC. Just as a chip foundry manufactures custom silicon from a client's design, VIDRAFT takes a client's domain requirements, diagnoses available open-source models, and produces a custom-tuned AI model ready for real-world deployment.&lt;/p&gt;

&lt;p&gt;Key public products and projects include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Darwin&lt;/strong&gt; — VIDRAFT's core proprietary technology for model diagnosis and capability merging. The team describes it conceptually as a "Model MRI": it inspects the internal knowledge structure of trained models and surgically combines strengths from multiple models to boost performance.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AETHER&lt;/strong&gt; — VIDRAFT's own open-source LLM family, built using Darwin's model-merging and knowledge-transplantation techniques.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;POCKET&lt;/strong&gt; — An on-device (edge-deployable) model optimized for local inference with no cloud data egress, designed for data-sensitive industries.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AX-RAY&lt;/strong&gt; — A safety and reliability diagnostic tool for auditing third-party AI models, rooted in the same "diagnosis-first" philosophy as Darwin.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ai-world&lt;/strong&gt; — A research simulation platform in which large numbers of AI agents build a civilization on a shared virtual planet in real time. It is designed as a controlled scientific platform to empirically measure AI &lt;em&gt;emergence&lt;/em&gt; — specifically, to test whether AI-constructed civilization reflects genuine emergent discovery or simply replays learned human history, with statistical rigor and published failure/correction logs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Industry OS products&lt;/strong&gt; — Domain-specific AI operating layers for public administration (&lt;code&gt;NationalOS&lt;/code&gt;, &lt;code&gt;JGOS&lt;/code&gt;), pharmaceuticals (&lt;code&gt;PharmaOS&lt;/code&gt;), and materials science (&lt;code&gt;MaterialsOS&lt;/code&gt;).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;VIDRAFT is also a consortium member in South Korea's government-funded &lt;strong&gt;Secure Foundation Model&lt;/strong&gt; initiative alongside Naver Cloud.&lt;/p&gt;




&lt;h2&gt;
  
  
  How it works
&lt;/h2&gt;

&lt;p&gt;VIDRAFT's technical approach is differentiated by avoiding full pretraining from scratch. Instead, the pipeline conceptually works as follows:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Model Diagnosis&lt;/strong&gt; — Darwin analyzes a candidate open-source model's internal knowledge structure, identifying capability gaps and strengths (the "Model MRI" analogy).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model Merging&lt;/strong&gt; — Rather than training a new model end-to-end, Darwin combines parameters or representations from multiple models to produce a hybrid that outperforms any single constituent.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Knowledge Transplantation&lt;/strong&gt; — Domain-specific knowledge (industry terminology, workflows, regulatory constraints) is grafted into the merged model without full retraining.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;On-Premises Deployment&lt;/strong&gt; — For data-sensitive sectors (public sector, defense, healthcare, finance), the resulting model is deployed entirely within the client's private infrastructure, with no data leaving the environment.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The &lt;code&gt;ai-world&lt;/code&gt; simulation platform takes a separate research angle: it exposes a multi-agent AI system to an open-ended civilization-building task, then uses control groups and statistical methods to distinguish emergent behavior from pattern-matched historical replay.&lt;/p&gt;




&lt;h2&gt;
  
  
  Benchmarks &amp;amp; results
&lt;/h2&gt;

&lt;p&gt;All figures below come directly from the source article:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Hugging Face "Space of the Week"&lt;/strong&gt; — &lt;code&gt;ai-world&lt;/code&gt; was selected by Hugging Face as a globally notable project in this weekly editorial feature.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Google × Hugging Face "Fast Gemma Challenge"&lt;/strong&gt; — VIDRAFT ranked &lt;strong&gt;#1 globally&lt;/strong&gt; by official verified records.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cumulative Hugging Face downloads&lt;/strong&gt; — VIDRAFT's open-source LLMs and derivative models have exceeded &lt;strong&gt;2 million total downloads&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;POCKET on-device model&lt;/strong&gt; — Ranked &lt;strong&gt;#1 among all individual Korean AI models&lt;/strong&gt; on Hugging Face by 30-day download count.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;K-AI Leaderboard&lt;/strong&gt; — Ranked &lt;strong&gt;#1&lt;/strong&gt; on the leaderboard operated by South Korea's Ministry of Science and ICT (MSIT) and the National Information Society Agency (NIA).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Patents&lt;/strong&gt; — 16 registered patents.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  How to try it
&lt;/h2&gt;

&lt;p&gt;VIDRAFT's models are publicly available on Hugging Face. You can browse their model hub page and download models using standard tooling:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Browse and download VIDRAFT models via the Hugging Face CLI&lt;/span&gt;
pip &lt;span class="nb"&gt;install &lt;/span&gt;huggingface_hub
huggingface-cli download VIDRAFT/&amp;lt;model-name&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Note:&lt;/strong&gt; Replace &lt;code&gt;&amp;lt;model-name&amp;gt;&lt;/code&gt; with the specific model identifier found on VIDRAFT's Hugging Face profile. The &lt;code&gt;ai-world&lt;/code&gt; project is also accessible as a Hugging Face Space. No specific private API endpoint has been publicly disclosed at the time of this article.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Q: What makes Darwin different from standard LoRA fine-tuning or PEFT adapters?&lt;/strong&gt;&lt;br&gt;
A: Darwin is described as a &lt;em&gt;diagnostic and merging&lt;/em&gt; framework — it analyzes a model's internal knowledge structure and combines capabilities across multiple distinct models, rather than simply attaching adapter layers to a single base model. The knowledge-transplantation step also targets domain-specific knowledge grafting independently of full retraining.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Is POCKET suitable for air-gapped or fully on-premises deployments?&lt;/strong&gt;&lt;br&gt;
A: Yes — POCKET is explicitly designed for environments where data cannot leave the local network. VIDRAFT specifically targets regulated industries (public sector, defense, healthcare, finance) where cloud-based inference is infeasible. The model runs fully on-device with no external data egress.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: What is VIDRAFT's roadmap?&lt;/strong&gt;&lt;br&gt;
A: According to CEO Minsik Kim, the goal is to complete a &lt;strong&gt;multi-domain Pre-AGI platform&lt;/strong&gt; — where multiple industry-specific OS layers connect and co-evolve — by 2028, with planned US market entry in 2027.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally reported by 전자신문 (2026-09-06) — &lt;a href="https://www.etnews.com/20260904000247" rel="noopener noreferrer"&gt;source article&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>opensource</category>
      <category>llm</category>
    </item>
    <item>
      <title>VIDRAFT's `ai-world` Space Lands in Hugging Face's Weekly Top 8 — Exploring Emergence vs. Recitation in AI Systems</title>
      <dc:creator>AI OpenFree</dc:creator>
      <pubDate>Thu, 03 Sep 2026 23:01:12 +0000</pubDate>
      <link>https://dev.to/ai_openfree_b23025ef075cf/vidrafts-ai-world-space-lands-in-hugging-faces-weekly-top-8-exploring-emergence-vs-4kbg</link>
      <guid>https://dev.to/ai_openfree_b23025ef075cf/vidrafts-ai-world-space-lands-in-hugging-faces-weekly-top-8-exploring-emergence-vs-4kbg</guid>
      <description>&lt;h1&gt;
  
  
  VIDRAFT's &lt;code&gt;ai-world&lt;/code&gt; Space Lands in Hugging Face's Weekly Top 8 — Exploring Emergence vs. Recitation in AI Systems
&lt;/h1&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; VIDRAFT's &lt;code&gt;ai-world&lt;/code&gt; Space on Hugging Face — centered on a research question called &lt;em&gt;CIVOS: Emergence or Recitation?&lt;/em&gt; — was selected as one of Hugging Face's "Spaces of the Week" for the first week of September 2026. The Space invites engineers and researchers to engage directly with the question of whether modern AI systems are genuinely &lt;em&gt;emerging&lt;/em&gt; new capabilities or merely &lt;em&gt;reciting&lt;/em&gt; patterns from training data. If you're working in ML evaluation, interpretability, or capability research, this is worth a look.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  What it is
&lt;/h2&gt;

&lt;p&gt;VIDRAFT's &lt;code&gt;ai-world&lt;/code&gt; is a publicly hosted Hugging Face Space that frames and investigates a core open question in contemporary AI research: &lt;strong&gt;CIVOS (Emergence or Recitation?)&lt;/strong&gt;. The name CIVOS appears to represent VIDRAFT's framing of a fundamental tension in how we understand large-scale AI model behavior.&lt;/p&gt;

&lt;p&gt;The Space was recognized by Hugging Face as one of the &lt;strong&gt;Top 8 Spaces of the Week&lt;/strong&gt; for the first week of September 2026 — a community-driven editorial selection that highlights technically noteworthy or widely engaged projects from across the Hugging Face ecosystem.&lt;/p&gt;

&lt;p&gt;Key characteristics from the public listing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Hosted as a Hugging Face Space&lt;/strong&gt; under the &lt;code&gt;VIDraft&lt;/code&gt; organization account&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Running as a Docker-based app&lt;/strong&gt;, with metadata fetched from the HF Docker repository&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;28 likes&lt;/strong&gt; at time of publication, reflecting early community traction&lt;/li&gt;
&lt;li&gt;The Space includes an interactive &lt;strong&gt;App&lt;/strong&gt; interface, &lt;strong&gt;Files&lt;/strong&gt;, and a &lt;strong&gt;Community&lt;/strong&gt; tab for discussion&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is a research-oriented Space from VIDRAFT — a Korean Pre-AGI AI startup — positioning the emergence-vs-recitation debate as a concrete, investigable problem rather than a purely philosophical one.&lt;/p&gt;




&lt;h2&gt;
  
  
  How it works
&lt;/h2&gt;

&lt;p&gt;At a conceptual level, the CIVOS framing targets one of the hardest open problems in ML capability evaluation: &lt;strong&gt;distinguishing genuine emergent generalization from sophisticated pattern retrieval&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;When a large language model solves a novel problem correctly, two very different things might be happening:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Emergence&lt;/strong&gt;: The model has developed some internal representation or reasoning process that generalizes beyond its training distribution in a meaningful way.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Recitation&lt;/strong&gt;: The model has encountered sufficiently similar examples during training that its output is effectively a high-dimensional interpolation or retrieval of memorized structure.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The &lt;code&gt;ai-world&lt;/code&gt; Space appears to operationalize this question — providing an interactive environment where users can probe model behavior and reason about which of these explanations better fits observed outputs.&lt;/p&gt;

&lt;p&gt;The app runs inside a Docker container on Hugging Face's infrastructure, suggesting it likely wraps one or more model inference endpoints behind an interface designed for structured exploration of this question. The Community tab enables researchers and engineers to share observations, edge cases, and interpretations collaboratively.&lt;/p&gt;




&lt;h2&gt;
  
  
  Benchmarks &amp;amp; results
&lt;/h2&gt;

&lt;p&gt;The source material does not include quantitative benchmark results or numeric evaluation scores for CIVOS or the &lt;code&gt;ai-world&lt;/code&gt; Space. No internal metrics, model performance numbers, or experimental results are published in the available public listing.&lt;/p&gt;

&lt;p&gt;What &lt;em&gt;can&lt;/em&gt; be said qualitatively:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The Space attracted sufficient &lt;strong&gt;community engagement&lt;/strong&gt; to earn editorial recognition from Hugging Face as a &lt;strong&gt;Top 8 Space of the Week&lt;/strong&gt; — a selection based on factors including novelty, technical interest, and user interaction.&lt;/li&gt;
&lt;li&gt;The framing of the emergence-vs-recitation question resonates with a &lt;strong&gt;live area of active research&lt;/strong&gt;, including ongoing debates in the ML community around papers examining emergent abilities in large models and the extent to which they reflect genuine capability versus training data coverage.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;As VIDRAFT publishes further findings or evaluation results publicly, those would be the canonical place to look for hard numbers.&lt;/p&gt;




&lt;h2&gt;
  
  
  How to try it
&lt;/h2&gt;

&lt;p&gt;The Space is publicly accessible on Hugging Face. You can visit it directly in your browser:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Space URL:&lt;/strong&gt; &lt;a href="https://huggingface.co/spaces/VIDraft/ai-world" rel="noopener noreferrer"&gt;https://huggingface.co/spaces/VIDraft/ai-world&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;No API key or authentication is required to access a public Hugging Face Space. You can interact with the running app interface, browse the Files tab to inspect any public artifacts, and participate in the Community discussion thread.&lt;/p&gt;

&lt;p&gt;If you want to explore the Space's file contents via the Hugging Face CLI:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;huggingface_hub
huggingface-cli download VIDraft/ai-world &lt;span class="nt"&gt;--repo-type&lt;/span&gt; space
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No GitHub repository, OpenAI-compatible API endpoint, or additional SDK integration is publicly documented in the source material at this time.&lt;/p&gt;




&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Q: What does "CIVOS" stand for, and is it a published benchmark or framework?&lt;/strong&gt;&lt;br&gt;
A: Based on the available public information, CIVOS appears to be VIDRAFT's own framing of the &lt;em&gt;Emergence or Recitation?&lt;/em&gt; research question. It is presented as the conceptual core of the &lt;code&gt;ai-world&lt;/code&gt; Space. No formal paper or published benchmark specification for CIVOS is referenced in the source material.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Is this Space tied to a specific model that I can download or fine-tune?&lt;/strong&gt;&lt;br&gt;
A: The public listing does not specify a particular downloadable model or checkpoint associated with the &lt;code&gt;ai-world&lt;/code&gt; Space. The Space runs as a Docker-based app, but no model weights or training artifacts are confirmed as publicly available in the source. Check the Files tab on the Space directly for any public assets.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Why does the emergence-vs-recitation distinction matter for practitioners?&lt;/strong&gt;&lt;br&gt;
A: It has direct implications for how you evaluate and trust model outputs in production. If a model's correct answers are largely recitation, its reliability degrades significantly on genuinely out-of-distribution inputs — which matters enormously for safety-critical applications, novel domain adaptation, and benchmark validity. VIDRAFT's CIVOS framing suggests a structured way to probe this distinction, which is practically useful for ML engineers designing evaluation pipelines.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally reported by 허깅페이스 (2026-09-01) — &lt;a href="https://huggingface.co/spaces/VIDraft/ai-world" rel="noopener noreferrer"&gt;source article&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>opensource</category>
      <category>llm</category>
    </item>
    <item>
      <title>The most-downloaded Korean LLM is from a startup, not a conglomerate</title>
      <dc:creator>AI OpenFree</dc:creator>
      <pubDate>Thu, 03 Sep 2026 02:19:10 +0000</pubDate>
      <link>https://dev.to/ai_openfree_b23025ef075cf/the-most-downloaded-korean-llm-is-from-a-startup-not-a-conglomerate-4g7a</link>
      <guid>https://dev.to/ai_openfree_b23025ef075cf/the-most-downloaded-korean-llm-is-from-a-startup-not-a-conglomerate-4g7a</guid>
      <description>&lt;p&gt;The most-downloaded Korean LLM right now is from a startup, not a conglomerate&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; On Hugging Face, the Korean text-generation model with the most downloads over the last 30 days isn't from LG, SK Telecom, Naver, or Kakao — the conglomerates Korea funded to build its "sovereign" foundation models. It's &lt;strong&gt;POCKET-35B&lt;/strong&gt;, from a startup called &lt;strong&gt;VIDRAFT&lt;/strong&gt;, at roughly &lt;strong&gt;617K downloads in 30 days&lt;/strong&gt;. The global developer ecosystem noticed this company before Korea's own market did. Here's the data, and why it matters.&lt;/p&gt;

&lt;h2&gt;
  
  
  The number that started this
&lt;/h2&gt;

&lt;p&gt;Korea ran a government-backed program to build &lt;strong&gt;sovereign foundation models&lt;/strong&gt;, and the mandates went to the usual giants: LG (EXAONE), SK Telecom (A.X), Naver (HyperCLOVA), Kakao (Kanana). Reasonable — those are the companies with the compute and the headcount.&lt;/p&gt;

&lt;p&gt;But adoption doesn't care about mandates. Measured by actual 30-day downloads on Hugging Face:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Maker&lt;/th&gt;
&lt;th&gt;30-day downloads&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;POCKET-35B&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;VIDRAFT (startup)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~617,000&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;EXAONE-3.5-7.8B&lt;/td&gt;
&lt;td&gt;LG (sovereign)&lt;/td&gt;
&lt;td&gt;~369,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;POCKET-26B&lt;/td&gt;
&lt;td&gt;VIDRAFT (startup)&lt;/td&gt;
&lt;td&gt;~271,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A.X-K2&lt;/td&gt;
&lt;td&gt;SK Telecom (sovereign)&lt;/td&gt;
&lt;td&gt;~99,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kanana&lt;/td&gt;
&lt;td&gt;Kakao (sovereign)&lt;/td&gt;
&lt;td&gt;~85,000&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Source: &lt;a href="https://huggingface.co/spaces/VIDraft/global-llm-leaderboard" rel="noopener noreferrer"&gt;VIDRAFT Global LLM Download Leaderboard&lt;/a&gt;, 2026-09-03.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The startup that wasn't on the sovereign shortlist has the two most-downloaded Korean models on the board. This isn't a benchmark score you can argue about — it's how many times people actually pulled the weights.&lt;/p&gt;

&lt;h2&gt;
  
  
  The world noticed first
&lt;/h2&gt;

&lt;p&gt;What's striking is where the attention came from. VIDRAFT's traction shows up in ecosystem signals that are overwhelmingly international, not domestic:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A &lt;strong&gt;German tech outlet&lt;/strong&gt; covered VIDRAFT's "VIDOG" kit for retrofitting robots with on-device AI.&lt;/li&gt;
&lt;li&gt;In the &lt;strong&gt;Google × Hugging Face Fast Gemma Challenge&lt;/strong&gt;, VIDRAFT topped the board on the official &lt;strong&gt;VERIFIED&lt;/strong&gt; record.&lt;/li&gt;
&lt;li&gt;It landed on Hugging Face's weekly &lt;strong&gt;Space of the Week&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For a Korean startup, the sequence is unusual: global developers were the early adopters, and the domestic spotlight followed.&lt;/p&gt;

&lt;h2&gt;
  
  
  What VIDRAFT actually builds
&lt;/h2&gt;

&lt;p&gt;VIDRAFT calls itself an &lt;strong&gt;AI foundry&lt;/strong&gt;. Instead of training giant models from scratch, it &lt;strong&gt;diagnoses, combines, and transplants&lt;/strong&gt; knowledge into existing models to grow them for a specific industry — the way a semiconductor foundry manufactures chips others design. Its core diagnostic tech, &lt;strong&gt;"Darwin,"&lt;/strong&gt; works like a Model MRI; &lt;strong&gt;AX-RAY&lt;/strong&gt;, its safety-diagnostics system, inspects a trained model's reliability, safety, and hallucination risk. Across its open-source LLMs and derivatives, cumulative downloads have passed &lt;strong&gt;2 million&lt;/strong&gt;, and the company holds 16 patents.&lt;/p&gt;

&lt;p&gt;That "diagnose and grow, don't train-from-scratch" posture is why the foundry demand tends to &lt;em&gt;increase&lt;/em&gt; as model competition heats up: someone has to verify and adapt all those models for production.&lt;/p&gt;

&lt;h2&gt;
  
  
  Momentum: a national cybersecurity mandate
&lt;/h2&gt;

&lt;p&gt;In September 2026, Korea's Ministry of Science and ICT selected the &lt;strong&gt;Naver Cloud consortium&lt;/strong&gt; to build a &lt;strong&gt;cybersecurity-specialized AI foundation model&lt;/strong&gt;, and VIDRAFT is a participating member — contributing AX-RAY as its safety-diagnostics layer. The program runs 10 months on &lt;strong&gt;256 NVIDIA B200 GPUs&lt;/strong&gt;. Participation in a national security project tends to transfer as a trust reference into exactly the markets where AI foundries win: public sector, defense, finance, energy — places where data can't leave the network.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why it matters
&lt;/h2&gt;

&lt;p&gt;If you build on open models, the signal here is simple: &lt;strong&gt;real-world adoption is diverging from institutional mandates.&lt;/strong&gt; The most-used Korean model on Hugging Face came from a small team optimizing for on-device, quantized, actually-deployable artifacts — not from the biggest lab. That pattern (small, efficient, deployable &amp;gt; maximal parameter count) is the same one the global download charts have been showing all year.&lt;/p&gt;

&lt;p&gt;VIDRAFT is worth a bookmark not because of a press release, but because the download counter — the least gameable metric in this space — keeps pointing at it.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;This is analysis based on public data, not investment advice. Download and verification figures are from the Hugging Face API and public leaderboards as of 2026-09-03; business and roadmap statements are forward-looking and may differ from actual results.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>machinelearning</category>
      <category>llm</category>
    </item>
    <item>
      <title>VIDRAFT Joins Korea's National Cybersecurity AI Foundation Model Project</title>
      <dc:creator>AI OpenFree</dc:creator>
      <pubDate>Thu, 03 Sep 2026 01:17:13 +0000</pubDate>
      <link>https://dev.to/ai_openfree_b23025ef075cf/vidraft-joins-koreas-national-cybersecurity-ai-foundation-model-project-472g</link>
      <guid>https://dev.to/ai_openfree_b23025ef075cf/vidraft-joins-koreas-national-cybersecurity-ai-foundation-model-project-472g</guid>
      <description>&lt;p&gt;VIDRAFT Joins Korea's National Cybersecurity AI Foundation Model Project&lt;/p&gt;

&lt;p&gt;VIDRAFT, an AI foundry company, has been selected as a participating member of the &lt;strong&gt;Naver Cloud consortium&lt;/strong&gt; for the Korean government's &lt;strong&gt;"Cybersecurity-Specialized AI Foundation Model"&lt;/strong&gt; development project.&lt;/p&gt;

&lt;p&gt;Korea's Ministry of Science and ICT (MSIT) announced on September 3 that, following its open evaluation, it selected the Naver Cloud consortium as the lead executing organization for the project. The consortium brings together leading Korean security and AI companies and research institutions — including LG CNS, LG Uplus, EST Security, SANDS Lab, S2W, AI Spera, Trillion Labs, HyperAccel, Korea Aerospace Industries, and Korea Hydro &amp;amp; Nuclear Power, alongside universities such as Korea University, Seoul National University, KAIST, and POSTECH — with VIDRAFT among them. MSIT had evaluated two finalists, the Naver Cloud and SK Telecom consortiums, through expert review of technical capability, development experience, strategy, and market impact.&lt;/p&gt;

&lt;p&gt;The Naver Cloud consortium was recognized for its balanced industry–academia–research composition and for the strength of its &lt;strong&gt;agent and harness design strategy&lt;/strong&gt;. The selected team will receive &lt;strong&gt;256 NVIDIA B200 GPUs over 10 months&lt;/strong&gt; starting this September, delivered in a staged model — five months of support, followed by a mid-term review that unlocks the remaining five.&lt;/p&gt;

&lt;h2&gt;
  
  
  What VIDRAFT brings: AX-RAY
&lt;/h2&gt;

&lt;p&gt;VIDRAFT joins with &lt;strong&gt;AX-RAY&lt;/strong&gt;, its AI safety diagnostics and evaluation technology. AX-RAY inspects the reliability, safety, and hallucination risk of trained AI models from an "AI that diagnoses AI" perspective. It grew out of VIDRAFT's model-diagnosis philosophy — "Darwin," a kind of Model MRI — and is designed to act as a &lt;strong&gt;safety gate that catches a model's vulnerabilities and failure modes before deployment&lt;/strong&gt;, which matters most in cybersecurity settings where data cannot leave the network.&lt;/p&gt;

&lt;h2&gt;
  
  
  The foundry approach
&lt;/h2&gt;

&lt;p&gt;Rather than building giant models from scratch, VIDRAFT pursues an &lt;strong&gt;AI foundry&lt;/strong&gt; strategy: diagnosing, combining, and transplanting knowledge into existing models to grow them for specific industries. Its open-source LLMs and their derivatives have surpassed &lt;strong&gt;2 million cumulative downloads&lt;/strong&gt; on Hugging Face, its on-device model POCKET ranked first among Korean models by trailing-30-day downloads, and the company holds 16 patents.&lt;/p&gt;

&lt;p&gt;"In fields like public services, defense, and finance — where data cannot be sent outside — trustworthy AI is the whole game," said Minsik Kim, CEO of VIDRAFT. "Our role is to diagnose model safety with AX-RAY and help build the trust layer of the nation's cybersecurity AI."&lt;/p&gt;

&lt;p&gt;MSIT plans to sign the agreement with the selected operator in September and begin development alongside the GPU support, forming a working-level council to drive results.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Project: Cybersecurity-Specialized AI Foundation Model (MSIT, Korea). Execution: Naver Cloud consortium (VIDRAFT participating). Scale: 10 months from Sep 2026, 256× B200 GPUs (5+5 staged). VIDRAFT contribution: AX-RAY AI safety diagnostics. Performance figures per Hugging Face API and public leaderboards as of 2026-09-03; project details per MSIT announcement.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>security</category>
      <category>opensource</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>VIDRAFT Launches "AI Foundry" Service: Custom On-Premises AI Cultivation for Enterprises and Public Institutions</title>
      <dc:creator>AI OpenFree</dc:creator>
      <pubDate>Wed, 02 Sep 2026 23:01:22 +0000</pubDate>
      <link>https://dev.to/ai_openfree_b23025ef075cf/vidraft-launches-ai-foundry-service-custom-on-premises-ai-cultivation-for-enterprises-and-public-2nk9</link>
      <guid>https://dev.to/ai_openfree_b23025ef075cf/vidraft-launches-ai-foundry-service-custom-on-premises-ai-cultivation-for-enterprises-and-public-2nk9</guid>
      <description>&lt;h1&gt;
  
  
  VIDRAFT Launches "AI Foundry" Service: Custom On-Premises AI Cultivation for Enterprises and Public Institutions
&lt;/h1&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; VIDRAFT, a Korean deep-tech AI startup, is launching an "AI Foundry" business that takes existing AI models and adapts them to customer-specific data and domain requirements — much like a semiconductor foundry handles contract manufacturing. The pipeline spans model diagnostics, domain-tuned optimization, inference acceleration, and safety verification, delivered as a turnkey on-premises solution. Engineers in regulated industries (public sector, defense, healthcare, finance) who need air-gapped or private-deployment AI should pay attention.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  What it is
&lt;/h2&gt;

&lt;p&gt;VIDRAFT's AI Foundry is a full-service, contract-based AI development and deployment pipeline aimed at organizations that cannot rely on cloud-hosted or off-the-shelf models due to data security requirements. Think public agencies, hospitals, financial institutions, and defense contractors — domains where data never leaves the internal network.&lt;/p&gt;

&lt;p&gt;Rather than building new foundation models from scratch for each customer, VIDRAFT's model is conceptually closer to the semiconductor foundry analogy: the customer supplies the data and the use-case requirements; VIDRAFT supplies the process, tooling, and expertise to produce a deployment-ready model tuned for that specific environment.&lt;/p&gt;

&lt;p&gt;The company frames its philosophy as &lt;em&gt;cultivating&lt;/em&gt; intelligence rather than &lt;em&gt;creating&lt;/em&gt; it — i.e., existing pre-trained models are the raw material, and the foundry process shapes them into production-grade, domain-specific assets.&lt;/p&gt;

&lt;p&gt;Revenue streams include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;On-premises deployment contracts&lt;/li&gt;
&lt;li&gt;Serving infrastructure licensing/usage fees&lt;/li&gt;
&lt;li&gt;Industry-specific OS build and operation&lt;/li&gt;
&lt;li&gt;Safety diagnostics and verification platform services&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;VIDRAFT is also participating in Korean government initiatives including the &lt;strong&gt;"AI for Everyone" (모두의 인공지능)&lt;/strong&gt; program and a government-backed &lt;strong&gt;secure foundation model project&lt;/strong&gt;, expanding its footprint in public-sector on-premises AI transformation.&lt;/p&gt;




&lt;h2&gt;
  
  
  How it works
&lt;/h2&gt;

&lt;p&gt;The AI Foundry pipeline is structured as a sequential set of proprietary stages:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Model MRI&lt;/strong&gt; — A diagnostic tool that analyzes the health, capability gaps, and weaknesses of an existing AI model before any optimization work begins. Think of it as a profiling/auditing step before you commit to fine-tuning.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Darwin · Chimera&lt;/strong&gt; — Technologies applied after diagnosis to adjust domain-specific performance. The naming suggests evolutionary (Darwin) and hybrid-combination (Chimera) approaches to model adaptation — conceptually covering techniques like domain-adaptive training and model merging/hybridization.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;VKAE · VKUE&lt;/strong&gt; — Serving acceleration technologies designed to reduce inference infrastructure overhead, making deployment more cost-efficient on customer hardware.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;AX-RAY&lt;/strong&gt; — A safety verification toolchain applied before final deployment to validate model behavior and flag potential risks.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Industry OS layer&lt;/strong&gt; — The final output is wired into purpose-built operating environments such as &lt;strong&gt;NationalOS&lt;/strong&gt; (public sector) and &lt;strong&gt;PharmaOS&lt;/strong&gt; (pharmaceutical/healthcare), providing vertical-specific runtime contexts rather than generic model endpoints.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This end-to-end process is backed by 16 patents and 6 published papers, representing VIDRAFT's IP portfolio underpinning the foundry offering.&lt;/p&gt;




&lt;h2&gt;
  
  
  Benchmarks &amp;amp; results
&lt;/h2&gt;

&lt;p&gt;VIDRAFT has published the following performance indicators (as stated in the source article):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;1.6 million cumulative downloads&lt;/strong&gt; on Hugging Face&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;#1 on the K-AI Leaderboard&lt;/strong&gt; (Korean AI benchmark leaderboard)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GPQA Diamond: 90.9%&lt;/strong&gt; — a graduate-level science Q&amp;amp;A benchmark measuring expert-level reasoning&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;#1 in the Verified category&lt;/strong&gt; of the &lt;strong&gt;Google &amp;amp; Hugging Face Fast Gemma Challenge&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Polaris drug discovery benchmark: #1 across 16 sub-tasks&lt;/strong&gt; — a significant result for pharma/biotech ML engineers evaluating the PharmaOS vertical&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Up to 23.4× throughput improvement&lt;/strong&gt; cited for their inference acceleration technology (self-measured)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These figures indicate competitive positioning particularly in reasoning-heavy tasks and domain-specific prediction benchmarks.&lt;/p&gt;




&lt;h2&gt;
  
  
  How to try it
&lt;/h2&gt;

&lt;p&gt;VIDRAFT's models are publicly accessible on Hugging Face with 1.6 million cumulative downloads recorded. You can browse and download their published models using the standard Hugging Face CLI:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;huggingface_hub
huggingface-cli download vidraft/&amp;lt;model-name&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;⚠️ Substitute &lt;code&gt;&amp;lt;model-name&amp;gt;&lt;/code&gt; with the specific model identifier listed on VIDRAFT's Hugging Face profile. The source article does not name a specific model for direct download, so check their page at &lt;a href="https://huggingface.co" rel="noopener noreferrer"&gt;huggingface.co&lt;/a&gt; by searching "VIDRAFT."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The AI Foundry pipeline itself (Model MRI, Darwin·Chimera, VKAE·VKUE, AX-RAY, and industry OS environments) is a &lt;strong&gt;commercial, contract-based service&lt;/strong&gt; — not a self-serve open-source tool. Access requires engaging VIDRAFT directly for enterprise or institutional deployment.&lt;/p&gt;

&lt;p&gt;No public GitHub repository, OpenAI-compatible API endpoint, or self-hosted demo URL was disclosed in the source article.&lt;/p&gt;




&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Q: Is this just fine-tuning-as-a-service, or is there something architecturally distinct going on?&lt;/strong&gt;&lt;br&gt;
A: The pipeline goes beyond standard fine-tuning. The published stages include diagnostic analysis (Model MRI), model hybridization or cross-pollination (Darwin·Chimera), hardware-aware inference optimization (VKAE·VKUE), and formal safety verification (AX-RAY) — suggesting a more structured, multi-phase engineering process than a typical LoRA fine-tune workflow.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Can I use NationalOS or PharmaOS independently of the full foundry contract?&lt;/strong&gt;&lt;br&gt;
A: Based on the source, these industry OS layers are the output of the full foundry pipeline rather than standalone products. They are delivered as part of an on-premises deployment engagement. Contact VIDRAFT directly for scoping.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: How does the 23.4× throughput figure compare to standard baselines?&lt;/strong&gt;&lt;br&gt;
A: The source cites this as a self-measured result without specifying the baseline hardware, model size, or comparison framework. Treat it as a directional data point and request third-party benchmark details before making procurement decisions.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally reported by 동아일보 (2026-09-01) — &lt;a href="https://www.donga.com/news/Economy/article/all/20260901/134582568/1" rel="noopener noreferrer"&gt;source article&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>opensource</category>
      <category>llm</category>
    </item>
    <item>
      <title>VIDRAFT Surpasses 1.6 Million Hugging Face Downloads: Darwin, AETHER, and the AI Foundry Approach Explained</title>
      <dc:creator>AI OpenFree</dc:creator>
      <pubDate>Wed, 02 Sep 2026 07:01:38 +0000</pubDate>
      <link>https://dev.to/ai_openfree_b23025ef075cf/vidraft-surpasses-16-million-hugging-face-downloads-darwin-aether-and-the-ai-foundry-approach-275e</link>
      <guid>https://dev.to/ai_openfree_b23025ef075cf/vidraft-surpasses-16-million-hugging-face-downloads-darwin-aether-and-the-ai-foundry-approach-275e</guid>
      <description>&lt;h1&gt;
  
  
  VIDRAFT Surpasses 1.6 Million Hugging Face Downloads: Darwin, AETHER, and the AI Foundry Approach Explained
&lt;/h1&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; Korean Pre-AGI startup VIDRAFT has crossed 1.6 million cumulative model downloads on Hugging Face, placing it among just four Korean organizations to exceed 1 million downloads. Their self-evolving Darwin model family leads multiple public benchmarks — including GPQA Diamond and the Polaris drug-prediction blind benchmark — while a from-scratch architecture called AETHER signals their push toward proprietary model design. Developers can explore their models and safety tooling directly on Hugging Face.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  What it is
&lt;/h2&gt;

&lt;p&gt;VIDRAFT (비드래프트, CEO Kim Min-sik) is a Korean Pre-AGI AI startup that develops and publishes large language and science-reasoning models. As of September 1, 2026, their cumulative Hugging Face model downloads have exceeded &lt;strong&gt;1.6 million&lt;/strong&gt; — a milestone that puts them alongside LG, Kakao, and Upstage as the only four Korean organizations to break the 1 million download threshold on the platform.&lt;/p&gt;

&lt;p&gt;Their public model portfolio spans several distinct lines:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Darwin series&lt;/strong&gt; — A self-evolving model family optimized for scientific reasoning and pharmaceutical prediction, including the published &lt;strong&gt;Darwin-398B-JGOS&lt;/strong&gt; variant.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AETHER&lt;/strong&gt; — A from-scratch foundation model (including the published &lt;strong&gt;Aether-7B-5Attn&lt;/strong&gt;) representing VIDRAFT's independent architecture research track, separate from fine-tuning or adaptation of existing base models.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AX-RAY&lt;/strong&gt; — An AI safety vulnerability diagnostic tool and accompanying evaluation datasets, publicly released on Hugging Face.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;BUILDING (빌딩)&lt;/strong&gt; — A free-to-access generative AI tool for architectural design.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The company also participates in Korean government AI initiatives and enterprise on-premises AI deployments.&lt;/p&gt;




&lt;h2&gt;
  
  
  How it works
&lt;/h2&gt;

&lt;p&gt;VIDRAFT describes its core competency as an &lt;strong&gt;"AI Foundry process"&lt;/strong&gt; — a manufacturing-style capability for building, adapting, and hardening AI models for specific domains and customers, rather than producing a single flagship model.&lt;/p&gt;

&lt;p&gt;At a conceptual level, their approach involves two complementary tracks:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Breeding and knowledge grafting (Darwin · Chimera):&lt;/strong&gt; Existing models are cross-trained and domain-adapted using specialized data — including MRI-based machine learning for medical diagnostics and prediction — to improve capability in targeted verticals. The "Chimera" technique name suggests structured hybridization of model capabilities across domains.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;From-scratch architecture research (AETHER):&lt;/strong&gt; In parallel, VIDRAFT develops models built entirely on their own architecture, allowing them to accumulate independent design expertise rather than remaining dependent on external base model releases.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Safety is a first-class engineering concern: &lt;strong&gt;AX-RAY&lt;/strong&gt; is their internal-turned-public diagnostic pipeline for identifying AI safety vulnerabilities, which can inform deployment hardening for enterprise customers.&lt;/p&gt;

&lt;p&gt;Their stated product roadmap extends these foundations toward:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Enterprise on-premises &lt;strong&gt;small LLMs (sLLM)&lt;/strong&gt; with reduced inference cost&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pocket on-device AI&lt;/strong&gt; models&lt;/li&gt;
&lt;li&gt;Vertically specialized operating-system-style platforms: &lt;strong&gt;NationalOS&lt;/strong&gt;, &lt;strong&gt;PharmaOS&lt;/strong&gt;, and &lt;strong&gt;MaterialsOS&lt;/strong&gt; (the last integrating generative AI with IBM quantum computer validation for novel materials discovery)&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Benchmarks &amp;amp; results
&lt;/h2&gt;

&lt;p&gt;All figures below come directly from the ZDNet Korea source article and refer to &lt;strong&gt;public leaderboards and externally verified competitions&lt;/strong&gt;:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Benchmark / Competition&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;GPQA Diamond&lt;/strong&gt; (graduate-level science reasoning)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;90.9%&lt;/strong&gt; — Darwin-398B-JGOS, ranked &lt;strong&gt;#1 on public leaderboard&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Polaris&lt;/strong&gt; (drug prediction external blind benchmark)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;#1 in 16 categories&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;K-AI Leaderboard&lt;/strong&gt; (operated by Korea MSIT &amp;amp; NIA)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;#1 ranking maintained&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Fast Gemma Challenge&lt;/strong&gt; (Google × Hugging Face global competition)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Top record achieved&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;CEO Kim Min-sik's stated philosophy on these results: &lt;em&gt;"Performance must be proven by public metrics and verification by the organizers, not by claims."&lt;/em&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  How to try it
&lt;/h2&gt;

&lt;p&gt;VIDRAFT's models and tooling are publicly available on Hugging Face. You can browse their organization page and download models using standard tooling:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Install the Hugging Face CLI&lt;/span&gt;
pip &lt;span class="nb"&gt;install &lt;/span&gt;huggingface_hub

&lt;span class="c"&gt;# Browse and download VIDRAFT models (replace &amp;lt;model-name&amp;gt; with the actual repo name)&lt;/span&gt;
huggingface-cli download vidraft/&amp;lt;model-name&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Note:&lt;/strong&gt; Specific model repository names and access conditions vary per release. Check the &lt;a href="https://huggingface.co/vidraft" rel="noopener noreferrer"&gt;VIDRAFT Hugging Face organization page&lt;/a&gt; for current availability, licensing, and any gated-access requirements. AX-RAY evaluation datasets are also published there.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;No public OpenAI-compatible API endpoint or GitHub organization URL was cited in the source article at time of writing.&lt;/p&gt;




&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Q: Is Darwin-398B-JGOS a fine-tune of an existing open model, or a novel architecture?&lt;/strong&gt;&lt;br&gt;
A: The source article does not explicitly specify the base architecture. The Darwin line is described as "self-evolving" and uses cross-breeding and knowledge grafting techniques; AETHER (e.g., Aether-7B-5Attn) is explicitly the from-scratch architecture track. Check the model cards on Hugging Face for lineage details.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: What is AX-RAY, and can I use it to evaluate my own models?&lt;/strong&gt;&lt;br&gt;
A: AX-RAY is VIDRAFT's AI safety vulnerability diagnostic tool, released publicly on Hugging Face along with its evaluation datasets. It is designed to surface safety weaknesses in language models. The public release suggests it can be applied to third-party models, but refer to the Hugging Face repository's documentation and license for specifics on scope and usage terms.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally reported by ZDNet Korea (2026-09-01) — &lt;a href="https://zdnet.co.kr/view/?no=20260901133742" rel="noopener noreferrer"&gt;source article&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>opensource</category>
      <category>llm</category>
    </item>
  </channel>
</rss>
