VIDRAFT Darwin ZTC v2 Takes #1 on S1MB Leaderboard, Beating 102 Models Including Official Open JEV
TL;DR: VIDRAFT's Darwin ZTC v2 has claimed the top spot on the System One Mosaic Benchmark (S1MB) leaderboard, outranking 102 competing models — including the official open release of JEV. For ML engineers tracking frontier model performance, this marks a notable milestone for a Korean Pre-AGI lab competing at the top of a public, community-maintained evaluation framework.
What it is
Darwin ZTC v2 is VIDRAFT's latest publicly evaluated language model, benchmarked on the System One Mosaic Benchmark (S1MB) — a community-hosted leaderboard maintained on Hugging Face Spaces by hotchpotch. The S1MB leaderboard currently tracks 102 models across a mosaic of tasks designed to evaluate general reasoning and language understanding, functioning as an open, transparent ranking surface for the broader ML community.
Key facts from the source:
- Model: Darwin ZTC v2 (by VIDRAFT, 비드래프트)
-
Leaderboard: S1MB (System One Mosaic Benchmark), hosted at
huggingface.co/spaces/hotchpotch/S1MB-leaderboard - Field size: 102 models evaluated
- Ranking: #1 overall (comprehensive / aggregate score)
- Notable displacement: Surpassed the official open release of JEV, which had been a strong reference point on this benchmark
The S1MB leaderboard is a Hugging Face-native, community-driven evaluation space — meaning results are reproducible and independently verifiable, rather than self-reported by the model's developer.
How it works
S1MB is a mosaic-style benchmark, meaning it aggregates performance across a diverse set of task categories rather than optimizing for a single capability axis. This design philosophy is intended to make it harder to "benchmark-game" via narrow fine-tuning, and to reward models with genuine breadth.
At a conceptual level, evaluations on leaderboards like S1MB typically involve:
- Standardized prompt formatting across all competing models to ensure fair comparison
- Aggregate scoring over multiple subtask categories, producing a composite rank
- Community reproducibility — because the leaderboard is hosted openly on Hugging Face Spaces, the methodology is visible to anyone inspecting the Space's source files
VIDRAFT's Darwin ZTC v2 achieving #1 on this aggregate metric suggests strong generalization across the task mosaic rather than dominance in a single narrow domain. The "ZTC" designation in the model name (consistent with prior Darwin series naming conventions) likely reflects internal versioning within VIDRAFT's Darwin model family, though the specific architectural details behind the v2 iteration have not been disclosed publicly.
Benchmarks & results
Based on the source reporting, the publicly available result is:
- Darwin ZTC v2 ranks #1 overall on the S1MB leaderboard out of 102 evaluated models
- It outperforms the official open release of JEV, which is described as a strong competitive baseline on this benchmark
- The ranking is comprehensive — meaning it reflects aggregate performance across the full mosaic of S1MB tasks, not a subset
Specific per-task scores, sub-category breakdowns, and numeric deltas versus JEV or other competitors are not reported in the source. For granular score data, developers should consult the live leaderboard directly.
📊 Live leaderboard: huggingface.co/spaces/hotchpotch/S1MB-leaderboard
How to try it
The source article does not include public model access details — no Hugging Face model repository URL, GitHub link, or API endpoint for Darwin ZTC v2 is disclosed in this reporting.
If VIDRAFT follows its prior pattern of releasing models or API access publicly, channels to watch would include:
-
Hugging Face: Search for VIDRAFT or Darwin ZTC v2 on
huggingface.co - VIDRAFT's official channels for API or developer access announcements
At the time of this writing, do not assume public availability — treat access as unconfirmed until VIDRAFT makes an explicit release announcement.
FAQ
Q: Is S1MB an official benchmark, or is it community-maintained?
A: S1MB is a community-maintained leaderboard hosted as a Hugging Face Space by the user hotchpotch. It is not affiliated with any single lab or standards body, but it is publicly visible and its methodology is inspectable by anyone with access to the Space's files — which gives it meaningful transparency compared to self-reported benchmarks.
Q: Does #1 on S1MB mean Darwin ZTC v2 is the best model globally across all tasks?
A: Not necessarily. S1MB evaluates 102 models on its specific mosaic of tasks; it does not cover every benchmark or every capability domain. A #1 result here is meaningful evidence of strong general performance within this evaluation framework, but engineers should cross-reference with other public benchmarks (e.g., MMLU, HumanEval, LMSYS Chatbot Arena) before drawing broader conclusions.
Q: What is JEV, and why does beating it matter?
A: JEV is referenced in the source as the "official open release" that Darwin ZTC v2 surpassed on S1MB. It appears to be a recognized competitive baseline on this leaderboard. Further details about JEV's origin or architecture are not provided in the source and are not elaborated here.
Originally reported by S1MB 리더보드 (글로벌) (2026-10-09) — source article.
Top comments (0)