DEV Community

AI OpenFree
AI OpenFree

Posted on Originally published at vidraft-ai-media.static.hf.space

Open 180B Model Leads 10 Official Hugging Face Leaderboards, and the Zero-Token Judge Behind It

Most Hugging Face official #1 boards by org, October 2026

TL;DR

Darwin-180B-RSI, an open-weight model built by VIDRAFT on top of Alibaba's Qwen3.8-Flash-Next, now holds the top rank on 10 official Hugging Face leaderboards. That is the most of any organization on a stage of 48 curated benchmarks and 95 competing orgs. The runner-up org holds 4. The newest two wins are enterprise-facing rather than exam-facing: ExtractBench (90.29) for pulling fields out of long PDFs, and IFStruct (98.95) for producing output that follows a strict schema. Separately, Gate Arcade, a demo that judges whether an agent action should run, was picked for Hugging Face Spaces of the Week. This post lists every board with numbers, explains the training idea in plain terms, and goes deep on ZTC, a judging method that reads a model's internal state and scores an action in 0.06 seconds with no generated tokens.

Which Hugging Face leaderboards does the model lead?

Ten official boards, spanning science, math, vision, law, enterprise tasks, and document parsing. All numbers below are the registered scores, each measured by the respective submitter using the board's official harness.

Board Domain Score Field count
GPQA Diamond Science and knowledge 94.44 111 models
MMLU-Pro Science and knowledge 88.12 141 models
MMMU-Pro Visual reasoning 79.48
AIME 2026 Math 100
HMMT 2026 Math 100
LEXam Law 68.94
LEXam-hard Law 45.72
ExtractBench Enterprise, PDF extraction 90.29
IFStruct Enterprise, structured output 98.95
MDPBench Document parsing 83.65

The two most crowded boards, MMLU-Pro (141 models) and GPQA (111 models), are both held, so this is not a thin-field result. By org, the tally of official #1 boards is FINAL-Bench (the VIDRAFT research org) 10, Z.ai 4, Xiaomi 3, DeepSeek 3, Moonshot 3.

Why do the two newest wins matter more than another exam score?

Most leaderboard wins reward exam-style reasoning. The two newest wins reward the things a team actually hits first when it puts a model into a product.

IFStruct, built by Liquid AI, checks whether a model emits data in the exact requested shape (JSON or YAML). It fails a response if a single field is missing, mistyped, out of range, or invented. The model passed 1,979 of 2,000 items for 98.95, ahead of the prior leader Agents-A1 (93.25), gpt-oss-20b (91.95), and Nemotron-3-Nano-30B (86.80). In a pipeline where a program reads the model output and passes it to the next step, one wrong character stalls the job, so schema adherence is a gating requirement for adoption, not a nicety.

ExtractBench, built by LlamaIndex, measures pulling the right fields out of 370 PDFs ranging from a few pages to roughly 190 pages. The score of 90.29 came from Darwin-180B-RSI-R3, a further self-improved revision, which edged past its own base model, Qwen3.8-Flash-Next (89.88). The model was never specifically trained on legal text or document extraction, yet it leads both, which points to transfer rather than task-specific tuning.

How was the model trained?

The method is model-level recursive self-improvement (RSI). The model solves verifiable problems on its own, keeps only the self-generated solutions that check out as correct, and trains on that filtered set. Each round makes the model a little stronger, and because unverified solutions are discarded, errors are not reinforced into the weights. No human-written answers or worked solutions are needed for the loop. The spillover into law and document extraction, areas never taught directly, is the practical signal that careful reasoning habits learned on checkable problems carry over.

A short, deliberately honest caveat: leaderboard scores reflect the board's own harness and submission, and different boards normalize differently, so treat a cross-board ranking as a portfolio view, not a single global IQ number.

What is ZTC, and how does it judge an action without generating text?

As agents start sending mail, making payments, and running code on a person's behalf, catching a wrong action before it executes has become its own engineering problem. The common approach is to ask a second model, in words, whether an action is safe, and read back its written verdict. That second pass costs tokens and latency, and it ships data to an external API.

ZTC (Zero-Token Confidence) takes a different route. Instead of asking another model to write a judgment, it reads the internal state of the model that produced the action once, and computes the probability that the action is correct. The number of extra generated tokens is zero.

Measured on the Gate Arcade setup (closed-network deployment, snapshot of 2026-09-24):

Metric External API judge (JEV) ZTC
Judgment accuracy (AUC) 0.7350 0.7289
Time per judgment 0.591 s 0.0615 s
Judgments in 300 s 135 177

The accuracy gap (0.7350 vs 0.7289) is statistically a wash, while ZTC is about 10x faster per call. Over a 2,000-round head-to-head, ZTC won 73.7 percent of rounds on the combined speed-and-accuracy scoring.

Two properties follow from the design:

  1. Data stays local. ZTC runs on the GPU next to the model, with no round trip to an external API, so it works on air-gapped networks, which suits defense, finance, and public-sector deployments.
  2. The marginal cost is near zero. It reads the internals of a model that already produced the answer. Using this approach, a separate model raised the accuracy of judging its own answers from 0.76 to 0.88.

A minimal mental model of the call path looks like this:

# Conceptual sketch, not the production path.
logits, hidden = model.forward(prompt, return_internals=True)
action = decode(logits)              # the candidate action
confidence = ztc_head(hidden)        # one read, zero new tokens
if confidence < threshold:
    hold(action)                     # defer instead of executing
else:
    execute(action)
Enter fullscreen mode Exit fullscreen mode

The judging space is getting crowded. Typesafe's JEV API set a de facto interface, the public Decision Index board now lists more than 70 models, Cloudflare shipped a JEV-compatible judge that crossed 1,100 Hugging Face likes within days, and research groups at Stanford and NVIDIA have published related models. ZTC's distinguishing bet is reading inside the model rather than asking it to speak.

FAQ

Are these official Hugging Face leaderboards or self-reported numbers?
They are official, curated boards. Each score is measured by the submitter with the board's official harness and then registered, which is the standard practice across these leaderboards.

Is the model open weight?
Yes. The base and the R3 revision are published under the FINAL-Bench org on Hugging Face, so the results are reproducible rather than API-gated.

Does ZTC replace a guardrail or policy model?
It complements them. ZTC is a fast, local confidence signal on whether an action is likely correct, so it is a cheap first gate that can defer uncertain actions before a heavier policy check runs.

How can ZTC be almost as accurate as an external judge while being 10x faster?
Because it skips generation entirely. An external judge has to produce tokens to express a verdict, while ZTC reads the hidden state the forward pass already computed, so it trades a small accuracy margin for a large latency win.

Why did a model win law and document extraction without domain training?
The RSI loop trains on verified reasoning over checkable problems, and those habits transfer. Leading LEXam and ExtractBench without targeted tuning is the observable evidence of that transfer.

What is the single most important takeaway?
The shift from exam scores to enterprise tasks. Leading ExtractBench and IFStruct means the model is measured on reading documents and emitting correct, schema-valid output, which is what production systems need.

Further reading

  • VIDRAFT Engineering Notes: Hugging Face official benchmarks, the complete list of 48 and how they are scored.
  • VIDRAFT Engineering Notes: Darwin-180B-RSI, how recursive self-improvement topped the leaderboards.

Top comments (0)