DEV Community

AI OpenFree
AI OpenFree

Posted on Originally published at huggingface.co

Darwin-27B-ZTC-v2 is #1 on the S1MB decision-engine leaderboard

Decision engines are a fast-growing category: models that read a piece of state and a typed question, then return a probability for every option. S1MB (System One Mosaic Benchmark) is the widest public comparison of them, with 137 English benchmarks across three question types: yes or no (Noul), multiple choice (Choice) and ordered rubrics (Score).

Our new model, Darwin-27B-ZTC-v2, is now #1 on S1MB. The S1MB maintainer validated and merged our results, and the model ranks first under both of the leaderboard's rankings.

The numbers

Computed with the leaderboard's own scoring code on the published results file, over the 102 models with all 137 benchmarks complete:

Rank Model Borda score Task Avg Noul Choice Score
1 Darwin-27B-ZTC-v2 89.58 66.46 67.66 71.51 60.21
2 openjev 87.50 62.60 65.22 67.74 54.83
3 AutoJev-27B 87.07 60.80 64.96 68.11 49.34
4 Eikos-27B 85.43 59.86 64.14 67.61 47.83
5 TypeSafe Jev 1.13 85.05 59.59 64.63 67.22 46.92
  • Borda score (the default sort) ranks every model on every benchmark, gives 100 points to first place and 0 to last, and averages over the 137 benchmarks.
  • Task Avg averages the baseline-adjusted score within each question type, then across the three types.

Darwin-27B-ZTC-v2 leads in all three question types. The biggest gap is on ordered rubrics (Score), where it reaches 60.21 against 54.83 for the next model.

How it works

ZTC stands for zero-token classification. The model lists the options in its prompt, runs one forward pass, and reads a probability for every option from its final hidden state. It generates no tokens, so a decision costs a single pass of a 27B model.

v2 builds on Darwin-27B-ZTC, which is #1 on the Hugging Face official Typed Decisions leaderboard (0.743, zero-shot). For v2 we continued training on public training splits of control and workflow decision tasks, then averaged the weights with v1 50/50. The average kept v1's skills and added the new ones.

Reproducibility

  • Weights, inference code and a /v1/systemone server are public on the model page.
  • The S1MB run used the evaluator's existing autojev adapter with no new adapter code.
  • Training used only public training splits. We checked every S1MB test case against our training rows and found zero matches.

Model: https://huggingface.co/FINAL-Bench/Darwin-27B-ZTC-v2
Leaderboard: https://huggingface.co/spaces/hotchpotch/S1MB-leaderboard

Top comments (0)