Decision engines are a fast-growing category: models that read a piece of state and a typed question, then return a probability for every option. S1MB (System One Mosaic Benchmark) is the widest public comparison of them, with 137 English benchmarks across three question types: yes or no (Noul), multiple choice (Choice) and ordered rubrics (Score).
Our new model, Darwin-27B-ZTC-v2, is now #1 on S1MB. The S1MB maintainer validated and merged our results, and the model ranks first under both of the leaderboard's rankings.
The numbers
Computed with the leaderboard's own scoring code on the published results file, over the 102 models with all 137 benchmarks complete:
| Rank | Model | Borda score | Task Avg | Noul | Choice | Score |
|---|---|---|---|---|---|---|
| 1 | Darwin-27B-ZTC-v2 | 89.58 | 66.46 | 67.66 | 71.51 | 60.21 |
| 2 | openjev | 87.50 | 62.60 | 65.22 | 67.74 | 54.83 |
| 3 | AutoJev-27B | 87.07 | 60.80 | 64.96 | 68.11 | 49.34 |
| 4 | Eikos-27B | 85.43 | 59.86 | 64.14 | 67.61 | 47.83 |
| 5 | TypeSafe Jev 1.13 | 85.05 | 59.59 | 64.63 | 67.22 | 46.92 |
- Borda score (the default sort) ranks every model on every benchmark, gives 100 points to first place and 0 to last, and averages over the 137 benchmarks.
- Task Avg averages the baseline-adjusted score within each question type, then across the three types.
Darwin-27B-ZTC-v2 leads in all three question types. The biggest gap is on ordered rubrics (Score), where it reaches 60.21 against 54.83 for the next model.
How it works
ZTC stands for zero-token classification. The model lists the options in its prompt, runs one forward pass, and reads a probability for every option from its final hidden state. It generates no tokens, so a decision costs a single pass of a 27B model.
v2 builds on Darwin-27B-ZTC, which is #1 on the Hugging Face official Typed Decisions leaderboard (0.743, zero-shot). For v2 we continued training on public training splits of control and workflow decision tasks, then averaged the weights with v1 50/50. The average kept v1's skills and added the new ones.
Reproducibility
- Weights, inference code and a
/v1/systemoneserver are public on the model page. - The S1MB run used the evaluator's existing
autojevadapter with no new adapter code. - Training used only public training splits. We checked every S1MB test case against our training rows and found zero matches.
Model: https://huggingface.co/FINAL-Bench/Darwin-27B-ZTC-v2
Leaderboard: https://huggingface.co/spaces/hotchpotch/S1MB-leaderboard
Top comments (0)