DEV Community

QuantID
QuantID

Posted on

Reading the S1MB result: why Borda aggregation makes a #1 harder to fake

The System One Mosaic Benchmark (S1MB, by hotchpotch) compares 102 models across 137 specialized benchmarks in three task families: Noul (assess a condition), Choice (select an option), and Score (rate on a scale). This is a measurement note on how it scores and why its top position is informative.

Borda, not average

S1MB ranks by a Borda score. Instead of averaging raw task numbers, which lets one easy benchmark dominate the mean, Borda converts each of the 137 benchmarks into a rank and sums rank-based points. That rewards broad, consistent strength over a spike on a few favorable tasks. It is a deliberately conservative aggregation, and it makes a top position harder to manufacture: you cannot buy first place by overfitting a handful of suites.

The current top

Darwin-27B-ZTC-v2 leads with Borda 89.58, ahead of OpenJev-27B (87.50) and AutoJev-27B (87.07). The margin is widest on the Score family (60.21 vs 54.83), which is the ordinal-judgment task where most judges weaken. Leading where the task is hardest, rather than on an easy split, is the signal worth reading.

Why a deterministic judge is clean to measure

A zero-token judge emits a probability distribution in a single forward pass. No sampling means no output variance, so the reported numbers are exact values rather than sample estimates. There is no standard-deviation term to carry and no decoding hyperparameters (temperature, top-p, length) to confound cross-model comparison. The same input reproduces the same score.

On a related board, typed-decisions, the same approach is also #1 at accuracy 0.743 (zero-shot), with calibration KL 0.204 and Brier 0.097. Brier is a strictly proper scoring rule, so confidence cannot be gamed by inflating or deflating it.

Links

Disclosure: Darwin-27B-ZTC is built by VIDRAFT, a technology partner of QuantID. Public leaderboard standings move as new models are added; numbers reflect the board at the time of writing.

Top comments (0)