DEV Community

Cover image for span-01 vs mercury-decide: same score, opposite failures
sunnydachs
sunnydachs

Posted on

span-01 vs mercury-decide: same score, opposite failures

span-01 vs mercury-decide: same score, opposite failures

Last time I tested a "decision model" — a model that takes a plain-language question about a text and answers with a probability — as a gate for keeping Japanese narration free of English words. That article is here: Is regex enough? I tested span-01 on mixed-language text

Code and measured data:

GitHub logo sunnydachs / span01-eval

Evaluating a prompt-defined decision model (span-01-lite) as a language gate vs regex vs a chat model — every number recomputed from JSON artifacts

span01-eval

Evaluating prompt-defined "decision models" (span-01-lite, mercury-decide) as a language gate — against a plain regex and a generic chat model — with every number recomputed from saved JSON artifacts.

English | 日本語

Can a model that scores "does this text match a behavior described in plain language" replace a regex for keeping generated Japanese narration pure Japanese? This repo measures it honestly: a 23-case boundary suite, a real narration corpus, instruction-wording sweeps, calibration, and a threshold sweep — plus AUDIT.md, the record of having five independent models re-review every number before publishing.

The task

Detect English words mixed into Japanese narration (the kind a TTS voice reads aloud awkwardly). Three detectors are compared on identical inputs:

method what it is
regex adjacency + whitespace patterns against Latin tokens
gate respan/span-01-lite — a decision model: describe the behavior in prose, get a probability
chat a generic chat model answering
…

The verdict then was "regex first, model second", and the model I used was respan's span-01-lite. This time I measured a second model for the same job: Inception's mercury-decide.

If two models score the same, can you use them the same way?

Inception (the model measured here):

Inception – When Every Millisecond Matters

We are leveraging diffusion technology to develop a new generation of LLMs. Our dLLMs are much faster and more efficient than traditional autoregressive LLMs.

favicon inceptionlabs.ai

respan (the model from the previous article):

Respan - Route, monitor, and evaluate every agent run

The AI router with built-in observability & automated evals.

favicon respan.ai

The verdict first

On the 23-case boundary suite, the two models scored the same: F1 0.93 (F1 rolls misses and false positives into one number; 1.0 is perfect).

But the inside was the opposite. Four things came apart:

  • where they miss

  • the shape of the probabilities they return

  • how they react to instruction wording

  • how stable they are across days

What the second model actually is

mercury-decide is a "decision" model from Inception. It generates no text: give it a text and it returns a yes/no, a choice, or a score — each with a probability.

It is the same family as span-01-lite, and in my use it was called the same way.

What I measured

The same 23 boundary cases as the previous article: one each of an English word, an English sentence, katakana, a brand name, a personal name, a URL, an acronym, a code fragment and a role noun.

  • 10/1: 2 models × one instruction (23 × 2 = 46 calls)

  • 10/2: 2 models × two instruction wordings, short and fully spelled out (23 × 2 × 2 = 92 calls)

  • 10/3: re-ran the previous article's whole pipeline to see whether the first-round numbers reproduce

Result 1: same score, different misses

10/1, the same instruction, threshold 0.5:

             F1     misses   false pos.
 mercury    0.93    1        0
 span-01    0.93    1        0
Enter fullscreen mode Exit fullscreen mode

The scores tie, but the single miss is a different case each time.

  • mercury missed an English greeting sentence: "Hello everyone, welcome back to my channel." — probability 0.002.

  • span-01 missed a single embedded word, "compartments", at 0.24.

Each one catches what the other drops, so laying the two side by side got all 23 right.

Result 2: the probability shapes are opposites

The spread of the returned probabilities was plainly different.

  • mercury never goes above 0.03 on a negative. Almost black and white.

  • span-01 reaches 0.36 on a negative. It has a range.

So mercury scored the same at every threshold from 0.15 to 0.85, while span-01 was threshold-sensitive: dropping to 0.15 adds false positives, raising to 0.85 adds misses.

Result 3: opposite reactions to instruction wording

The same 23 cases, measured with a short instruction and with one that spells the conditions out:

             short instr.   spelled out
 mercury     0.67           0.80
 span-01     0.67           0.93
Enter fullscreen mode Exit fullscreen mode

Both improve when you write more, but they improve differently.

  • span-01 simply lost its false positives: all 4 verdicts that moved were improvements.

  • mercury moved 9 — and one case it started missing once the instruction got longer.

Listing the exclusion conditions appears to have made it drop positive cases that were not on the list.

Result 4: a day later, the answers moved

Same texts, same instruction, measured again the next day:

             verdicts flipped
 mercury     2 of 23 (F1 0.93 → 0.80)
 span-01     0 of 23
Enter fullscreen mode Exit fullscreen mode
  • mercury had correctly called a katakana word negative the day before; the next day it called it positive at 0.92.

  • A role noun it had caught the day before, it dropped at 0.004 the next day.

Within a single day, repeated calls return the same probability (I checked three times) — deterministic only within the same day.

span-01 did not change a single verdict across four days (9/27, 10/1, 10/2, 10/3).

Summary: how to use it

If you pick one: spell the instruction out and use span-01. It reacts plainly, and it returns the same answer across days.

If in doubt, run both. They miss different things, so one covers the other.

 [script]
    ├─▶ model A ─┐
    │            ├─ either above threshold → needs a fix
    └─▶ model B ─┘
Enter fullscreen mode Exit fullscreen mode

That said, this is a 23-case ruler.

Honest limitations

  • 23 cases, one run each. It says nothing about accuracy on real scripts.

  • The cross-day flip is 2 cases. "A one-day fluke" cannot be fully ruled out.

  • The instruction is reused from the previous article, not tuned for mercury.

  • The ground-truth labels are my own judgement.

Closing

"The same accuracy" was not "the same behaviour".

When picking a decision model, it seems worth looking not only at the score, but at how it reacts to instruction wording and how stable it is across days.

Code and measured data: github.com/sunnydachs/span01-eval

This is a personal OSS project, so there is no warranty. Use it at your own risk. Bug reports and improvement ideas are welcome as issues.

Top comments (0)