span-01 vs mercury-decide: same score, opposite failures
Last time I tested a "decision model" — a model that takes a plain-language question about a text and answers with a probability — as a gate for keeping Japanese narration free of English words. That article is here: Is regex enough? I tested span-01 on mixed-language text
Code and measured data:
sunnydachs
/
span01-eval
Evaluating a prompt-defined decision model (span-01-lite) as a language gate vs regex vs a chat model — every number recomputed from JSON artifacts
span01-eval
Evaluating prompt-defined "decision models" (span-01-lite, mercury-decide) as a language gate — against a plain regex and a generic chat model — with every number recomputed from saved JSON artifacts.
English | 日本語
Can a model that scores "does this text match a behavior described in plain language" replace a regex for keeping generated Japanese narration pure Japanese? This repo measures it honestly: a 23-case boundary suite, a real narration corpus, instruction-wording sweeps, calibration, and a threshold sweep — plus AUDIT.md, the record of having five independent models re-review every number before publishing.
The task
Detect English words mixed into Japanese narration (the kind a TTS voice reads aloud awkwardly). Three detectors are compared on identical inputs:
| method | what it is |
|---|---|
| regex | adjacency + whitespace patterns against Latin tokens |
| gate |
respan/span-01-lite — a decision model: describe the behavior in prose, get a probability |
| chat | a generic chat model answering |
The verdict then was "regex first, model second", and the model I used was respan's span-01-lite. This time I measured a second model for the same job: Inception's mercury-decide.
If two models score the same, can you use them the same way?
Inception (the model measured here):
respan (the model from the previous article):
The verdict first
On the 23-case boundary suite, the two models scored the same: F1 0.93 (F1 rolls misses and false positives into one number; 1.0 is perfect).
But the inside was the opposite. Four things came apart:
where they miss
the shape of the probabilities they return
how they react to instruction wording
how stable they are across days
What the second model actually is
mercury-decide is a "decision" model from Inception. It generates no text: give it a text and it returns a yes/no, a choice, or a score — each with a probability.
It is the same family as span-01-lite, and in my use it was called the same way.
What I measured
The same 23 boundary cases as the previous article: one each of an English word, an English sentence, katakana, a brand name, a personal name, a URL, an acronym, a code fragment and a role noun.
10/1: 2 models × one instruction (23 × 2 = 46 calls)
10/2: 2 models × two instruction wordings, short and fully spelled out (23 × 2 × 2 = 92 calls)
10/3: re-ran the previous article's whole pipeline to see whether the first-round numbers reproduce
Result 1: same score, different misses
10/1, the same instruction, threshold 0.5:
F1 misses false pos.
mercury 0.93 1 0
span-01 0.93 1 0
The scores tie, but the single miss is a different case each time.
mercury missed an English greeting sentence: "Hello everyone, welcome back to my channel." — probability 0.002.
span-01 missed a single embedded word, "compartments", at 0.24.
Each one catches what the other drops, so laying the two side by side got all 23 right.
Result 2: the probability shapes are opposites
The spread of the returned probabilities was plainly different.
mercury never goes above 0.03 on a negative. Almost black and white.
span-01 reaches 0.36 on a negative. It has a range.
So mercury scored the same at every threshold from 0.15 to 0.85, while span-01 was threshold-sensitive: dropping to 0.15 adds false positives, raising to 0.85 adds misses.
Result 3: opposite reactions to instruction wording
The same 23 cases, measured with a short instruction and with one that spells the conditions out:
short instr. spelled out
mercury 0.67 0.80
span-01 0.67 0.93
Both improve when you write more, but they improve differently.
span-01 simply lost its false positives: all 4 verdicts that moved were improvements.
mercury moved 9 — and one case it started missing once the instruction got longer.
Listing the exclusion conditions appears to have made it drop positive cases that were not on the list.
Result 4: a day later, the answers moved
Same texts, same instruction, measured again the next day:
verdicts flipped
mercury 2 of 23 (F1 0.93 → 0.80)
span-01 0 of 23
mercury had correctly called a katakana word negative the day before; the next day it called it positive at 0.92.
A role noun it had caught the day before, it dropped at 0.004 the next day.
Within a single day, repeated calls return the same probability (I checked three times) — deterministic only within the same day.
span-01 did not change a single verdict across four days (9/27, 10/1, 10/2, 10/3).
Summary: how to use it
If you pick one: spell the instruction out and use span-01. It reacts plainly, and it returns the same answer across days.
If in doubt, run both. They miss different things, so one covers the other.
[script]
├─▶ model A ─┐
│ ├─ either above threshold → needs a fix
└─▶ model B ─┘
That said, this is a 23-case ruler.
Honest limitations
23 cases, one run each. It says nothing about accuracy on real scripts.
The cross-day flip is 2 cases. "A one-day fluke" cannot be fully ruled out.
The instruction is reused from the previous article, not tuned for mercury.
The ground-truth labels are my own judgement.
Closing
"The same accuracy" was not "the same behaviour".
When picking a decision model, it seems worth looking not only at the score, but at how it reacts to instruction wording and how stable it is across days.
Code and measured data: github.com/sunnydachs/span01-eval
This is a personal OSS project, so there is no warranty. Use it at your own risk. Bug reports and improvement ideas are welcome as issues.
inceptionlabs.ai
Top comments (0)