DEV Community

ryuya
ryuya

Posted on Originally published at qiita.com

I swapped LLM scoring for a non-generative model. The score now moves by 0.01 out of 5

TL;DR

  • I replaced LLM-based scoring with a split: judgment = Jev, arithmetic = code, feedback text = LLM
  • Run the same answer 10 times and the score moves by 0.01 on a 0-5 scale (SD, mean of 30 items)
  • Median latency 249 ms. Cost per 1,000 scorings dropped from $0.16 to $0.045
  • All numbers below come from one log: 30 answers x 10 runs = 300 scorings on jev-1.13.0

Why not an LLM

I run a small vocabulary app. It asks you to write a short English answer, then scores it and estimates a CEFR level.

For scoring, being stable matters more than being right. A learner who submits the same sentence twice and sees two different levels stops trusting the app immediately.

LLM scoring gave me three problems.

Problem What it looked like
Scores drift Lowering temperature and adding few-shot examples did not stop it
Scores are coarse Asked for 0-5, got a pile at 3 and 4. No difference between 2.6 and 3.1
Slow Five criteria in five calls meant five round trips. One call made the criteria drag each other around

Jev is a System One model from TypeSafe AI. It does not generate text. You send state and typed questions, and it returns typed answers with probabilities and a confidence value.

It has three question types: Choice, Score and Noul. Score takes a rubric with ordered levels and returns a continuous value. You can ask for several criteria in a single round trip, which killed all three problems at once.


What I built

Five criteria, scored on every answer:

grammatical accuracy / lexical range / syntactic complexity / task achievement / naturalness

The weights are tuned for my content, so I am keeping those private. The structure is the part worth sharing.

Design

The rule I settled on: Jev judges, code computes, the LLM writes prose.

flowchart LR
    A[Learner answer] --> B["Jev<br/>score 5 criteria"]
    A --> E["LLM<br/>write feedback"]
    B --> C["App code<br/>weighted sum<br/>map to CEFR band"]
    C --> D["Shown instantly<br/>~0.2 s"]
    E --> F["Appended<br/>a few seconds later"]
const res = await jev.systemOne({
  model: "jev-1.13.0",
  input: { task: taskPrompt, answer: answerText },
  primitives: [
    { type: "score", name: "grammar_accuracy",     rubric: RUBRIC.grammar },
    { type: "score", name: "vocabulary_range",     rubric: RUBRIC.vocabulary },
    { type: "score", name: "syntactic_complexity", rubric: RUBRIC.syntax },
    { type: "score", name: "task_achievement",     rubric: RUBRIC.task },
    { type: "score", name: "naturalness",          rubric: RUBRIC.naturalness },
  ],
});

// the sum and the band mapping stay in code, always
const total = weightedSum(res.scores);
const band  = toBand(total);
Enter fullscreen mode Exit fullscreen mode

Mapping a score to a CEFR band is deterministic. There is no reason to spend a model call on arithmetic, and every reason not to: code gives you the same answer forever.

Scoring and feedback are separate endpoints. Scoring returns in about a quarter of a second, so the learner sees a number long before the prose arrives.


The numbers

I wrote 30 reference answers, 5 per CEFR level from A1 to C2, and ran each one 10 times.

Do the levels come out in order

Level n Mean SD Min Max
A1 50 0.86 0.103 0.66 0.99
A2 50 1.79 0.127 1.60 1.96
B1 50 3.09 0.102 2.90 3.22
B2 50 4.13 0.130 3.91 4.35
C1 50 4.63 0.086 4.51 4.78
C2 50 4.84 0.052 4.72 4.89

Every one of the 5 question sets climbed from A1 to C2 without a single inversion.

The gaps between neighbouring levels:

Pair Gap Distributions overlap
A1 - A2 0.93 no
A2 - B1 1.30 no
B1 - B2 1.04 no
B2 - C1 0.50 no
C1 - C2 0.21 yes

C1 and C2 do not separate. Advanced writing pins every criterion near the ceiling: C2 grammar scores landed between 4.92 and 4.98 across all 50 runs.

How much does the same answer move

This is the part I cared about.

Metric Measured
SD of the total, mean over 30 answers 0.0105
SD of the total, worst answer 0.0215
Max minus min, mean 0.032
Max minus min, worst 0.070

0.01 on a 0-5 scale. I never got close to that with an LLM.

Worth saying plainly: not one of the 30 answers returned an identical value 10 times out of 10. Jev is not deterministic. It barely moves, which is a different property.

Per criterion the movement is larger:

Criterion Mean SD
Grammatical accuracy 0.0166
Lexical range 0.0181
Syntactic complexity 0.0190
Task achievement 0.0194
Naturalness 0.0211

Around 0.02 each, and 0.0105 once they are weighted and summed. The criteria wobble in uncorrelated directions and the sum cancels most of it.

Latency

300 calls:

Measured
p50 249 ms
p90 419 ms
p95 580 ms
Max 1,956 ms
Under 300 ms 236 / 300 (79%)

Fast enough that the score is on screen before the learner looks up.


Three things that bit me

1. Short correct sentences scored too well

I like apples. was getting a high score. Nothing is wrong with it, and that was the problem: my rubric asked "is this grammatically correct".

I rewrote the rubric levels from "what is being measured" to "what difficulty was attempted, and was it pulled off".

To check it, I wrote pairs of answers to the same prompt, one deliberately simple (S) and one deliberately complex (C):

Prompt Simple (S) Complex (C) Delta
Q1 1.68 2.37 +0.70
Q2 1.87 2.70 +0.83
Q3 1.87 2.08 +0.21
Q4 1.66 2.32 +0.66
Q5 1.98 2.25 +0.28
Mean +0.54

Every prompt moved the right way.

2. A stable score still produces an unstable label

The score moves by 0.01, and yet 2 of the 30 answers changed band across the 10 runs.

Answer Band flip Score range
An A2 answer A2.1 to A2.2 1.88 - 1.91
A C2 answer C1 to C2 4.72 - 4.77

Both sit almost exactly on a band boundary. A 0.01 wobble is irrelevant to the score and decisive to the label the learner reads.

The fix is not in the model. I added hysteresis: moving up and moving down use different thresholds, so the displayed level stops flickering at a boundary. I have not re-run the full calibration since adding it, so treat that as a design change rather than a measured result.

3. Pin the model version

If you point at a floating latest tag, everyone's score shifts on the day the weights change, and your history becomes incomparable. For a learning app that is fatal. I pin jev-1.13.0.

On top of that, a script replays the reference set and fails the build if any answer lands in a different band than the recorded baseline.


When it breaks

A wrong score is worse than a missing one, so everything degrades toward showing nothing.

flowchart TD
    A[Answer received] --> B{Written in the target language?}
    B -- no --> X[Skip scoring]
    B -- yes --> C{Jev responded?}
    C -- no --> Y["Hide the score<br/>show feedback only"]
    C -- yes --> D{Any low-confidence criterion?}
    D -- yes --> E["Hide that criterion<br/>keep it in the total"]
    D -- no --> F[Show everything]

Confidence comes back from Jev and drops when its judgment is spread out. Across the 300 calls (1,500 criterion scores) 45, or 3%, were hidden.

This is the other reason scoring and feedback are separate endpoints. Scoring can fail without taking the feedback down with it.


Cost

Per 1,000 scorings
Scoring with Jev ~$0.045
Asking an LLM for the score ~$0.16
Speech recognition $0 (on-device Web Speech API)

Jev returns numbers, so there is no output token bill. The whole 300-call calibration run cost about one US cent, which is why re-running everything 10 times per answer was an easy decision.

My billing is in yen, so these are converted at roughly ¥155 to the dollar.

Before this I capped how often a free user could be scored. That cap is gone.


Takeaway

Text that varies between runs is fine. A score that varies between runs is not.

I was asking one model to do both and getting the worse half of each. Splitting the judgment out into Jev and keeping the arithmetic in code fixed it. If you have an LLM producing numbers in production, it is worth looking at.


I build Chunkbook, the vocabulary app this scoring runs in. Originally published in Japanese on Qiita.

Top comments (0)