DEV Community

Vijay Kumar Sridharan
Vijay Kumar Sridharan

Posted on

Critical-entity error rate: the metric WER can't see

Part 2 of the CAFA-IVR series.

In part 1, WER scored "peas" and a ten-thousand-dollar mis-transfer as the same error. This post is about the machinery for telling them apart.

The errors that break IVR tasks don't distribute evenly across the vocabulary. They cluster in a tiny governed set: amounts, dates, account numbers, yes/no decisions, negation, cancel/confirm actions. A recognizer can mangle every filler word in a call and the task still completes; corrupt one governed token and the task completes wrong — which is worse than not completing at all.

CEER — critical-entity error rate — counts corruption in that governed set, and nothing else.

What it measures

For each trial, the framework compares two multisets: the critical entities the reference should have produced against the ones the speech path actually produced. Entities are (type, normalized value) pairs — ("amount", "fifty dollars") — and normalization runs before scoring, so 500 and five hundred count as the same entity rather than two different failures.

The count itself is simple: every expected entity that never arrived (missing) plus every observed entity that was never expected (spurious), divided by the number of expected entities.

Two properties worth noticing before we go further:

  1. It's unbounded. A substituted amount counts twice — once for the amount you lost, once for the amount you invented. CEER can exceed 1.0, and does: the framework's own demo output reports CEER 2.0, which the docs call expected behavior for an unbounded error count. That's not a bug. It's the framework refusing to pretend that acting on the wrong amount is the same as acting on no amount.

  2. It's null when there's nothing to govern. Trials with no annotated entities don't get a CEER of zero — they get no CEER at all. Zero would mean "no corruption," which is a claim. Null means "not measured," which is the truth.

Three vignettes

Hand-built illustrative inputs — not measurements. But I ran them through the actual reference scorer instead of doing the arithmetic by hand, because the arithmetic is the easy part to get wrong. (Vignette 3 is the proof.)

Vignette 1: the amount. Reference: "Transfer fifty dollars to savings." ASR: "Transfer fifteen dollars to savings." One substitution in five words — WER 0.20, the kind of number that sails through a regression gate without a second look. The entity inventory tells the other story: one expected amount, one observed amount, and they disagree. Missing one, spurious one: CEER 2.0. WER says one word in five was wrong. CEER says the entire governed payload of the utterance was destroyed and replaced with a fiction. Your billing-disputes team would recognize the second description.

Vignette 2: the date. "Pay my bill on the fifth" → "on the ninth." WER 0.17 — smaller still, safer-looking still. CEER 2.0 again. The payment lands four days late, the late fee posts, a human agent spends six minutes unwinding it. WER 0.17 will never explain that ticket. CEER 2.0 at least points at the right word.

Vignette 3: the negation — and the inversion. "Please do not cancel my card" → "Please cancel my card." Two words dropped: WER 0.33, the highest WER of the three. But the entity inventory here has two entries — negation and action — and only one was corrupted. CEER 0.5, the lowest of the three.

Read that again, because it's the whole argument: the trial with the worst WER has the best CEER, and the trials with the best WER have the worst CEER. Across all three, mean WER is 0.23 and CEER is 1.25 — the two numbers aren't even pointing in the same direction. CEER is not WER with the volume turned up. It's a different unit measuring a different thing: corruption of governed information, not corruption of words. Sort your ASR error backlog by WER and you fix vignette 3 first. Sort by CEER and you fix vignettes 1 and 2 first — and your compliance team agrees with the second ordering.

What's specified, what's yours

The framework specifies the mechanism: the entity columns (expected_entities_json, predicted_entities_json), the normalization, the missing-plus-spurious count, the null-not-zero rule. It deliberately does not specify the inventory. What counts as critical — which entity types, which values — is a governance decision your team makes, because "critical" depends on what your system is allowed to do with the words. A music IVR and a payments IVR should not share an entity inventory, and the framework doesn't ask them to.

That also means CEER is only as honest as the annotations. Unannotated trials don't silently pass; they report null, and the summary aggregates only annotated trials. If your CEER coverage is 12% of the suite, the number means "of the 12% we bothered to govern." The framework won't let you forget the other 88%.

How it composes with attribution

Part 1's four-state table answers whether the speech path broke the task. CEER answers what governed information it corrupted along the way. All three vignettes above came back SPEECH_ATTRIBUTABLE — text control passed, audio path failed. But the CEER column splits that single verdict into two different tickets: vignettes 1 and 2 are entity failures (the recognizer corrupted governed tokens — an acoustic/model problem for the ASR backlog), while a SPEECH_ATTRIBUTABLE trial with CEER 0 would be a lexical failure the downstream survived everywhere except the task oracle (a robustness problem for the NLU backlog). Same attribution, different owner, different fix. That's why you want both numbers.

The honest boundary

Stated plainly: the 360-trial pilot shipped with the framework carries no critical-entity annotations, so its reproduced summary reports ceer: null. CEER is specified, implemented, and covered by the test suite — but it is not empirically exercised by that dataset. The vignettes above are illustrative inputs run through the reference scorer, not measurements. The first real CEER numbers will come from annotated suites, which don't exist in the repo yet. (A worked CEER walkthrough is on the project's issue list.)

The operational takeaway

Gate releases on entity corruption, not word corruption. The compare gate already supports it:

cafa-ivr compare \
  --baseline baseline/summary.csv \
  --candidate candidate/summary.csv \
  --max-ceer-delta 0.01
Enter fullscreen mode Exit fullscreen mode

A candidate that moves WER by a point but moves CEER not at all is a lexical reshuffle — ship it. A candidate that holds WER flat while CEER ticks up has learned to corrupt the words that matter while looking better on the dashboard. That one doesn't ship, and WER would never have told you.

The question to take back to your team: of the ASR errors in your last release, how many touched a governed entity — and would you even know? If the answer is no, you're gating on WER while your entities go unguarded.

Next in the series: the ASR-IFR deep-dive — from attribution verdicts to the release-gate metric, and why the denominator matters more than the numerator.

Top comments (0)