DEV Community

Javier Leandro Arancibia
Javier Leandro Arancibia

Posted on

Your oldest head is probably your worst head (a 7M router post-mortem)

A few weeks ago I wrote about building a 7M-parameter typed dispatcher in pure machin/MFL — a tiny model that reads a request, picks a route, reports calibrated confidence, and delegates instead of guessing.

Since then I spent a sprint upgrading every decision head in the stack to multi-tap features. The biggest lesson wasn't the upgrade: the head I never re-benchmarked was the worst one.

The one-line version

  • The core routing head was still the original single-tap classifier. Everything else had been improved around it.
  • Re-benchmarking it honestly: 91.9% live holdout. Swapping in a multi-tap head trained on the same data: 99.5%.
  • That's +7.6 points on the front door of the product, from a 55KB file.

What "multi-tap head" means

A linear head reads the frozen trunk's hidden state. The original format pooled one tap — one layer, one pooling function. The new mhd3 format concatenates N pooled taps in a single forward pass:

head = LR([ mean@-3 || max@-4 ])   // 576-dim feature, one prefill
Enter fullscreen mode Exit fullscreen mode

All taps accumulate in-flight during the batched GEMM prefill — no extra forward passes, no framework, still a single static binary. The head format is ~30 bytes of header + weights.

The numbers that mattered (all live-verified)

head before after
decide (16-route core) 91.9% mhd1 99.5% mhd3 (796/800, 0 confident-wrong)
gate (domain dispatch) 95.4% 100% (149/149)
noul (OOD/abstain) ~97% 100% (150/150)
score (complexity tier) ~98% 100% (150/150)
municipal expert 79.0% 82.5% → 92.1% on merged taxonomy

Two honest caveats, because this post is also a benchmark discipline post:

  1. Every offline number was wrong until verified live. Twice an offline sweep produced a head that scored 99%+ on cached features and broke in production: a final-layer tap that doesn't map to the runtime's, and a last-position tap that isn't expressible in the head format. Rule now: no emitted head is trusted until a live eval matches.
  2. The municipal "92.1%" is a taxonomy finding, not a model win. Over half the 15-class head's residual errors were one ambiguous route pair — collecte_om vs collecte_selective — that citizens literally can't distinguish in phrasing. Merging them is a business decision, and it turns out that's worth more than any classifier trick we tried (ensembling, label-noise filtering, sub-heads: all null).

The performance side

Pooled heads initially forced every request through a sequential per-token prefill (~130ms). Two fixes:

  • Taps now accumulate in-flight inside the batched prefill — the pooled path rides the same GEMM pass as generation.
  • Prefix-cache reuse got clamped at the pooled-span boundary (reusing KV inside the span silently corrupted features — that's the bug that took 99.5% to 68.5% and almost shipped).

Net: ~4× faster than the first mhd3 build, same accuracy, still a single static binary + 8MB int8 model.

The actual lesson

When you stack improvements on a system, the component you stopped measuring rots quietly. The decide head worked, demos looked fine, and it sat at 91.9% while everything around it got better — because "it works" and "it's still the best available" are not the same question.

Sweep your oldest artifact against your newest technique, on live traffic, before assuming it's fine.

Links

Top comments (0)