A few weeks ago I wrote about building a 7M-parameter typed dispatcher in pure machin/MFL — a tiny model that reads a request, picks a route, reports calibrated confidence, and delegates instead of guessing.
Since then I spent a sprint upgrading every decision head in the stack to multi-tap features. The biggest lesson wasn't the upgrade: the head I never re-benchmarked was the worst one.
The one-line version
- The core routing head was still the original single-tap classifier. Everything else had been improved around it.
- Re-benchmarking it honestly: 91.9% live holdout. Swapping in a multi-tap head trained on the same data: 99.5%.
- That's +7.6 points on the front door of the product, from a 55KB file.
What "multi-tap head" means
A linear head reads the frozen trunk's hidden state. The original format pooled one tap — one layer, one pooling function. The new mhd3 format concatenates N pooled taps in a single forward pass:
head = LR([ mean@-3 || max@-4 ]) // 576-dim feature, one prefill
All taps accumulate in-flight during the batched GEMM prefill — no extra forward passes, no framework, still a single static binary. The head format is ~30 bytes of header + weights.
The numbers that mattered (all live-verified)
| head | before | after |
|---|---|---|
| decide (16-route core) | 91.9% mhd1 | 99.5% mhd3 (796/800, 0 confident-wrong) |
| gate (domain dispatch) | 95.4% | 100% (149/149) |
| noul (OOD/abstain) | ~97% | 100% (150/150) |
| score (complexity tier) | ~98% | 100% (150/150) |
| municipal expert | 79.0% | 82.5% → 92.1% on merged taxonomy |
Two honest caveats, because this post is also a benchmark discipline post:
-
Every offline number was wrong until verified live. Twice an offline sweep produced a head that scored 99%+ on cached features and broke in production: a final-layer tap that doesn't map to the runtime's, and a
last-position tap that isn't expressible in the head format. Rule now: no emitted head is trusted until a live eval matches. -
The municipal "92.1%" is a taxonomy finding, not a model win. Over half the 15-class head's residual errors were one ambiguous route pair —
collecte_omvscollecte_selective— that citizens literally can't distinguish in phrasing. Merging them is a business decision, and it turns out that's worth more than any classifier trick we tried (ensembling, label-noise filtering, sub-heads: all null).
The performance side
Pooled heads initially forced every request through a sequential per-token prefill (~130ms). Two fixes:
- Taps now accumulate in-flight inside the batched prefill — the pooled path rides the same GEMM pass as generation.
- Prefix-cache reuse got clamped at the pooled-span boundary (reusing KV inside the span silently corrupted features — that's the bug that took 99.5% to 68.5% and almost shipped).
Net: ~4× faster than the first mhd3 build, same accuracy, still a single static binary + 8MB int8 model.
The actual lesson
When you stack improvements on a system, the component you stopped measuring rots quietly. The decide head worked, demos looked fine, and it sat at 91.9% while everything around it got better — because "it works" and "it's still the best available" are not the same question.
Sweep your oldest artifact against your newest technique, on live traffic, before assuming it's fine.
Links
- Router repo: github.com/javimosch/mtlm-router
- Runtime: machin-anvil (pure MFL inference server)
- Weights: Hugging Face
- Live demo + honest eval table: mtlm-router.intrane.fr
- Previous post: A 7M-parameter model that knows when to say no
Top comments (0)