DEV Community

Cover image for Our task categorizer routed 46% of tasks. A $0.04/MTok decision model routes 97%.
Marcell Havlik
Marcell Havlik

Posted on Originally published at todofor.ai

Our task categorizer routed 46% of tasks. A $0.04/MTok decision model routes 97%.

Cosine similarity between a task title and a group description is not a classifier.

Bench, cases, raw verdicts: backend/bench/categorization in the TODOforAI repo. Total
cost ~$0.50.

Setup

TODOforAI boards have groups (SEO, Paid, Frontend…); new tasks get a suggested
group. V1: embed title and group name + description, nearest cosine wins. Embedder
Qwen/Qwen3-Embedding-4B via DeepInfra, 512 dims, unit-normalized, query instruction
"Given a task, retrieve the workstream whose purpose best matches the task". Abstain if
best < 0.50 or margin to runner-up < 0.06.

100% precision on 18 hand cases. Most of my real board in Unsorted.

Judges

Same 138 titles, same 8 groups, same description text, no example tasks — exactly what
production sees. Titles written like real todos: short, half Hungarian, 15 nonsense
(call mom, asdf), 5 ambiguous.

  • embedding — production, unchanged
  • jevTypeSafe Jev, non-autoregressive decision model on Vercel AI Gateway. One choice question, 8 groups + none. $0.04/MTok in, output free.
  • sonnet — Claude Sonnet 4.6, temp 0, one slug or none
  • opus — Claude Opus 5, same prompt. The truth. My labels are one more opinion.

Trap: Opus 5 reasons before answering; max_tokens: 20 looked like abstaining on 88/138.
Give it 400.

Numbers

Judge agrees with Opus wrong group missed 138 tasks
embedding 77/138 1 60/112 8 s
jev 130/138 5 3/112 24 s
sonnet 4.6 131/138 7 0/112 57 s

Opus routed 112, called none on 26 (15 nonsense + 11 vague). Embedding: wrong once,
missed 54%. Jev: 97% routed, 5 wrong. Sonnet's 7 wrong are coin flips with Opus.
On my real board Jev and Opus agree 21/26.

Why the embedding abstains

Not the margin:

                         agree   wrong   missed
margin 0.06  floor 0.50    76      2       60     ← production
margin 0     floor 0.50    79     41       18
margin 0     floor 0       76     62        0
Enter fullscreen mode Exit fullscreen mode

Drop the gate and Unsorted becomes wrong group. Top-1 is wrong on ~45% of tasks.

bun test flaky on CI          frontend 0.540   development 0.537
call mom                      email 0.540      plg 0.525
asdf                          development 0.663  frontend 0.659
Enter fullscreen mode Exit fullscreen mode

Scores sit in 0.42–0.72; asdf outscores a real task, so no floor finds nonsense.
Frontend/Development, Paid/SEO, Enterprise/Paid overlap, so real tasks sit inside any
margin. Jev and the LLMs know outcomes: a CI flake is Development because of what fixing
it achieves. Similarity can't say that.

What shipped

Jev routes; embedding is fallback when the gateway is unset or down. Batches of 25,
none explicit, accept probability ≥ 0.5. ~0.3 s per task.

Not Sonnet: 57 s vs 24 s and ~100× the price, for a suggestion 400 ms after you stop typing.

Router per 1,000 tasks per 1M tasks
embedding (~15 tok, $0.01/MTok) $0.0002 $0.15
jev (~260 tok catalog per question, $0.04/MTok in, out free) $0.01 $10
sonnet 4.6 (~350 in / 10 out) $1.20 $1,200
opus 5 (~350 in / ~200 out with reasoning) $7 $7,000

List prices, 8 groups. Jev costs ~70× the embedding and routes twice as much.

Jev's 5 wrong are channel-vs-outcome (sponsor a newsletter, measure signups → email,
not paid). Confidence is lower when wrong (0.70 vs 0.94), but floor 0.6 trades 2 wrong
for 7 Unsorted. We keep the 5.

Takeaways

  • Precision alone lies. Print recall against something you trust.
  • Score against a strong model, not your own labels.
  • Check the reference's output budget. Truncated looks cautious.

138 cases is small. It's enough to see a 54% recall gap; that's the only claim.

Top comments (0)