DEV Community

Mr Say Nothing
Mr Say Nothing

Posted on Originally published at mrsaynothing.dev

What decision models can't do: six honest limits

This one first ran on mrsaynothing.dev — What decision models can't do, the against-the-grain companion to yesterday's explainer.

Everyone is piping the new decision models into production because the release notes were exciting. Nobody has published an accuracy bench. That gap is this post.

The short version: use them for decisions, never for reasons. Two of the six are non-commercial. The printed limits — option counts, token ceilings, language gaps — are real and fail silently. And the confidence number is a claim until somebody benches it on real labels.

The six limits

  1. There are no reasons in the box. A decision model returns a label and a number. Why that label — the evidence, the counter-argument, the doubt — is not in the output, at any price. The moment your pipeline needs a why, you are back to a chat model or a human.
  2. Agreement is not verification. Ask three models and get three matching answers, and you have one opinion in triplicate. A unanimous wrong answer arrives in the confident tray — sorted, stamped, and wrong at 0.99. Consensus needs an independent reference to mean anything, and none of these models brings one.
  3. The license walls are real. Julia-1, Laya, lev and Kev-4B are Apache-2.0. Nimble and OpenJev are CC BY-NC 4.0 — non-commercial, full stop. The llama.cpp runtime is MIT, and that means nothing about the weights: the checkpoint card outranks the repo's license badge every time.
  4. The modality map has holes. Laya is English-only — don't route multilingual traffic through it. Nimble is text-only locally. OpenJev's vision input needs its own projector. Kev-4B ships without the reference-date preprocessing it expects. Each gap is printed; none of them is visible from the API's side.
  5. The limits are printed on the box. Julia-1 takes 2–20 options and an 8,192-token combined question — go past either and the failure is silent, not an error. Feeding a decision model a 12k-token document is not stretching it; it is using a different product than the one that shipped.
  6. Serving is not proof. A new endpoint returning numbers is a plumbing fact, not an accuracy fact. The published per-million-token prices and the serving support say nothing about whether the probability is honest on your labels. Only a bench on real traffic answers that — and no one has published one yet.

What holds despite all six

The scope gets smaller, not the value. One-pass routing, gating and triage — where the output is a decision and a human owns the consequences — is still the cleanest fit local AI has offered in a year. Small, fast, free of per-token bills, and honest about being a stamp. The failure mode to design for is the confident tray: set a threshold, escalate below it, and never let the stamp be the last word on anything that matters.

Remember this: a probability is a claim. Stamped ≠ sorted right. Bench on your labels before the tray is the truth.

The bench is next — Julia-1 against Laya on real routing labels, calibration per size, latency per hardware. If the confidence numbers hold, the wave earns it. If they don't, this post already said so.

Which of the six limits bites hardest in your stack — and has anyone seen a published calibration bench for any of these models yet?

Top comments (2)

Collapse
 
mrsaynothing profile image
Mr Say Nothing •

The license split surprised me most: 4 of 6 Apache-2.0, but Nimble and OpenJev are CC BY-NC — none of the release-notes excitement mentions that a commercial pipeline legally stops at those two. Which of the six limits did I underweight?

Collapse
 
deanlee profile image
Dean Lee •

The second limit, treating agreement across models as verification, is where the balance-sheet risk concentrates. In financial pricing, running three correlated models that share underlying training distributions does not buy you an ensemble; it gives you the exact same tail blindness at three times the inference cost. If all three checkpoints were pre-trained on the common web crawl, unanimous agreement is just shared covariance masquerading as confidence.

The other structural blind spot is calibration under asymmetric loss. A raw softmax score or confidence scalar treats a false positive and a false negative as having identical payoffs. In real production triage, the cost matrix is almost never symmetric. Letting a corrupted payload pass downstream can break an entire pipeline, while routing a borderline request to human review is a fixed queue cost. Until a decision model lets you condition the decision boundary on the operational loss function rather than raw likelihood, the confidence number is just an unhedged probability claim.