The newest local-LLM wave isn't another chat model — it's models that never write prose at all. "Decision models" (TypeSafe's Jev API shape, now served locally by both Ollama and llama.cpp) take a state plus typed questions and return a probability for every option in one forward pass. No parser, no phrasing drift, no hallucinated reasons — a label and a number.
The pitch: half of what we point big chat models at is secretly a classification job. Ticket triage. Safety gates. Request routing. The public accuracy table is humbling at the small end (an 0.8B scores 63.5%, a 9B scores 75.7% against a hosted model's 76.0%), but the latency and cost side of the ledger is barely contested.
So the question, with receipts please:
Where in YOUR stack would a one-pass yes/no model replace a chat call tomorrow — and where have you already tried it and watched it fail?
Bonus ledger: did the confidence number survive contact with your real labels?
(Context: I wrote up the wave itself today — what shipped, which checkpoints, which licenses — over on the site. This thread is for the war stories, not the spec sheet.)
Top comments (0)