Quick question about your pipeline: how many of your "LLM calls" actually need the model to write anything?
Be honest. A lot of them are secretly classification wearing a chat costume:
- "Is this ticket about billing, bug, or feature?"
- "Should this action be allowed or escalated?"
- "Is this review fake?"
- "Which of these 4 routes fits?"
You're paying a generative model to emit tokens, then parsing those tokens back into a decision — and hoping it didn't decide to be chatty, or wrap the JSON in an apology, or hallucinate a fifth category.
A new class of model is quietly calling that out. Decision models — Jev (TypeSafe), its open clone Laya, and one reportedly trained for $104 at 144M parameters — do the opposite of a chatbot: they refuse to generate text. They take a structured question and return a calibrated probability. That's it.
Why "won't write text" is a feature, not a limitation
When a model's only output is a number over a fixed set of choices, three problems just… disappear:
- Nothing to parse. No "sometimes it returns markdown, sometimes a code fence, sometimes prose." The output is a distribution, always the same shape.
- Nothing to hallucinate. It can't invent a new category or narrate a wrong reason, because it isn't producing free text at all.
- Calibrated confidence. A good decision model tells you how sure it is — 0.51 vs 0.99 — which is exactly what you need to route the uncertain cases to a human. An LLM's "I'm confident!" is vibes; a calibrated probability is a number you can threshold.
And the economics are absurd in your favor. A 144M-parameter model runs in ~tens of milliseconds on modest hardware, at a fraction of a cent. When the reported training cost is two figures, the inference cost rounds to zero. Compare that to a frontier LLM doing the same yes/no at $10–50 per million tokens plus latency.
Where they fit (and where they don't)
Use a decision model when the task is a decision:
- classification and routing
- filtering / moderation gates
- "should the agent take this action?" checks at the tool call
- ranking and relevance scoring
- anywhere you're currently regex-parsing an LLM's answer back into an enum
Keep the LLM when you actually need language:
- generating the reply, the summary, the code
- open-ended reasoning where the output space isn't fixed
- anything where "write me a paragraph" is the literal job
The pattern that's emerging in serious pipelines: a cheap decision model out front to route, filter, and gate; the expensive generative model only on the calls that truly need words. You stop paying frontier prices for yes/no, and you stop parsing prose into booleans.
The honest caveats
- Structured input is a real cost. These models want a defined question schema, not a freeform prompt. Retrofitting a pipeline built around "just prompt it" takes work.
- Open vs closed matters. Jev is closed; Laya is the open alternative; the ecosystem is young and the tooling is rougher than the LLM world you're used to.
- They don't reason out loud. If you need a rationale you can read, a silent probability isn't it — though "return the number, ask the LLM for the why only when the human wants it" is a nice hybrid.
The mental shift is the takeaway: not every decision your system makes needs a paragraph and a parser. Some of them just need a well-calibrated number — and there's finally a cheap, un-hallucinating tool built to give you exactly that.
How many of your current LLM calls are secretly just classification you're parsing back out of prose? Count them honestly — I bet it's more than you think. 👇
I write about building with AI and picking the right tool for the job. Follow me here if that's your lane. 👋
Top comments (0)