DEV Community

Rudratosh Shastri
Rudratosh Shastri

Posted on

Someone trained a decision model for $104. The price isn't the interesting part — it's that it refuses to write text.

Quick question about your pipeline: how many of your "LLM calls" actually need the model to write anything?

Be honest. A lot of them are secretly classification wearing a chat costume:

  • "Is this ticket about billing, bug, or feature?"
  • "Should this action be allowed or escalated?"
  • "Is this review fake?"
  • "Which of these 4 routes fits?"

You're paying a generative model to emit tokens, then parsing those tokens back into a decision — and hoping it didn't decide to be chatty, or wrap the JSON in an apology, or hallucinate a fifth category.

A new class of model is quietly calling that out. Decision models — Jev (TypeSafe), its open clone Laya, and one reportedly trained for $104 at 144M parameters — do the opposite of a chatbot: they refuse to generate text. They take a structured question and return a calibrated probability. That's it.

Why "won't write text" is a feature, not a limitation

When a model's only output is a number over a fixed set of choices, three problems just… disappear:

  • Nothing to parse. No "sometimes it returns markdown, sometimes a code fence, sometimes prose." The output is a distribution, always the same shape.
  • Nothing to hallucinate. It can't invent a new category or narrate a wrong reason, because it isn't producing free text at all.
  • Calibrated confidence. A good decision model tells you how sure it is — 0.51 vs 0.99 — which is exactly what you need to route the uncertain cases to a human. An LLM's "I'm confident!" is vibes; a calibrated probability is a number you can threshold.

And the economics are absurd in your favor. A 144M-parameter model runs in ~tens of milliseconds on modest hardware, at a fraction of a cent. When the reported training cost is two figures, the inference cost rounds to zero. Compare that to a frontier LLM doing the same yes/no at $10–50 per million tokens plus latency.

Where they fit (and where they don't)

Use a decision model when the task is a decision:

  • classification and routing
  • filtering / moderation gates
  • "should the agent take this action?" checks at the tool call
  • ranking and relevance scoring
  • anywhere you're currently regex-parsing an LLM's answer back into an enum

Keep the LLM when you actually need language:

  • generating the reply, the summary, the code
  • open-ended reasoning where the output space isn't fixed
  • anything where "write me a paragraph" is the literal job

The pattern that's emerging in serious pipelines: a cheap decision model out front to route, filter, and gate; the expensive generative model only on the calls that truly need words. You stop paying frontier prices for yes/no, and you stop parsing prose into booleans.

The honest caveats

  • Structured input is a real cost. These models want a defined question schema, not a freeform prompt. Retrofitting a pipeline built around "just prompt it" takes work.
  • Open vs closed matters. Jev is closed; Laya is the open alternative; the ecosystem is young and the tooling is rougher than the LLM world you're used to.
  • They don't reason out loud. If you need a rationale you can read, a silent probability isn't it — though "return the number, ask the LLM for the why only when the human wants it" is a nice hybrid.

The mental shift is the takeaway: not every decision your system makes needs a paragraph and a parser. Some of them just need a well-calibrated number — and there's finally a cheap, un-hallucinating tool built to give you exactly that.


How many of your current LLM calls are secretly just classification you're parsing back out of prose? Count them honestly — I bet it's more than you think. 👇

I write about building with AI and picking the right tool for the job. Follow me here if that's your lane. 👋

Top comments (0)