DEV Community

jamilxt
jamilxt

Posted on

TypeSafe's Jev Costs 400x Less Than an LLM. These 25 Lines of Python Do the Same Job for Free.

Your agent loop has a dirty secret: most of what it does all day is make tiny decisions. Is this input spam? Is this tool call dangerous? Should this ticket go to billing or support? And for every one of those coin-flip decisions, your harness fires up a full frontier LLM, generates a paragraph of reasoning, and charges you for the privilege.

A startup called TypeSafe AI built a whole product around that observation. Their model, Jev, skips text generation entirely and returns typed decisions with probabilities attached. The company reports up to 200x faster inference and 400x lower cost than comparable LLMs on classification tasks. The launch hit the front page of Hacker News, LangChain shipped an integration within days, and the usual hype cycle spun up.

Then a group called NobodyWho published a blog post with a title designed to start fights: "Jev in 25 lines of Python." It rebuilds the core trick with a 0.6-billion-parameter Qwen model running locally through llama.cpp. That post also went viral, sitting near the top of Hacker News with over 450 points, and the comment section turned into a genuine argument about what Jev actually is.

I have not called the Jev API myself. Everything below comes from TypeSafe's announcements, the LangChain writeup, the parody post, and two independent critiques that stress-tested the calibration claims. But the comparison between the two approaches is the most useful 30 minutes you can spend if you build anything with agents, because it exposes a decision you will have to make sooner or later: pay for fast decisions, or build your own.

What Jev actually is

Traditional LLMs are text generators. You prompt, they generate tokens until they decide to stop. Jev is what TypeSafe calls a System One model. It does not generate text at all. You hand it a state, you define the set of possible answers in advance, and it returns each answer with a probability attached. Their framing: a frontier-intelligence function call. Unstructured state in, typed probabilistic decisions out.

The training story matters to their pitch. TypeSafe says Jev was trained with reinforcement learning for calibrated decisions (RLCD), meaning the training signal rewards not just picking the right answer but attaching honest probabilities to it. That is the "calibrated" part of the marketing: if Jev says 0.7, they claim you can treat that as a real 70 percent chance rather than a vibes-based score.

Why would anyone want this? Three use cases keep coming up:

  • Agent safety gates. Coding harnesses like Claude Code and Cursor classify tool calls as dangerous or safe before executing them. That classifier step has historically been locked inside the closed-source parts of the harness. LangChain's AutoModeMiddleware uses Jev for exactly this, checking every tool call before it runs.
  • Routing. Read an incoming request, decide which agent or model should handle it, without burning a full LLM call on the decision.
  • Bulk classification. Sentiment, category, or risk scoring across thousands of rows, where a frontier model's cost and latency are absurd for the job at hand.

On paper it is compelling. The question the internet asked within hours was whether there is actually a there there.

The 25-line counterargument

The NobodyWho post is labeled a parody, but the code is real, it runs, and it demonstrates something important. Here is their core example, from their public blog post:

import numpy
from llama_cpp import Llama

model = Llama.from_pretrained(
    repo_id="Qwen/Qwen3-0.6B-GGUF",
    filename="Qwen3-0.6B-Q8_0.gguf",
    n_ctx=512,
    logits_all=True,
    verbose=False,
)

labels = ["A", "B", "C"]
choices = ["Legitimate", "Spam", "Phishing"]
email = "Payroll asks for your password on a non-company sign-in page."
options = "\n".join(
    f"{label}. {choice}" for label, choice in zip(labels, choices, strict=True)
)
prompt = f"""<|im_start|>system
Choose one option.<|im_end|>
<|im_start|>user
Email: {email}

{options}<|im_end|>
<|im_start|>assistant
<think>

</think>

"""
model.eval(tokens=model.tokenize(text=prompt.encode(), add_bos=False, special=True))

logits = model.scores[model.n_tokens - 1]
token_ids = [model.tokenize(text=label.encode(), add_bos=False)[0] for label in labels]
choice_logits = numpy.asarray([logits[token_id] for token_id in token_ids])
logprobs = choice_logits - numpy.logaddexp.reduce(choice_logits)
probabilities = numpy.exp(logprobs)
Enter fullscreen mode Exit fullscreen mode

For their sample phishing email, the output was: Legitimate 0.031, Spam 0.084, Phishing 0.885. A correct answer with a confidence score, from a model small enough to run on a laptop, with no API call and no data leaving the machine.

What the post is really arguing is an architecture claim, not a product claim. The trick Jev performs, reading probabilities directly off the logits for a fixed set of choice tokens instead of generating text, is a well-known technique. Constrain the output space, read the distribution, skip the generation. You can do it with any instruct model that exposes logits, which is most of them.

The parody authors are upfront about what they did not replicate: no System One branding, no RLCD training run, no calibration fine-tuning. Their point is narrower and sharper. The primitive, fast local classification with probabilities, is not proprietary. Whether you need TypeSafe's trained version depends on how much you trust its numbers and how much value you place on not running your own inference.

There are also serious open implementations now. OpenJev maintains an open reimplementation, there is an sglang-based variant, and DataCamp's roundup of alternatives counts at least seven open-source Jev-style options. The ecosystem voted with its feet within two weeks.

The calibration problem nobody marketed

Here is where it gets uncomfortable, and where the two most useful critiques landed.

Alex Molas, in a post titled "Jev can't be calibrated," makes a statistical argument that deserves to be quoted in spirit: calibration is not just a property of a model. It is a property of a model on a data distribution. A model is calibrated when, across all cases where it predicts 70 percent, roughly 70 percent of them turn out true. But your production data distribution is not TypeSafe's training distribution. Two companies can define spam identically and still see very different traffic. Jev will return the same probability to both for the same input, and it cannot be calibrated for both unless their distributions happen to match. Even if the RLCD training worked perfectly, the calibration guarantee does not transfer to you.

He also found sharper evidence. Prompted about a fair coin, Jev reportedly said heads comes up with probability 0.92. That is not a distribution shift problem. A fair coin has one right answer, and 0.92 is nowhere near it.

Arcturus Labs ran the systematic version of that test. They swept the stated bias of a coin from 0 to 100 percent in steps of a quarter point, 401 requests in all, and plotted Jev's stated probability against the true one. The result showed the outputs tracking the ideal line poorly across large stretches of the range. Their practical conclusion matches Molas': treat Jev's outputs as scores for ranking and thresholding, not as probabilities you can multiply against dollar amounts.

To be fair to TypeSafe, this critique applies to their marketing claim, not to the product's usefulness. Molas himself concedes the core value proposition: if you do not have labeled training data, Jev gives you a universal classifier on day one, where a fine-tuned BERT needs a dataset before it does anything. That is real. But it means the honest way to use Jev, or any of its open clones, is with your own validation set and a reliability curve, not with the training-on-paper calibration baked into your risk math.

Choosing between them: a decision matrix

So which one do you reach for? Both routes share the same architecture: constrained decoding over a fixed choice set, logits instead of generated text. The differences are about operations and trust.

  • No training data and you need answers today. Jev's strongest case. Zero-shot universal classification with structured output is exactly what it was built for. The open implementations work here too, just with less polish.
  • Data cannot leave your perimeter. The local script, no contest. The whole NobodyWho point was "we like not sending your data anywhere else." If you are classifying internal tickets or security logs, a 0.6B GGUF model on existing hardware removes the entire compliance conversation.
  • Scale where every millisecond and fraction of a cent compounds. This is Jev's reported 200x speed and 400x cost claim, and assuming those numbers hold on your workload, a managed endpoint with dedicated serving wins over self-hosting for most teams. Verify on your own traffic before believing the multiple.
  • You need honest probabilities for risk decisions. Neither, as-is. Build a small labeled set from your own data, measure calibration yourself, and pick whichever model's reliability curve is acceptable. A score you have validated beats a probability you were promised.
  • You already run an inference stack. The 25-line pattern is 20 minutes of work on top of llama.cpp or any logits-exposing endpoint. Adding a managed dependency for a trick you can own outright rarely pays.

The general lesson is bigger than this particular spat. When a well-known technique gets repackaged behind an API with a new name, the parody post is not just mockery, it is a price check on the abstraction. Sometimes the product earns its margin with training, serving quality, and ergonomics. Sometimes it is a 25-line script wearing a suit. Both of those things were true here, in different proportions, and the HN argument was people disagreeing about the ratio.

What I would actually do

If I were wiring a decision gate into an agent pipeline this week, I would start with the local script on real traffic, log the probabilities, and build a labeled set from whatever the pipeline got right and wrong. That costs an afternoon and produces the one artifact that matters: your own calibration curve. If a 0.6B model's curve is good enough, done, and no invoice ever arrives. If it is not, then TypeSafe's version, or one of the larger open reimplementations, has to beat the curve you measured, not the one in the launch blog post.

That order, measure first, buy second, inverts the usual demo-driven adoption path, and it works because classification is one of the rare ML problems where ground truth is cheap to collect after the fact.

I write about AI infrastructure, agents, and backend engineering every week. Subscribe, it is free.

Do you run classification inside an agent loop today? Are you paying an API for it or rolling your own logits-reading trick? Tell me what your thresholds look like in the comments.

Top comments (0)