DEV Community

thekami
thekami

Posted on

Meet Jev: The Fast, Cheap AI That Only Picks, Never Writes

TL;DR: Jev is a model from TypeSafe AI that cannot write text. You send it a state and a few typed questions, and it returns typed answers with probabilities in 70 to 500 ms. It is built for small, repeated decisions (classify, route, score, extract), not for chat. The numbers below are TypeSafe's own, so test on your data before you build on them.

Topic first shared on LinkedIn by Abubakar Siddiq Sazzad.

Most AI launches in the last few years have been about models that write better, reason longer, or sound more human. On September 15, 2026, TypeSafe AI released one that cannot write at all, and that is the whole idea.

The model is called Jev. It arrived after two years of stealth, together with a $40 million seed round led by DCVC, and it is the first model in a class TypeSafe calls System One. We spent a few days reading the launch material, the documentation, and the independent write-ups. This post is what we took from it: what Jev does, which claims look solid, and which ones need a second look before anyone builds on them.

At a Glance

Detail Info
Maker TypeSafe AI, San Francisco. Founded in 2024 by Diogo Almeida (formerly at OpenAI), Erik Gafni, and Sasha Sheng
Released September 15, 2026, in limited early access
Funding $40 million seed round led by DCVC
What it returns Typed answers with probabilities. No text
Reported speed 70 to 500 ms end to end
Reported price $0.042 per million input tokens, output free

Every speed, price, and accuracy figure in this post comes from TypeSafe or from people reporting TypeSafe's numbers. We have not run Jev on a workload of our own, and we will say so again where it matters.

The Problem It Goes After

Look inside an AI agent and you will find a loop. A model decides what to do, a tool runs, something checks the result, and the cycle repeats.

A surprising number of the steps in that loop are not open-ended thinking. They are questions with a short list of valid answers:

  • Is this customer message urgent?
  • Which of these two models should handle this request?
  • Is this shell command safe to run?
  • Did the scraper return a real page or a captcha?

Today, each of those goes to a large language model. The model writes a sentence, and then your code parses the sentence to extract the one word it needed. You pay for an essay to read a single word from it, and you wait seconds for the privilege.

Jev is built around that observation. If the valid answers are known before the call, the model does not need to write anything. It can just pick.

System One and System Two

The name comes from Daniel Kahneman's Thinking, Fast and Slow. System 1 is fast, intuitive judgment. System 2 is slow, deliberate reasoning. Chat models behave like System 2: they work through a problem one token at a time, each token depending on the last.

TypeSafe describes the differences this way:

Typical LLM Jev (System One)
Output Free text that software has to parse and validate Typed values whose shape is fixed before the call
Generation One token at a time All answers produced in parallel, in one pass
Confidence Often overconfident, even when asked to estimate A calibrated probability with every answer
Training goal Responses people prefer, or outputs a program can verify Answers whose stated probabilities are honest
Best at Conversation, writing, code, open-ended reasoning Classify, route, score, extract, branch

That last training goal is the interesting one. TypeSafe calls its method Reinforcement Learning for Calibrated Decisions. The target is not "give the answer a human likes" but "when you say 90%, be right about 90% of the time." A model that knows when it is unsure is more useful inside software than one that sounds equally confident about everything.

What a Call Looks Like

You send Jev a state, which is whatever context your software holds right now, and a set of questions about it. This is a simplified version of the example in TypeSafe's quickstart:

{
  "model": "jev-latest",
  "state": "I've been trying to connect my Stripe account for 3 days and it keeps failing. I'm losing sales. Please help ASAP.",
  "questions": {
    "is_urgent": {
      "type": "noul",
      "instructions": "The message conveys urgency or time-sensitivity"
    }
  }
}
Enter fullscreen mode Exit fullscreen mode

What comes back is not a sentence:

{
  "is_urgent": { "type": "noul", "noul": 0.999 }
}
Enter fullscreen mode Exit fullscreen mode

The real response wraps this in a few extra fields (an answers object, the model version, token usage), and the exact probability can shift a little between runs, so treat the number above as an illustration. The point is the shape. Your code gets a number it can threshold, log, and compare against last month, not the word "yes".

Three kinds of question

Everything you can ask Jev fits one of three types:

  • Choice picks from a list of options and returns a probability for each, plus an overall confidence.
  • Score rates something against ordered levels such as low, medium, and high.
  • Noul answers a yes-or-no question and returns the probability that it is true.

Why parallel questions matter

You can attach many questions to one state, and Jev evaluates all of them at once, each independently. Adding a question barely changes the response time. You only pay for the extra question tokens.

That sounds like a convenience, but it changes how you design things. When every check costs a full model call, you ration them and ask only the two questions you are sure you need. When extra checks are nearly free, you ask all seven and ignore the ones you did not need. The checks people skip are often the ones that would have caught the problem.

TypeSafe's documentation includes a worked example that puts numbers on this: thirteen compliance questions about a long document. Asked as one batched request, it cost roughly twelve times less and ran about ten times faster than thirteen separate calls, with the answers staying the same. The reason is plain. The document is the expensive part, and separate calls pay for it thirteen times.

What "Type-Safe" Does and Does Not Mean

Because the possible answers are defined in advance, Jev cannot return a value outside that set. No malformed JSON, no invented category, no field that does not exist. TypeSafe treats this as true by construction, not as something they measured. For anyone who has written retry logic around broken model output, that alone is worth attention.

But it is a narrower promise than it first sounds:

  • It guarantees the shape of the answer.
  • It does not guarantee the answer is right. Jev can pick the wrong option with high confidence.

TypeSafe is open about this distinction. Its broader claims about hallucination are flagged as not empirical. The probability attached to each answer is what lets you decide how far to trust it, which is why calibration matters so much to the design.

The Numbers, and Whose Numbers They Are

The headline claims are large. TypeSafe says Jev responds in 70 to 500 ms, against several seconds for frontier models on comparable decision tasks, and its homepage cites 193.6 times faster and 444.6 times cheaper on its own workflow comparisons.

TypeSafe published the evals behind those multipliers, and the details are more useful than the headline. As summarized by Width.ai, an agent-development consultancy, the setup works like this: each task is split into narrow typed questions, code makes the final decision from the answers, and accuracy means agreement with a reference built from two frontier models (one from OpenAI, one from Anthropic) answering the same questions at high reasoning effort. That is a reasonable proxy for correctness, but it is not the same thing as correctness.

Averaged over four workflows (customer service, agent-trace review, security incidents, and invoice processing):

Model Accuracy Cost per case Time per case
Sol (most accurate) 74.1% $0.0836 23.3 s
Opus 5 73.1% $0.1761 37.8 s
Sonnet 5 67.8% $0.1174 78.1 s
Jev 67.8% $0.0004 0.4 s

Jev is not the most accurate model on the board, and TypeSafe does not claim it is. What stands out is that it matches Sonnet 5's accuracy at a tiny fraction of the cost and latency. Averages also hide a lot. On customer service, Jev landed within about two points of the best model. On invoice processing, where the task means holding a bill, an order, and a delivery record against each other, it trailed by more than seventeen points.

The pattern is easy to read: Jev does well when the job is to read one situation and pick from known responses, and it struggles when the job is to reconcile several documents. That is useful guidance for choosing where to start.

There was one more finding in the same evals that has nothing to do with Jev. Every model tested did better when the decision was broken into small questions than when the whole policy was handed over as one big prompt. One smaller model jumped from about 18% to about 54% from decomposition alone. Splitting the decision up helped more than upgrading the model did, and you can do that today with whatever you already use.

Where Jev Fits

Nobody is suggesting Jev replaces language models. The realistic picture is two tiers: a System One model handling many small, repeated decisions, and an LLM handling reasoning and anything that has to be written in words. Based on TypeSafe's documentation and the LangChain integration that shipped two days after launch, the main uses look like this:

  • Model routing. Decide whether a request needs a powerful model or a cheap one before spending money on either. The routing call costs almost nothing, but it is only worth adding if enough of your traffic can actually move to the cheaper model.
  • Tool-call gating. Check a proposed action, such as a shell command, before the agent executes it. This is the pattern coding agents use to decide what is safe to run unattended. A fast, cheap classifier makes it practical for any team.
  • Guardrails. Screen messages going into and out of an LLM application, and decide in code whether to pass, review, or block.
  • Retrieval cleanup. Score retrieved passages, re-rank candidates, or check whether a cited source really supports a claim. These are checks most RAG systems would like to run on every request and rarely can.
  • Taxonomy classification. Walk a category tree with one Choice question per node, keeping several candidate paths alive so one early mistake does not doom the result.

For anything that sits in front of an irreversible action, the sensible structure is layered. A hard-coded deny list in ordinary code goes first. A probabilistic gate like Jev goes behind it. Human review catches the denials. A model that is right 98% of the time is still wrong 2% of the time, and that 2% adds up on something that runs on every action.

Limits Worth Knowing

  • It is in early access. The API and the integrations built on it can still change.
  • It does not write text. Drafting, summarizing, code, and explanations stay with a language model.
  • Its input is limited. Reports of TypeSafe's documentation put the budget for a state plus all its questions at about 64k tokens, and the model does not read images.
  • TypeSafe publishes its own list of weak spots. Counting and arithmetic belong in ordinary code. Dates are read as text, not as ordered quantities. The model reads instructions very literally. And accuracy drops when the state is full of details irrelevant to the question, so filter first and send only what the question needs.
  • The performance numbers are vendor numbers. TypeSafe's own team built the evals, and TypeSafe says openly that it cannot prove the low pricing is sustainable. Test on your own inputs before relying on any of it.
  • Some tutorials are not about Jev. A few circulating walkthroughs use the name for a generic structured-output pattern. A quick check helps: the real interface takes a state and typed questions, does not take an LLM model name, never returns text, and attaches a probability to every answer.

One more caution applies to any decision that fails quietly. A misrouted support ticket is noticed within minutes. A wrongly scored eligibility decision looks identical to a correct one until someone audits it. That kind of step should keep a strong model or a person on it, however neatly it fits the question types.

What We Take From This

We are a small team, and most teams in Bangladesh work under similar constraints: tight budgets, few engineers, and no appetite for a feature that adds a slow, expensive API call to every action. That is why the economics here interest us more than the model itself. When a small decision costs a fraction of a cent and returns in a fraction of a second, a lot of ideas that were not worth building become reasonable.

A few practical lessons seem worth keeping, even if Jev itself turns out not to be the model you use:

  1. Split decisions into small questions. The evals suggest this helps every model, not just this one.
  2. Log the probability, not only the decision. You will need the distribution later to set sensible thresholds, and it is painful to reconstruct.
  3. Keep the model name in configuration. If each decision point names its model in config rather than in code, trying a System One model, or the next one, is a small change.
  4. Start with a low-stakes, high-volume step. You learn how the model behaves on your data somewhere a mistake is cheap.

Whether Jev becomes the standard or just the first of many, the question it raises is the right one. For years the question was whether AI could generate better text. The more interesting one now is how much of the software around us can learn to decide quickly, cheaply, and predictably.

FAQ

Is Jev a large language model?
No. It does not generate text. It takes a state and typed questions and returns typed values with probabilities, and all answers are produced in parallel instead of token by token.

Can Jev hallucinate?
It cannot return a value outside the answers you defined. It can still choose the wrong one, sometimes with high confidence, so the probabilities and a fallback path still matter.

How much does it cost?
TypeSafe lists $0.042 per million input tokens, with output free. That figure is the company's own, and it has acknowledged that long-term pricing will need to prove itself.

Can anyone use it today?
Not fully. It launched in limited early access on September 15, 2026, and developers are being let in gradually.

Sources

  • TypeSafe AI's launch announcement and documentation, including its published evals and model-limitations page
  • Launch coverage by WOWTALE and TAO Media
  • Width.ai, "What Is Jev AI?" (September 29, 2026), for the eval breakdown and LangChain integration details
  • Explainers from Nexos, Maxim, and others on pricing and design

All performance, pricing, and accuracy figures are as stated by TypeSafe AI and may change as Jev leaves early access.


Originally published on the TheKami blog. We are TheKami, a remote-first software company from Bangladesh. Follow us on GitHub for our open-source work.

Top comments (0)