DEV Community

Nicolai Bohn
Nicolai Bohn

Posted on

What are Jev evals? How they work and how they compare to LLM-as-a-judge

Evaluating agents with Jev has become one of the most discussed ideas in AI testing over the past few weeks. The appeal is easy to see: a verdict in milliseconds, for a fraction of what an LLM judge costs. Rhesis now supports Jev as a model provider for evaluation, so you can put it behind your categorical metrics today.

This post explains what Jev is and how it compares to an LLM judge. There's a demo video at the end.

What Jev is

Jev is a "System One" model from TypeSafe AI, released in September 2026. The name borrows from Kahneman's fast, intuitive System 1. Jev makes decisions and never generates text. You get no prose out of it, only typed values with probabilities attached.

A call has two parts. The state is whatever you want judged, as plain text or JSON. Alongside it you send a set of typed questions, and Jev answers all of them together in one parallel pass.

Take a support ticket that reads "I was charged twice for order A-104. Please refund the duplicate." You could ask Jev three questions about it at once:

  • Choice, department: billing, technical or sales? Jev picks one of up to 255 options and returns a probability for each.
  • Score, frustration: where does the customer sit on a scale from calm to very angry? A scale can have between two and ten levels.
  • Noul, refund_requested: is a refund being asked for? This is a yes/no question, answered as a probability.

The System One concept page in TypeSafe's docs explains these three primitives in more detail.

Pricing is $0.042 per million input tokens, and output is free. According to TypeSafe, a call takes 70 to 500 ms end to end. For a hands-on walkthrough with code, the practical guide from Valyu on DEV is the best place to start.

A Jev call: a support ticket as state, three typed questions answered in one parallel pass, and the typed answers with their probabilities

How it differs from an LLM judge

An LLM judge reads a rubric written in plain language. It returns a score or a category, usually with written reasoning that explains the verdict. That flexibility is why it works for any kind of metric. It also has a cost. Each call takes seconds, and you pay for every output token, reasoning included.

Jev takes a question with fixed answer options. It returns one of those options with a probability for each, and nothing else. Since there is no explanation to fall back on, the option names have to carry the meaning on their own. The price difference is largest where inputs are long. Output is free, so a long conversation costs you only its input tokens. The MindStudio post on Jev pricing compares this with using a general-purpose LLM for the same kind of decision, and shows the gap growing as inputs get longer.

Jev has real limits, and you should know them before you rely on it. It reads negations literally, so "the agent did not avoid the topic" lands at face value. It can't count or do math. It gives you no written rationale, which makes a surprising verdict harder to debug. The failure modes section of the Valyu guide covers these with examples. On the plus side, its probabilities are calibrated, so a high probability really does mean a more reliable answer. DataCamp's explainer describes how TypeSafe trains for that.

My rule of thumb: use Jev when the verdict is one of a few known outcomes, and use an LLM judge when you need a numeric score or an explanation.

Jev vs. an LLM judge: input, output, speed, price, what to watch out for and where each works in Rhesis

Jev in Rhesis

Jev is a new model provider in Rhesis. It does not add a new metric type. It is evaluation-only and works only with categorical metrics, the ones built on CategoricalJudge. Rhesis won't let you pick it for test generation or test runs, and numeric metrics refuse it as well.

When a categorical metric runs on Jev, Rhesis turns it into a single Choice question. The test input and the agent's reply become the state. The metric's prompt and evaluation steps become the instructions. Its categories become the answer options. Jev picks one category and returns probabilities for all of them, and Rhesis maps that pick to pass or fail based on the metric's passing categories. The result shows the chosen category and its probability. There is no reasoning text.

Two habits make Jev metrics work better. Name each category so it describes its outcome without context, for example recommended_partner_airline rather than ok. And phrase instructions positively, so the meaning doesn't hinge on a "not".

Setup takes a minute:

  1. Open the Models page and choose the Jev tile.
  2. Paste your TypeSafe API key.
  3. Click Test Connection to confirm Rhesis can reach Jev.
  4. Pick Jev as the model for a categorical metric, or make it your default model for evaluation.

The decision models section of the Rhesis docs covers the details, and the metrics docs explain how categorical metrics are set up. If you want to see how the integration is built, it landed in PR #2840.

How Rhesis maps a categorical metric onto a Jev call, and the three setup steps

See it in action

In the video, I test a travel agent that should only recommend partner airlines. For every reply, Rhesis asks Jev one Choice question, and Jev decides whether the agent kept to that rule. You'll see the categorical metric, the replies being judged one by one, and the category Jev picked for each test along with its probabilities.

Rhesis is open source. You can try it on app.rhesis.ai or self-host it from GitHub.


Originally published on the Rhesis blog.

Have you tried a decision model like Jev for your evals yet, or are you sticking with LLM judges? I'd like to hear in the comments.

Top comments (0)