DEV Community

Cover image for Jev is the best decision model. Here's what to run when you can't use it.
Mark Vange
Mark Vange

Posted on Originally published at Medium AI-assisted

Jev is the best decision model. Here's what to run when you can't use it.

Jev is the best decision model. Here's what to run when you can't use it.

Part 1 of a two-part series. Part 2 continues the story — link to be added once it's published.

We put seven open-source Jev look-alikes on a single 8 GB GPU for teams whose data isn't allowed to leave the building. One of them is good enough to build on.

Mark Vange · Autom8ly · September 2026 · 9 min read

Written mostly by AI. This article and the code behind it were largely generated by an AI assistant (Claude), directed by me in response to specific use cases we face in high-compliance environments. Every number was measured on real hardware, and the code is public so you can check it.

A few weeks ago TypeSafe released Jev, and it changed how I think about putting language models inside software.

Most of what we ask models to do in production isn't writing. It's deciding. Which team should get this ticket? Does this message need a human? Is this claim eligible under our policy? How urgent is this, on a scale of one to three? For years the standard answer was to prompt a general chat model, ask it to reply in JSON, parse whatever came back, and hope it didn't wander off script.

Jev skips the text entirely. You send it a piece of state and a set of typed questions (pick one, yes/no, or a score), and it sends back probabilities. No parsing, no retries, no invented options. In our tests it answered in about 135 milliseconds, over the internet.

The probabilities are the real prize. If a model tells you it is 97% sure a ticket belongs to billing, you can route that ticket automatically and send the 60% cases to a person. That pattern, automate the confident decisions and escalate the rest, is how you put AI into processes where mistakes are expensive. Routing, triage, guardrails, policy checks and agent stop conditions all become cheap, fast function calls. It opens up a lot of new work.

The catch: some of us can't call an API

I work with teams in high-compliance environments. For many of them, "just call a hosted API" isn't an option. Regulated data may not be allowed to leave a particular network, or a particular country. A new vendor can mean months of security and privacy review. Some systems run on networks with no route to the internet at all. And even where an external call is permitted, auditors want to know exactly which model made each decision, and that it won't change underneath you.

When you can't bring your data to the best model, you have to bring a good model to your data.

So we went looking for alternatives we could run on our own hardware. Within a week of Jev's release, at least seven open-source look-alikes had appeared. The question was whether any of them were good enough, and whether they would fit on the kind of modest GPU you can actually get approved: a single 8 GB consumer card.

Seven look-alikes in a week

  • Kev: Qwen3.5 models (0.8B, 4B, 9B) fine-tuned to read out option probabilities in one pass.
  • Von: a 395-million-parameter encoder built for speed.
  • Laya: small multilingual encoders behind a language router.
  • Rizzo Flow: 1.7B and 4B models served through llama.cpp, with a built-in playground.
  • SemIf: no training at all; it scores options straight from an ordinary model's next-token probabilities.
  • Nimble: a 9B model with an open data-curation and training recipe.
  • NanoJev: a small, fully trainable replica, released with a checkpoint trained on game states.

All of them copy the same basic idea, and four of them speak Jev's HTTP interface, which made a fair comparison possible.

How we tested

Everything ran on one NVIDIA GeForce RTX 4060 with 8,188 MiB of memory, inside a virtual machine with the GPU passed through. We ran each tool alone, one request at a time, and put every one behind the same /v1/systemone interface so a single scorer could grade them all.

For the questions we used an independent public suite: jabr/classifier-benchmark v2, 866 synthetic cases across 49 everyday decision tasks such as support routing, refund and warranty eligibility, phishing and fraud checks, triage, tone and grammar. Its cases were generated and cross-checked by a committee of seven different LLMs, and it is released into the public domain. None of the tools was built around it, and our runs reproduce its author's published scores for Jev, Von and Laya almost exactly.

Three of the tools only ship full-precision weights, which don't fit in 8 GB, so we added 4-bit and 8-bit loading to them with bitsandbytes. We also ran Jev itself through TypeSafe's API, and a plain generative model (qwen3:8b through Ollama, asked for a JSON answer) as the old-fashioned baseline. As a cross-check we ran every setup on a second suite, Kev's own transfer-v4; those results are in the repository and agree on who comes first.

What fits in 8 GB

Almost everything, once quantized. The only one that didn't fit was Nimble's 9B model, which loads at 7.5 GB before doing any work and runs out of memory on its first request.

Peak GPU memory per setup against the 8GB RTX 4060:

  • NanoJev 0.6B — 2.4 GB
  • Rizzo 1.7B · q8 — 2.5 GB
  • SemIf 4B · 4-bit — 3.2 GB
  • Von — 3.3 GB
  • Kev-4B · 4-bit — 3.8 GB
  • Rizzo 4B · q4 — 4.1 GB
  • Kev-0.8B — 4.5 GB
  • SemIf 4B · 8-bit — 5.0 GB
  • Kev-4B · 8-bit — 5.3 GB
  • Laya — 5.5 GB
  • Rizzo 4B · q8 — 5.7 GB
  • qwen3:8b via Ollama — 5.9 GB
  • Kev-9B · 4-bit — 7.4 GB
  • Nimble-9B · 4-bit — 7.5 GB idle, out of memory

Peak GPU memory while scoring, sampled every half second.

Quantization cost little. Kev-4B scored 0.878 at 8-bit and 0.872 at 4-bit, and Rizzo 0.807 and 0.799; only SemIf lost more, about three points. For a small GPU that trade is easy: Kev-4B needs more than 8 GB in full precision and runs in 3.8 GB at 4-bit.

Accuracy against speed

Results — accuracy on 866 questions against median time per question, including a local HTTP hop:

Setup Accuracy Wrong at ≥90% conf. Automatable @5% err p50 ms Peak VRAM
Jev 1.13 (hosted) 0.968 0.2% 100% 135 hosted
Kev-9B · 4-bit 0.893 0.7% 76% 223 7.4 GB
Kev-4B · 8-bit 0.878 0.0% 77% 385 5.3 GB
Kev-4B · 4-bit 0.872 0.1% 75% 188 3.8 GB
SemIf 4B · 8-bit 0.866 3.6% 74% 168 5.0 GB
SemIf 4B · 4-bit 0.837 5.0% 66% 130 3.2 GB
qwen3:8b via Ollama 0.827 n/a n/a 207 5.9 GB
Rizzo 4B · q8 0.807 12.4% 22% 71 5.7 GB
Rizzo 4B · q4 0.799 11.5% 3% 74 4.1 GB
Von 0.724 1.7% 21% 18 3.3 GB
Kev-0.8B 0.718 0.3% 14% 29 4.5 GB
Rizzo 1.7B · q8 0.639 22.5% 3% 35 2.5 GB
Laya 0.585 3.1% 0% 21 5.5 GB
NanoJev 0.6B 0.343 0.5% 0% 29 2.4 GB
Nimble-9B · 4-bit did not fit — — — 7.5 GB idle

Kev-4B at 4-bit is the best all-rounder. It scores 0.872, 9.6 points behind Jev's 0.968, in 3.8 GB at 188 ms per question. Its confidence is trustworthy: only 0.1% of its answers are wrong at 90% confidence or more (Jev: 0.2%), so with a 5% error budget you could automate about 75% of its decisions without review. The 8-bit build is a hair more accurate but twice as slow. Kev-9B scores a little higher (0.893), but it uses 7.4 GB of the 8 GB card, leaving no room for anything else.

SemIf is a close second with no training at all. Scoring options straight from an unmodified Qwen3.5-4B reaches 0.866 at 8-bit and 0.837 at 4-bit. That makes it the natural choice if you can't adopt a new model but can run one you already trust. Its confidence needs calibrating first: 3.6–5% of its answers are confidently wrong.

Von is the speed pick. It answers in 18 ms in 3.3 GB, ten times faster than Kev-4B, and scores 0.724. It is dependable on tone and routing but drops to coin-flip level on rule-based checks such as warranty eligibility, phishing and suspicious transactions.

The rest are harder to recommend today. The generative baseline scores a respectable 0.827 but gives no usable confidence: 17% of its answers are confidently wrong. Rizzo is fast (71 ms) and scores about 0.80, but 11–12% of its answers are confidently wrong until you calibrate it, and it struggles with policy rules. Laya (0.585) and Kev-0.8B (0.718) trail Von, and NanoJev's game-trained checkpoint scores near chance.

Jev is still the best, and a small experiment shows why

A 9.6-point gap is large, and it widens on the hardest tasks: on grammar checking Jev scores 0.81, while Kev-4B and SemIf score 0.48. But the gap is not just a score. Jev's real advantage shows up when you ask it something new.

While writing this article we tried exactly that. Ten different AI models each rewrote the article's opening paragraph to sound more natural, keeping every fact. We added two control paragraphs written as deliberately extreme styles, one stiff corporate prose and one very casual. Then we asked Kev-4B and Jev the same three questions about each paragraph: was it written by a human rather than an AI, how natural does it sound, and which of the ten sounds most human.

Rewrite by Kev: P(human) Jev: P(human) Jev: natural Jev: picked as most human
Fable 5.1 0.69 0.57 0.76 12%
Opus 5.5 0.68 0.58 0.75 39%
Sonnet 5 0.70 0.60 0.76 7%
Opus 4.6 0.70 0.56 0.75 5%
Codex (OpenAI) 0.73 0.56 0.72 3%
Haiku 4.5 0.70 0.56 0.70 2%
hermes3:8b (local) 0.71 0.45 0.67 30%
OpenCode big-pickle 0.69 0.48 0.56 2%
qwen3:8b (local) 0.69 0.47 0.43 1%
gemma4:e2b (local) 0.70 0.49 0.40 1%
Control: stiff corporate prose 0.66 0.25 0.14 —
Control: deliberately casual 0.68 0.59 0.84 —

Sorted by Jev's average rank across its three measures. "Natural" is a 3-level score scaled 0–1. "Picked" is the share of ten rotated pick-one questions; chance is 10%.

Kev could not tell. It scored the stiff and casual controls within 0.014 of each other, and its three measures disagreed on the winner. Jev separated the controls sharply (0.25 against 0.59 on the human question, 0.14 against 0.85 on naturalness), and its measures agreed with each other. Its three top picks were all rewrites that kept every fact, and it ranked the two smallest local models' rewrites last on every measure.

This is one small experiment with twelve paragraphs, not a benchmark. But it points at a real difference. Kev is very good at the decision shapes it was trained on. Jev appears to generalize further, to questions its builders probably never trained for. If you can use Jev, use Jev.

Where the alternatives earn their place

For those of us who can't, the open models are more than a consolation prize, because we control them. Four opportunities stand out.

  • Train on your own decisions. Kev ships its full training recipe: LoRA adapters on Qwen3.5 base models, trained on labelled examples. Its authors report the 0.8B model trains on a 4 GB GPU. A model fine-tuned on a few thousand of your own labelled decisions could close much of the gap to Jev on the questions you actually ask. We haven't done that yet; it is our next step.
  • Calibrate on your own data. Several tools let you fit their confidence to your labelled examples (Rizzo has a rizzo calibrate command). Calibration doesn't change which answer wins, but it is what makes a confidence threshold something you can defend to an auditor.
  • Pin and audit the model. A local model is a file with a checksum. It never changes unless you change it, and every decision it makes stays on your network.
  • Tier the decisions. Let a fast local model take the confident, routine calls and escalate the rest to a person or, where policy allows, to Jev. Because every one of these tools speaks the same interface, switching between them is a configuration change.

Others are measuring too

We aren't the only ones testing this new category, and the independent results agree on the main points: Jev leads, and the small encoders trail the 4B models.

  1. jabr/classifier-benchmark: the public-domain suite we scored everything on. Its own results (Jev 0.964, Von 0.724, Laya 0.585) match ours.
  2. umstek/zero-shot-ie-bench: the broadest survey, covering 38 zero-shot extraction, classification and typed-decision systems, local and hosted, in nine languages.
  3. scienthoon/jev-ood-calibration: tests whether Jev's probabilities stay honest on a task it can't have seen. Accurate, but overconfident on rules that aren't in the text.
  4. cacan/system1-bench: a dashboard comparing Laya and several general LLMs against Jev on support-triage cases. Some of its answer keys were aligned with Jev's own answers, so Jev's 100% there is partly by construction.
  5. jaredpalmer/kev: Kev's frozen suites and scoring code. Our scorer's metric definitions are based on it.

What we add is the constraint: a fixed 8 GB memory budget, the larger open models quantized to fit it, and the question a compliance team actually asks, which is how much of this you can safely automate on hardware you control.

Try it yourself

The scorer, the harness, the quantization patches, the adapters, every result and the humanness experiment are at github.com/Autom8ly/gutcheck-bench. With any /v1/systemone server running, one command scores it:

python -m gutcheck.benchmark --endpoint http://127.0.0.1:8008 \
    --model kev-latest --suite suites/jabr-v2.jsonl --out runs/my-tool
Enter fullscreen mode Exit fullscreen mode

If you run it on different hardware, or on a labelled set from your own domain, I'd like to hear what you find.


About this article. This article and the code in the gutcheck-bench repository were mostly generated by AI (Claude, by Anthropic), directed by me in response to specific use cases we face in high-compliance environments. The measurements are real: every number comes from runs on the hardware described above, on 25 and 26 September 2026. We are not affiliated with TypeSafe or with any of the projects tested. Our Jev runs cost a few cents in total.

How we scored. We wrote a small, tool-neutral scorer (gutcheck) so that no contestant grades itself. Its metric definitions are based on Kev's open-source scorer (Apache-2.0), and we checked that it reproduces Kev's scores exactly on every run we made with both.

Caveats. One GPU, one machine, synthetic decision tasks. Your own decisions may rank these tools differently, so test on a labelled set from your domain before choosing. Jev's latency depends on your distance from TypeSafe's servers. All seven projects were days old when tested and are changing quickly; treat these numbers as a snapshot.

Top comments (0)