DEV Community

Cover image for Jev Open Source Alternatives: Understanding System One Before Building It Yourself
Ayush kumar
Ayush kumar

Posted on

Jev Open Source Alternatives: Understanding System One Before Building It Yourself

What Jev is, why it is different from an LLM, how its architecture works, and the open-source projects trying to recreate the idea

There is a pattern I keep noticing in AI applications.

We use a large language model for almost everything.

Need to classify a support ticket? Call an LLM.

Need to decide whether something is urgent? Call an LLM.

Need to route a request to the right workflow? Call an LLM.

Need to decide whether an agent should execute a tool? Call an LLM again.

It works, but there is something slightly strange about it.

The model is often being asked to produce a few words just so our application can throw those words away and turn them back into a boolean, a label, a score, or a branch in our code.

That is the problem TypeSafe is trying to approach differently with Jev.

In September 2026, TypeSafe introduced Jev as its first public System One Model: a model designed not primarily to write text for humans, but to make structured decisions that software can consume directly.

And that changes the way we think about the model.

Instead of:

Input
  ↓
LLM
  ↓
Generated text
  ↓
Parse JSON
  ↓
Validate
  ↓
Decision
Enter fullscreen mode Exit fullscreen mode

The idea becomes:

Input
  ↓
Jev
  ↓
Typed decision + probability
  ↓
Application
Enter fullscreen mode Exit fullscreen mode

That sounds like a small change.

It isn't.

What Is Jev?

Jev is TypeSafe's first public System One Model.

The simplest way to understand it is to think of Jev as something closer to a programmable intelligence primitive than a conventional chatbot.

You provide:

  • some state
  • one or more questions
  • the possible answers when necessary

and Jev returns structured decisions and probabilities.

For example, imagine a customer support application receives this:

"I was charged twice for my subscription, and I need a refund."

Instead of asking a language model to write:

This seems to be a billing issue.

and then parsing that response, we can ask a typed question:

Which team should handle this?

billing
technical
sales
Enter fullscreen mode Exit fullscreen mode

The result is something closer to:

{
  "choice": "billing",
  "probabilities": {
    "billing": 0.97,
    "technical": 0.02,
    "sales": 0.01
  }
}
Enter fullscreen mode Exit fullscreen mode

The application doesn't need to interpret prose.

It already knows the possible outputs.

That is the basic idea behind System One.

TypeSafe describes Jev as taking unstructured state and typed questions and returning typed probabilistic decisions. The company also describes its broader System One approach as a separate class of models built for automation rather than chat.

Why Did TypeSafe Build This?

This becomes much easier to understand when we look at how normal LLM applications work.

Suppose we need to answer a simple question:

Is this customer requesting a refund?

With a normal LLM, the model may generate:

Yes, the customer appears to be asking for a refund.
Enter fullscreen mode Exit fullscreen mode

Then the application has to turn that output into something useful:

if "yes" in response.lower():
    ...

Enter fullscreen mode Exit fullscreen mode

Or we force structured output:

{
  "refund_requested": true
}

Enter fullscreen mode Exit fullscreen mode

That is already much better.

But the model is still fundamentally generating a string representation of a decision.

And there is another problem.

Suppose the model says:

{
  "refund_requested": true,
  "confidence": 0.97
}
Enter fullscreen mode Exit fullscreen mode

Where did that 0.97 come from?

It is still text generated by the model.

TypeSafe argues that the probability used by software should be a property of the decision itself, rather than another sentence the language model happens to generate. Their System One approach is therefore built around typed outputs and calibrated probabilities.

That leads to a very different model interface:

State + Question
        ↓
   Decision model
        ↓
Probability distribution
        ↓
Typed answer

Enter fullscreen mode Exit fullscreen mode

System One vs Traditional LLMs

The name comes from the familiar distinction between System 1 and System 2 thinking.

TypeSafe uses “System One” as the name for a class of models intended to make fast, focused decisions inside software.

A traditional LLM is designed around language generation.

A simplified view looks like this:

Prompt
  ↓
Token 1
  ↓
Token 2
  ↓
Token 3
  ↓
Token 4
  ↓
...
  ↓
Final response
Enter fullscreen mode Exit fullscreen mode

Every generated token depends on what came before it.

That is incredibly useful when the output itself is language.

For a decision like:

Is this urgent?
Enter fullscreen mode Exit fullscreen mode

we don't necessarily need a paragraph.

We need:

true
Enter fullscreen mode Exit fullscreen mode

or perhaps:

P(urgent) = 0.91
Enter fullscreen mode Exit fullscreen mode

Jev is designed around that kind of output.

TypeSafe describes its sampler as parallel, rather than the sequential token generation used by conventional LLMs. It also says its training approach, called Reinforcement Learning for Calibrated Decisions (RLCD), is designed around producing calibrated decisions.

So the conceptual difference looks like this:

Traditional LLM

State
 ↓
Generate text
 ↓
Parse
 ↓
Validate
 ↓
Decision
Enter fullscreen mode Exit fullscreen mode

versus:

System One

State + Question
 ↓
Direct decision readout
 ↓
Probability distribution
 ↓
Typed value

Enter fullscreen mode Exit fullscreen mode

The Three Basic Decision Types

The Jev API revolves around three basic question types.

TypeSafe's public API currently exposes:

Choice
Noul
Score
Enter fullscreen mode Exit fullscreen mode

Choice

Use choice when exactly one option should be selected from a set.

For example:

Which team should handle this request?

billing
technical
sales
Enter fullscreen mode Exit fullscreen mode

The model returns the selected option and probabilities over the available options.

Conceptually:

billing    → 0.92
technical  → 0.05
sales      → 0.03

Enter fullscreen mode Exit fullscreen mode

Noul

noul is Jev's yes/no style decision.

The question is essentially:

Is this statement true?
Enter fullscreen mode Exit fullscreen mode

For example:

Is the customer asking for a refund?

Enter fullscreen mode Exit fullscreen mode

The result is a probability between 0 and 1.

P(true) = 0.94
Enter fullscreen mode Exit fullscreen mode

This is particularly useful in automation because the application can make the policy explicit:

if probability > 0.90:
    automate()
else:
    review()
Enter fullscreen mode Exit fullscreen mode

That is much closer to normal programming logic.

Score

The third primitive is score.

Instead of selecting one categorical answer, we provide an ordered scale.

For example:

How urgent is this request?

0 → low
1 → medium
2 → high
3 → critical

Enter fullscreen mode Exit fullscreen mode

The model can return both the probability distribution and an expected score.

So instead of:

urgency = "high"
Enter fullscreen mode Exit fullscreen mode

we can get something conceptually like:

low      0.02
medium   0.12
high     0.65
critical 0.21

expected score = 2.05
Enter fullscreen mode Exit fullscreen mode

That extra information matters when the surrounding application needs to make a threshold-based decision.

The Interesting Part: Probabilities Are the Output

This is probably the biggest conceptual difference between Jev and a normal chat model.

A normal LLM is ultimately optimized to produce a sequence of tokens.

Jev is designed around making decisions and exposing the probability distribution associated with those decisions.

So imagine this question:

Which workflow should handle this ticket?

A → billing
B → technical
C → account
Enter fullscreen mode Exit fullscreen mode

A normal language model might generate:

The correct answer is billing.

Enter fullscreen mode Exit fullscreen mode

Jev's interface is closer to:

A → 0.91
B → 0.06
C → 0.03
Enter fullscreen mode Exit fullscreen mode

There is no need to generate an explanation just to recover the decision.

That is why TypeSafe describes System One models as producing typed values directly, rather than strings that software subsequently has to parse.

What Does a Jev Request Look Like?

The API makes the design especially clear.

A simplified request looks like:

{
  "state": "I was charged twice for my subscription.",
  "questions": {
    "department": {
      "type": "choice",
      "instructions": "Which team should handle this?",
      "criteria": {
        "billing": "Payments and refunds",
        "technical": "Bugs and technical problems",
        "sales": "Purchases and pricing"
      }
    }
  }
}
Enter fullscreen mode Exit fullscreen mode

The important thing here is that the application defines the question and answer space.

The model isn't asked to invent a schema.

It is asked to answer within one.

The public TypeSafe API exposes this through POST /v1/systemone.

The Architecture Behind the Idea

Now we get to the really interesting part.

There are two levels of architecture we need to separate.

What TypeSafe has officially disclosed

TypeSafe says Jev uses:

  • a new model architecture
  • a parallel sampler
  • RLCD, or Reinforcement Learning for Calibrated Decisions

The company also explains that Jev does not need conventional autoregressive text generation for its decision outputs.

But TypeSafe has not published the complete internal architecture diagram or implementation.

So we should be careful about presenting deeper details as confirmed facts.

What independent research suggests

In September 2026, Archer Hume published a reverse-engineering investigation based on roughly 10,000 API calls to Jev. The investigation is explicitly presented as a reconstruction from observed behavior, not a disclosure from TypeSafe.

That investigation points toward an architecture with three important ideas:

Shared state
     ↓
Question-specific branches
     ↓
Direct probability readout

Enter fullscreen mode Exit fullscreen mode

This is where the architecture gets interesting.

A Simplified Jev Architecture

A useful conceptual representation is:

This is a conceptual model, not a published Jev source-code diagram.

The key idea is that the same state can be used for several independent questions.

For example:

State:
"Customer was charged twice, is angry and needs help today."

Questions:

1. Which team?
2. Is this urgent?
3. Is a refund requested?
4. Should this be escalated?
Enter fullscreen mode Exit fullscreen mode

Instead of running an entire LLM generation cycle four times, a System One design can treat these as separate decisions against the same state.

TypeSafe specifically highlights multiple independent questions as an important part of its workflow approach.

Shared State + Independent Questions

This is one of the most interesting properties reported in the independent architecture analysis.

Think about a request like this:

STATE
────────────────────────────
Customer:
"My payment failed twice..."
────────────────────────────

QUESTION 1
Which team should handle this?

QUESTION 2
Is the customer frustrated?

QUESTION 3
Does this require escalation?
Enter fullscreen mode Exit fullscreen mode

Conceptually:

The important idea is shared context with separate decisions.

The independent investigation found behavioral evidence consistent with question isolation and shared state, including tests where information inserted into one question did not behave like information placed in the common state. Again, this should be described as an experimental inference, not as an officially confirmed implementation detail.

There Is No Need to Generate a Paragraph

This is where the computational difference becomes easier to visualize.

A normal autoregressive model may need to produce:

"The"
"best"
"department"
"for"
"this"
"request"
"is"
"billing"
"."
Enter fullscreen mode Exit fullscreen mode

Even though the application ultimately only needs:


billing

Enter fullscreen mode Exit fullscreen mode

A System One design can instead read the decision directly from the model's internal computation.

Conceptually:

This is one of the central ideas behind the System One approach: don't generate language when the software only needs a decision.

TypeSafe says this parallel decision design is a major reason Jev can achieve much lower latency and cost for the tasks it targets.

What About the Model Backbone?

This is where we need to be especially careful.

You will find claims online about the exact Jev backbone, MoE configuration, parameter count, tokenizer internals, attention structure, and other implementation details.

Some of those come from reverse engineering.

They are interesting, but they are not equivalent to TypeSafe publishing the model architecture.

For this article, I would use this wording:

TypeSafe has publicly described Jev at the system level, including its new architecture, parallel sampler, and RLCD training method. The company has not published the complete internal model design. Independent researchers have therefore used black-box experiments to infer pieces of the implementation, including shared-state processing, isolated question branches, and direct probability readouts.

That keeps the article technically interesting without turning an inference into a fact.

Why This Architecture Makes Sense

Once you stop thinking of Jev as a chatbot, the design becomes much easier to understand.

An LLM is excellent when the output itself is language:

Write an email.
Explain this bug.
Generate Python code.
Summarize this document.

Enter fullscreen mode Exit fullscreen mode

But many software decisions look more like:

Should I?
Which one?
How urgent?
How risky?
Which route?
How confident?

Enter fullscreen mode Exit fullscreen mode

Those are much closer to functions.

For example:

is_urgent(ticket)
Enter fullscreen mode Exit fullscreen mode

or:

route_ticket(ticket)
Enter fullscreen mode Exit fullscreen mode

or:

risk_score(transaction)

Enter fullscreen mode Exit fullscreen mode

The TypeSafe vision is essentially to make AI intelligence usable in this form.

The company's own description is that System One models are meant to make intelligence usable directly inside software, with typed outputs and calibrated confidence.

Jev Is Not “Just a Smaller LLM”

This is probably the most important misconception to avoid.

If we reduce Jev to:

“It's just a smaller LLM that returns JSON.”

we miss the point.

The change is not simply model size.

The interface is different.

The objective is different.

The sampling strategy is different.

The output is different.

And the way the model is expected to interact with software is different.

A useful mental model is:

LLM
=
Language generation primitive

Jev / System One
=
Decision intelligence primitive
Enter fullscreen mode Exit fullscreen mode

That's the idea we need to keep in mind as we look at open-source projects.

And This Leads to the Open-Source Question

Once Jev appeared, another question became unavoidable:

What if we want this kind of decision model, but we don't want to depend on a proprietary hosted API?

That is where the ecosystem gets interesting.

There are now projects trying to reproduce or extend different parts of this idea:

But these projects are not all the same thing.

Some are complete decision models.

Some are Jev-compatible API implementations.

Some are serving layers.

Some are research replicas.

Some are ports of another decision model.

Part 2: The Open-Source Jev Ecosystem

Understanding Jev is one thing.

Running something similar yourself is another.

Once we started looking at the open-source ecosystem around System One decision models, we found something more interesting than a single “Jev alternative.”

There are multiple projects exploring the same basic problem from different directions.

Some train their own decision models.

Some reproduce Jev's API.

Some build a complete inference server around open weights.

Some focus on Apple Silicon.

Others are research projects trying to understand how far a very small model can go.

So before installing anything, it is worth putting the ecosystem into perspective.

                         Jev
                          │
                 System One idea
                          │
          ┌───────────────┼────────────────┐
          │               │                │
          ▼               ▼                ▼
    Decision Models   Compatible APIs   Runtimes / Ports
          │               │                │
      ┌───┼────┐       OpenJev          laya-mlx
      │   │    │       openjev-server
      │   │    │       SiliconLabAI
      ▼   ▼    ▼
    Laya Kev  Von
      │    │    │
      └────┼────┘
           │
      Nimble / SemIf
      NanoJev / Rizzo
Enter fullscreen mode Exit fullscreen mode

1. OpenJev

Repository: razorback16/openjev

Let's start with the project whose name makes the intention obvious.

OpenJev describes itself as an open-source System One decision server. It accepts a state and typed questions such as noul, choice, and score, then returns structured probabilities instead of generated text.

What makes this project particularly interesting is that it isn't tied to only one backend.

The current repository supports a larger open System One ecosystem, including:

OpenJev
├── DiffusionGemma
├── Laya
├── Verdict
├── CLM
└── JevK5
Enter fullscreen mode Exit fullscreen mode

The primary OpenJev model uses DiffusionGemma 26B-A4B, with 26B total parameters and about 4B active parameters, while the server can also route requests to smaller decision models such as Laya.

That makes OpenJev more than just another model.

It is really a model + serving + compatibility layer.

OpenJev architecture

The project uses a very different approach from a standard autoregressive LLM.

The repository describes a read-only diffusion-style canvas where the answer slots are masked, and the model reads the probabilities associated with those slots.

Conceptually:

                State
                  │
                  ▼
        ┌───────────────────┐
        │  Decision Canvas  │
        │                   │
        │ q1: [?]           │
        │ q2: [?]           │
        │ q3: [?]           │
        └─────────┬─────────┘
                  │
                  ▼
          DiffusionGemma
                  │
          ┌───────┼────────┐
          ▼       ▼        ▼
        q1 probs q2 probs q3 probs
          │       │        │
          └───────┼────────┘
                  ▼
          Structured answers
Enter fullscreen mode Exit fullscreen mode

The important detail is that the model does not need to generate a sentence and then have the server parse it.

Instead, OpenJev reads probabilities directly from the answer positions.

That gives us:

State
+
Question
+
Allowed answers
        ↓
Probability distribution
Enter fullscreen mode Exit fullscreen mode

rather than:

State
+
Question
        ↓
Generated text
        ↓
Parser
        ↓
Decision
Enter fullscreen mode Exit fullscreen mode

The repository also reports that uncertain questions can be reread with fresh noise and that multiple questions can be processed in parallel.

Installing OpenJev

For an NVIDIA machine, the project provides Docker-based deployment.

git clone https://github.com/razorback16/openjev
cd openjev

docker compose up -d
Enter fullscreen mode Exit fullscreen mode

Once the model has loaded:

curl localhost:8080/v1/models
Enter fullscreen mode Exit fullscreen mode

The repository also provides a direct Docker invocation:

docker run -d --gpus all --ipc=host -p 127.0.0.1:8080:8080 \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  razorback16/openjev:0.5.0
Enter fullscreen mode Exit fullscreen mode

The first launch downloads the model weights, so this is not a lightweight laptop setup for the main DiffusionGemma model.

Apple Silicon

OpenJev also provides an MLX backend:

pip install -e '.[mlx]'
OPENJEV_BACKEND=mlx python -m openjev
Enter fullscreen mode Exit fullscreen mode

The repository says the 4-bit DiffusionGemma weights require roughly 16 GB of memory, although the practical amount of free memory depends on the rest of the workload. (OpenJev repository)

That immediately makes the project interesting for Apple Silicon users.

2. OpenJev Server

Repository: abhishekgahlot2/openjev-server

This project solves a slightly different problem.

It is better thought of as a serving layer for open System One models.

The repository describes itself as:

One-pass decisions from any open language model, served as an API.

The important phrase is “any open language model.”

The server can sit on top of OpenJev models, but its architecture is designed to support other backends as well, including models running with vLLM and MLX. (openjev-server repository)

Architecture

The idea looks like:

                 Client
                   │
                   ▼
        /v1/systemone endpoint
                   │
            OpenJev Server
                   │
        ┌──────────┴──────────┐
        ▼                     ▼
      vLLM                   MLX
        │                     │
        ▼                     ▼
    Open weights         Apple Silicon
Enter fullscreen mode Exit fullscreen mode

The server handles several things around the model:

prompt construction
model probing
candidate readout
calibration
API compatibility
health checks
metrics
serving

Enter fullscreen mode Exit fullscreen mode

This is useful because a model and the server become separate components.

Installation

The project can be installed directly from GitHub:

pip install "openjev-server @ git+https://github.com/abhishekgahlot2/openjev-server"
Enter fullscreen mode Exit fullscreen mode

For a Mac, the repository currently documents:

pip install "openjev-server[mlx] @ git+https://github.com/abhishekgahlot2/openjev-server"

hf download openjev/openjev-MLX-4bit \
  --local-dir openjev-MLX-4bit

openjev serve \
  --backend mlx \
  --model openjev-MLX-4bit \
  --profile openjev \
  --port 3000
Enter fullscreen mode Exit fullscreen mode

The resulting server exposes:

POST /v1/systemone
Enter fullscreen mode Exit fullscreen mode

For example:

curl -s http://localhost:3000/v1/systemone \
  -H 'Content-Type: application/json' \
  -d '{
    "state": "Customer says their payment failed twice.",
    "questions": {
      "route": {
        "type": "choice",
        "instructions": "Which team should handle this?",
        "criteria": {
          "billing": "Payments and refunds",
          "technical": "Bugs and outages",
          "support": "General support"
        }
      }
    }
  }'
Enter fullscreen mode Exit fullscreen mode

This is where Jev compatibility becomes useful.

A client written against the TypeSafe System One API can be redirected to a locally hosted implementation without necessarily rewriting the whole application.

3. SiliconLabAI OpenJev

Repository: SiliconLabAI/OpenJev

This project is more of a playground and integration environment.

Its README describes three modes:

parallel
oneshot
decider
Enter fullscreen mode Exit fullscreen mode

Parallel

The project sends small scoring requests for individual options and combines the results into a probability distribution.

Oneshot

A single structured request is sent to an OpenAI-compatible model.

Decider

This mode connects OpenJev to an actual open System One model, currently through the Mapika/decider family.

So the architecture looks like:

       OpenJev UI
                      │
          ┌───────────┼───────────┐
          ▼           ▼           ▼
       parallel     oneshot     decider
          │           │           │
         LLM         LLM       System One
Enter fullscreen mode Exit fullscreen mode

Installation

For the LLM-backed playground:

git clone https://github.com/SiliconLabAI/OpenJev
cd OpenJev

npm install
cp .env.example .env

npm run dev

Enter fullscreen mode Exit fullscreen mode

Then open:

http://localhost:3001
Enter fullscreen mode Exit fullscreen mode

For the decider backend, the project documents starting the open model server separately:

pip install "git+https://github.com/Mapika/decider#egg=decider[serve]"
Enter fullscreen mode Exit fullscreen mode

Then:

scripts/serve.sh Mapika/decider-2b 8000
Enter fullscreen mode Exit fullscreen mode

And point OpenJev at:

http://localhost:8000
Enter fullscreen mode Exit fullscreen mode

This is a nice example of how the System One interface can become a common boundary between different implementations.

4. Laya

Repository: NandhaKishorM/laya

Laya is the project we have already explored in depth, and it deserves a central place in this ecosystem.

Unlike a server that simply wraps another model, Laya provides its own open-weight decision engine.

The project describes itself as a:

Multilingual, non-autoregressive System 1 decision engine.

Its current model family includes:

Laya
├── ModernBERT-large
├── mmBERT-base
└── typed-decision checkpoint
Enter fullscreen mode Exit fullscreen mode

The published checkpoints range from roughly 322M to 421M parameters.

Laya architecture

The simplified architecture is:

       State
                  │
                  ▼
            Tokenization
                  │
                  ▼
        Bidirectional Encoder
                  │
                  ▼
          Decision Layers
             /    |    \
            /     |     \
           ▼      ▼      ▼
        Choice   Noul   Score
           │      │      │
           ▼      ▼      ▼
      Probabilities / typed output
Enter fullscreen mode Exit fullscreen mode

This is quite different from the diffusion-based OpenJev design.

Laya is built around a bidirectional encoder and decision heads.

The input contains the state, instructions, and options. The network then scores the available decisions.

There is no need for the model to generate a paragraph such as:

"The customer appears to be dealing with a billing issue..."

The output can simply be:

billing

with probabilities attached to the candidate labels.

Installation

python -m pip install laya
Enter fullscreen mode Exit fullscreen mode

Then:

import laya

print(laya.__version__)
Enter fullscreen mode Exit fullscreen mode

For example:

from laya import Router

router = Router(preload=True)

state = """
I was charged twice for my subscription and want a refund.
"""

questions = {
    "department": {
        "type": "choice",
        "instructions": "Which department should handle this?",
        "criteria": {
            "billing": "Payments, invoices and refunds",
            "technical": "Bugs and technical problems",
            "other": "Anything else",
        },
    }
}

result = router.predict(state, questions)

print(result)
Enter fullscreen mode Exit fullscreen mode

That gives us a complete local decision workflow.

And because Laya has its own HTTP server, a compatible API layer can also expose the same model to applications that already understand the System One request format.

Interested in trying Laya yourself? I’ve written a full hands-on guide covering installation, usage, and my local experiments with the 421M-parameter decision engine.

Read the full Laya guide here

5. Kev

Repository: jaredpalmer/kev

Kev is one of the projects that most explicitly positions itself as a Jev alternative.

The repository calls it:

Small Jev-like decision models you can train and run yourself.

The project takes a different route from Laya.

Instead of a compact encoder-only model, Kev builds decision models on top of Qwen3.5 and Qwen3.8 with a LoRA adapter and pointer-style decision head.

The available model family currently spans:

Kev-0.5B
Kev-0.8B
Kev-4B
Kev-9B
Kev-27B

That gives Kev a fairly wide range of deployment targets.

Kev architecture

A simplified view is:

    State
                   │
                   ▼
              Qwen backbone
                   │
              LoRA adapter
                   │
                   ▼
             Pointer Head
                   │
          ┌────────┼────────┐
          ▼        ▼        ▼
        Choice    Noul     Score
          │        │        │
          └────────┼────────┘
                   ▼
             Probabilities
Enter fullscreen mode Exit fullscreen mode

The repository explains that multiple questions can share the state while remaining isolated from each other's answers.

This is important.

Suppose we ask:

Which team?
Is it urgent?
How frustrated is the customer?

Enter fullscreen mode Exit fullscreen mode

The questions can use the same state, but one question does not simply receive the answer from another as additional context.

That makes the design much closer to independent decision functions than a normal conversational interaction.

Installing Kev

Kev currently uses uv.

First:

git clone https://github.com/jaredpalmer/kev.git
cd kev
Enter fullscreen mode Exit fullscreen mode

Then:

uv sync --extra serve
Enter fullscreen mode Exit fullscreen mode

And launch a model:

uv run --extra serve \
  python -m kev.serve \
  --run jaredpalmer/kev-4b \
  --port 8009

Enter fullscreen mode Exit fullscreen mode

Now we can send a System One-compatible request:

curl -s localhost:8009/v1/systemone \
  -H 'content-type: application/json' \
  -d '{
    "state": "I was charged twice. Please fix this ASAP.",
    "model": "kev-latest",
    "questions": {
      "billing": {
        "type": "noul",
        "instructions": "Is this ticket about billing?"
      },
      "tone": {
        "type": "choice",
        "instructions": "What is the customer tone?",
        "criteria": {
          "calm": null,
          "frustrated": null,
          "angry": null
        }
      },
      "urgency": {
        "type": "score",
        "instructions": "How urgent is this ticket?",
        "criteria": [
          "can wait",
          "this week",
          "today"
        ]
      }
    }
  }'
Enter fullscreen mode Exit fullscreen mode

One of the interesting properties of Kev is that the project also includes fine-tuning and deployment workflows rather than stopping at inference. (Kev repository)

So Kev is not just:

download model → run inference

It can also be:

your labelled data
       ↓
fine-tune Kev
       ↓
calibrate
       ↓
serve
Enter fullscreen mode Exit fullscreen mode

That makes it especially relevant when the generic model isn't good enough for a company's own decision schema.

6. Nimble

Repository: Bespoke Labs Nimble

Nimble takes a very focused approach.

The repository describes it as:

Data, Model, Recipe for an open Jev.

The current model is Bespoke-Nimble-9B, based on Qwen3.5-9B.

The interesting part isn't only the model.

The repository publishes the data curation and training recipe as well.

That is important because one of the questions surrounding System One models is not simply:

“Which backbone are you using?”

but:

“How do you train a model to make decisions instead of generating answers?”

Nimble tries to make that process reproducible.

Nimble architecture

Its implementation follows a simple pattern:

      Context
                  │
                  ▼
             Qwen3.5-9B
                  │
             LoRA training
                  │
                  ▼
             Answer tokens
                  │
                  ▼
           Logit readout
                  │
                  ▼
            Probabilities
Enter fullscreen mode Exit fullscreen mode

Nimble explicitly says its scorer reads the logits for the one-token answer codes.

For example:

A → billing
B → technical
C → sales
Enter fullscreen mode Exit fullscreen mode

The server can read the probability associated with those candidate tokens directly.

So the application gets:

billing    → 0.95
technical  → 0.03
sales      → 0.02
Enter fullscreen mode Exit fullscreen mode

without requiring the model to generate an entire JSON structure.

Installation

The repository currently recommends Python 3.12.

git clone https://github.com/bespokelabsai/nimble.git nimble
cd nimble
Enter fullscreen mode Exit fullscreen mode

Create an environment:

python3.12 -m venv .cache/venvs/nimble
source .cache/venvs/nimble/bin/activate
Enter fullscreen mode Exit fullscreen mode

Then install the required packages:

python -m pip install torch==2.8.0 -r requirements/training.txt
Enter fullscreen mode Exit fullscreen mode

The latest 9B model is large enough that memory requirements matter, especially when merging LoRA weights.

The repository documents both Apple Silicon and NVIDIA GPU workflows. (Nimble repository)

The thing I particularly like here is that the repository not only shows a final model card.

It gives us access to:

training data
training recipe
evaluation code
latency benchmarks
Jev comparisons
Enter fullscreen mode Exit fullscreen mode

That makes Nimble interesting from a research perspective as much as an application perspective.

7. SemIf

Repository: TheoLeeCJ/SemIf-OpenJev

SemIf is another particularly interesting project because of its history.

It was originally called OpenJev and was later renamed SemIf.

The authors explicitly describe it as an independent project that reproduces the interface pattern of Jev rather than its undisclosed proprietary model or training process.

That distinction is worth preserving.

Jev implementation
       ≠
SemIf implementation
Enter fullscreen mode Exit fullscreen mode

Instead:

Jev interface idea
        ↓
open model
        ↓
native option readout
Enter fullscreen mode Exit fullscreen mode

SemIf architecture

The core idea is very easy to visualize:

             State
                │
                ▼
         Qwen / open model
                │
        ┌───────┼────────┐
        ▼       ▼        ▼
      Option A Option B Option C
        │       │        │
      logits  logits   logits
        │       │        │
        └───────┼────────┘
                ▼
           softmax
                │
                ▼
          probabilities
Enter fullscreen mode Exit fullscreen mode

The project emphasizes that the options are scored directly.

The model doesn't have to produce:

"The answer is A because..."

The system only needs the relevant option scores.

Installation

For a normal GPU setup:

python -m venv .venv
. .venv/bin/activate

pip install -e '.[test]'
Enter fullscreen mode Exit fullscreen mode

For Apple Silicon, the repository provides MLX support:

pip install -e '.[test,mlx]'
Enter fullscreen mode Exit fullscreen mode

There are also PyTorch/MPS and llama.cpp paths.

A basic scoring example looks like:

CUDA_VISIBLE_DEVICES=0 semif-score \
  --mode direct \
  --model Qwen/Qwen3.5-4B \
  --input examples/decisions.jsonl \
  --output results.jsonl

Enter fullscreen mode Exit fullscreen mode

One of the most useful things in SemIf is the amount of reproducibility work around the implementation.

The repository records model revisions, prompt hashes, timing data, and benchmark fixtures, which makes it easier to understand what exactly was tested. (SemIf repository)

8. Von

Repository: wfzyx/von

Von is another full System One implementation.

Its repository describes it as:

An Open-Source, Non-Autoregressive System One Decision Model.

Von exposes three core decision primitives:

Choice
Noul
Score
Enter fullscreen mode Exit fullscreen mode

But the project goes beyond a simple API wrapper.

It also includes calibration work, Python and TypeScript SDKs, local inference, serving, and benchmark tooling.

Von architecture

The conceptual pipeline looks like this:


                     State
                       │
                       ▼
             Bidirectional processing
                       │
              ┌────────┼────────┐
              ▼        ▼        ▼
            Choice    Noul     Score
              │        │        │
              ▼        ▼        ▼
          Softmax   Sigmoid   Distribution
              │        │        │
              └────────┼────────┘
                       ▼
                Structured result
Enter fullscreen mode Exit fullscreen mode

One feature of the current Von implementation is particularly interesting.

The project has worked on option-order invariance.

Why does that matter?

Imagine:

A = billing
B = technical
C = sales

Enter fullscreen mode Exit fullscreen mode

and then we reorder it:

A = sales
B = billing
C = technical
Enter fullscreen mode Exit fullscreen mode

A good decision model should not suddenly change its mind just because the labels moved.

Von's recent architecture changes explicitly target this behavior by isolating option representations so that one option does not gain an advantage simply because of its position in the list. (Von repository)

That's exactly the kind of issue worth testing in System One models.

Installation

Python:

pip install von-sdk

Enter fullscreen mode Exit fullscreen mode

Or directly from GitHub:

pip install git+https://github.com/wfzyx/von.git

Enter fullscreen mode Exit fullscreen mode

Then:

import von

result = von.decide(
    state="Database replication lag exceeded 45 seconds.",
    choices={
        "infrastructure": "Database, hardware, network, or server failures",
        "billing": "Invoices, payments and refunds",
        "feature_request": "Requests for new platform capabilities",
    },
    instructions="Classify the root cause domain of this incident.",
)

print(result.choice)
print(result.confidence)
print(result.probabilities)
Enter fullscreen mode Exit fullscreen mode

Von also provides a JavaScript/TypeScript SDK.

That makes it one of the projects where the developer experience around the decision model is almost as important as the model itself.

9. NanoJev

Repository: TianyuCodings/NanoJev

NanoJev takes the name literally.

The project describes itself as:

A nano replica of Jev.

The current model is only 0.6B parameters, built around Qwen3-0.6B with decision heads.

That is interesting because one of the biggest questions around System One isn't:

“Can a huge model make decisions?”

Of course it can.

The more interesting question is:

How small can a useful decision model become?

NanoJev is directly exploring that question.

NanoJev architecture

The project processes:

state
+
question
+
candidate actions
Enter fullscreen mode Exit fullscreen mode

and produces:

probability distribution

Enter fullscreen mode Exit fullscreen mode

without normal output-token generation.

A simplified diagram:

    State
                  │
                  ▼
             Qwen3-0.6B
                  │
          Decision heads
          /      |      \
         ▼       ▼       ▼
      Choice   Boolean   Score
         │       │        │
         └───────┼────────┘
                 ▼
           Probability
             outputs
Enter fullscreen mode Exit fullscreen mode

One of the things that makes NanoJev unusual is its game-based evaluation.

The project currently evaluates the decision model on environments such as:

Maze
Snake
ViZDoom
Enter fullscreen mode Exit fullscreen mode

where the model has to choose actions instead of generating natural-language answers.

That makes the concept much easier to visualize.

An agent doesn't necessarily need the model to tell it:

"I recommend moving north because..."

It can simply receive:

North  → 0.78
East   → 0.10
South  → 0.04
West   → 0.08
Enter fullscreen mode Exit fullscreen mode

and act.

Installation

git clone https://github.com/TianyuCodings/NanoJev.git
cd NanoJev

python -m pip install -r requirements-toy.txt huggingface_hub
Enter fullscreen mode Exit fullscreen mode

Then download the model:

from huggingface_hub import snapshot_download

snapshot_download(
    repo_id="C-Tianyu/NanoJev",
    revision="unified-games-v1",
    local_dir="checkpoints/NanoJev-unified",
    allow_patterns=[
        "best.safetensors",
        "config.json",
        "tokenizer/*",
        "backbone_config/*",
    ],
)

Enter fullscreen mode Exit fullscreen mode

Start the inference server:

python scripts/serve_decisions.py \
  --checkpoint-dir checkpoints/NanoJev-unified \
  --web-root web \
  --port 8765 \
  --disable-native-triton
Enter fullscreen mode Exit fullscreen mode

The API is exposed at:

POST http://127.0.0.1:8765/api/evaluate

Enter fullscreen mode Exit fullscreen mode

10. Rizzo Flow

Repository: Rizzo-AI-Academy/rizzo-flow

Rizzo Flow is interesting because it explicitly separates its own decision API from its Jev-compatible API.

The project states that it does not reproduce Jev's proprietary architecture or TypeSafe's RLCD implementation.

Instead, it reproduces the interface pattern using an open model stack. The current implementation is based on Spark-X2.5 and includes a LoRA fine-tuning path. (Rizzo Flow repository)

That makes it a useful example of a broader idea:

You don't have to reproduce the exact Jev internals to build a System One-style software interface.

Rizzo Flow architecture

The project exposes two conceptual API layers:

                Rizzo Flow
                      │
          ┌───────────┴───────────┐
          │                       │
          ▼                       ▼
   Native decisions        Jev-compatible API
          │                       │
          ▼                       ▼
 boolean / choice /         noul / choice /
 score / numeric             score
Enter fullscreen mode Exit fullscreen mode

The native API also includes capabilities such as abstention and policy metadata that aren't part of the basic Jev-compatible contract.

This makes Rizzo Flow less of a one-to-one replica and more of an independent runtime inspired by the same design space.

11. Laya-MLX

Repository: mizorewww/laya-mlx

This one needs a separate category.

It is not an independent Jev alternative in the same sense as Kev, Von, or NanoJev.

It is an independent MLX implementation of Laya designed specifically for Apple Silicon.

The repository says it preserves Laya's:

weights
prompt format
calibration
output schema
Enter fullscreen mode Exit fullscreen mode

while moving inference to the MLX stack.

So the relationship is:

Laya
 │
 └── laya-mlx
      │
      └── Apple Silicon / MLX

Enter fullscreen mode Exit fullscreen mode

rather than:

Jev
 │
 └── completely different model
Enter fullscreen mode Exit fullscreen mode

Architecture

The high-level path becomes:

State + Question
       │
       ▼
MLX tokenizer
       │
       ▼
Laya decision model
       │
       ▼
Decision head
       │
       ▼
Probabilities
Enter fullscreen mode Exit fullscreen mode

The current repository provides pre-converted checkpoints and reports local benchmarks on Apple Silicon.

Installation is straightforward:

pip install laya-mlx

Enter fullscreen mode Exit fullscreen mode

Then:

import laya_mlx as laya

agent = laya.load("aac6fef/laya-mlx")

result = agent.predict(
    "I was billed twice. Please refund the duplicate.",
    {
        "department": {
            "type": "choice",
            "instructions": "Who should handle this?",
            "criteria": [
                "billing",
                "technical",
                "sales",
            ],
        }
    },
)

print(result["answers"]["department"])
Enter fullscreen mode Exit fullscreen mode

For people with Apple Silicon, this is one of the more practical ways to explore the decision-model idea without bringing in a large cloud dependency.

What We Actually Have

After looking through all of these projects, the ecosystem starts to make more sense.

They aren't all competing in the same category.

A cleaner way to think about them is:

| Project                  | Category                           | Main idea                                             |
| ------------------------ | ---------------------------------- | ----------------------------------------------------- |
| **OpenJev**              | Full server + open model ecosystem | Open System One server with multiple backends         |
| **openjev-server**       | Serving layer                      | Host open decision models behind a Jev-compatible API |
| **SiliconLabAI/OpenJev** | Playground / integration           | Explore multiple System One backends                  |
| **Laya**                 | Decision model                     | Small non-autoregressive decision engine              |
| **Kev**                  | Decision model family              | Qwen-based Jev-like models                            |
| **Nimble**               | Decision model + recipe            | Open model, data and training methodology             |
| **SemIf**                | Research implementation            | Direct option-logit readout                           |
| **Von**                  | Decision model + SDK               | Calibrated System One runtime                         |
| **NanoJev**              | Research model                     | Very small 0.6B Jev-style system                      |
| **Rizzo Flow**           | Decision runtime                   | Native decision API + Jev-compatible API              |
| **Laya-MLX**             | Laya port/runtime                  | Apple Silicon implementation of Laya                  |

Enter fullscreen mode Exit fullscreen mode

Which Architecture Is Everyone Exploring?

Despite all the differences, there is a common pattern underneath many of these projects:

              Unstructured state
                      │
                      ▼
             Typed question
                      │
                      ▼
              Decision model
                      │
          ┌───────────┼───────────┐
          ▼           ▼           ▼
        Choice       Noul        Score
          │           │           │
          ▼           ▼           ▼
      probability probability probability
          │           │           │
          └───────────┼───────────┘
                      ▼
               Application logic
Enter fullscreen mode Exit fullscreen mode

The model does not need to be the application.

It doesn't need to generate the whole workflow.

It can simply become the small intelligence component sitting between unstructured input and deterministic software logic.

That's probably the most important idea we took away from looking at this ecosystem.

Conclusion

Jev introduces a different way of thinking about where AI belongs inside software.

Instead of asking a model to generate text for every decision, the System One approach treats many of those decisions as their own primitive: give the model a state, define the question, specify the possible outcomes, and let the application consume the resulting decision and probabilities.

The open-source ecosystem shows that this idea is already being explored in several different ways.

Projects like Laya, Kev, Nimble, SemIf, Von, NanoJev, and Rizzo Flow experiment with different model architectures, training approaches, runtimes, and deployment strategies. Other projects, such as OpenJev and openjev-server, focus more heavily on making these models usable through a Jev-compatible API and serving layer. Laya-MLX, meanwhile, shows how the same decision-model idea can be adapted to a specific hardware ecosystem such as Apple Silicon.

What makes this ecosystem interesting is that there isn't a single implementation of the idea.

Some projects focus on very small models.

Others use larger language-model backbones.

Some optimize for calibration.

Some focus on latency.

Some focus on compatibility.

And some are primarily research projects trying to understand how these decision systems should actually work.

The common thread is simpler:

Unstructured information
        ↓
Typed question
        ↓
Decision model
        ↓
Probability
        ↓
Software action
Enter fullscreen mode Exit fullscreen mode

That may become an important building block for AI applications where the model doesn't need to talk to the user at every step.

It just needs to make the right decision at the right time.

And that is what makes the open-source Jev ecosystem worth watching.

In the next article, we'll move from concepts to code and actually install and run these open-source Jev alternatives, explore their architectures in practice, and test how they behave on the same decision workloads.

Thank you so much for reading

Like | Follow | Subscribe to the newsletter.

Connect with me on:

GitHub: https://github.com/Ayush7614

LinkedIn: https://www.linkedin.com/in/ayush-kumar-984443191/

Twitter: https://x.com/AYUSHKUMAR82274

Substack: https://substack.com/@felixayush

Dev.to Blog: https://dev.to/ayush7614

Personal Blog: https://neural-verse-peach.vercel.app/

Website: https://ayushbuilds-dev.vercel.app/

Top comments (0)