Absolutely bhai. I checked the latest llama.cpp decision-model implementation, and your original post is actually outdated in one important way: the current approach is no longer the old phi-2-decision-model + ./main + GBNF workflow. As of October 2026, llama.cpp has a native /v1/systemone decision-model endpoint with models such as Julia-1, Laya, Kev-4B, lev, and OpenJev. Hugging Face
Below is a fresh DEV.to-ready article. I deliberately removed Liquid syntax, raw/endraw, outdated commands, and unnecessary escaping so you should be able to paste it directly into DEV.
Run Local AI Decision Models with llama.cpp: Fast, Private Classification Without an API
What if your application doesn't need an LLM to write an answer?
Sometimes you only need a decision.
For example:
- Which team should handle this support ticket?
- Is this message spam?
- Is the customer angry?
- Is this request urgent?
- Did an AI agent successfully complete its task?
- Should this request be approved or rejected?
Using a full conversational LLM for every one of these tasks can be unnecessary.
You don't always need generated text.
You need a reliable decision.
That's where the new decision model support in llama.cpp becomes interesting.
What Are Decision Models?
A traditional LLM generates text token by token.
For example:
User: I was charged twice for my order.
LLM:
I'm sorry to hear that. It looks like this may be a billing issue...
But if your application only needs:
{
"route": "billing"
}
generating a paragraph first is unnecessary work.
A decision model approaches the problem differently.
Instead of generating a long response, it scores the options you provide and returns the most likely decision along with probabilities.
For example:
Input:
"I was charged twice for my order."
Options:
billing
shipping
technical
Result:
billing → 90.49%
shipping → 2.75%
technical → 6.76%
This makes decision models particularly interesting for:
- Classification
- Routing
- Content moderation
- Agent evaluation
- Intent detection
- Approval workflows
- Priority detection
- Yes/no decisions
- Structured automation
The latest llama.cpp implementation exposes this through the /v1/systemone endpoint. Hugging Face
Why Not Just Use a Normal LLM?
You absolutely can.
But there is an important difference.
A normal chat model might need to generate:
The customer appears to have a billing-related problem because
they were charged twice for the same order...
Your application then has to parse that response.
A decision model can instead directly answer:
{
"choice": "billing",
"probabilities": {
"billing": 0.9049,
"shipping": 0.0275,
"technical": 0.0676
}
}
There is no generated explanation to parse.
The output is already structured for a decision-making system.
The llama.cpp implementation can also answer multiple questions about the same state, which makes it useful for agentic workflows. Hugging Face
Decision Models Available for llama.cpp
The current llama.cpp ecosystem already includes several decision models.
Some examples include:
| Model | Approx. Size | Use |
|---|---|---|
| Julia-1 | 144M | Lightweight classification |
| Laya | 421M | English classification |
| Kev-4B | 4B | More capable text decisions |
| lev | 4B | Zero-shot classification |
| OpenJev | 27B | Multilingual + vision |
| Clef | 27B | Multimodal decisions |
For example, the published benchmark reports very low median decision latency for the smaller models on an NVIDIA RTX PRO 6000, although your actual performance will depend heavily on your hardware and configuration. Hugging Face
The important idea isn't simply "smaller model = faster."
It's:
Don't use a generative model when your application only needs a decision.
Getting Started
The latest llama.cpp CLI can download and run compatible GGUF models directly.
For example:
llama serve -hf ggml-org/Kev-4B-GGUF
This starts the llama.cpp server with the Kev-4B decision model. Hugging Face
If you prefer building llama.cpp yourself, the project uses CMake:
git clone https://github.com/ggml-org/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build --config Release
The current project uses the llama- command naming convention, such as llama-cli and llama-server; the old main executable name is no longer the current interface. GitHub
Your First Decision
Let's build a simple customer-support router.
Imagine this incoming message:
I was charged twice for my order last week
and nobody has replied.
We want the system to answer three questions:
- Which team should handle it?
- Is the customer angry?
- How urgent is the issue?
This is where decision models become much more interesting than simple classification.
Using the llama.cpp Decision API
Once the server is running, send a request to:
http://localhost:8080/v1/systemone
Here's a complete example:
curl http://localhost:8080/v1/systemone \
-H "Content-Type: application/json" \
-d '{
"state": "Customer message: I was charged twice for my order last week and nobody has replied.",
"questions": {
"route": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "payments, charges, refunds, invoices",
"shipping": "delivery, tracking, lost or late parcels",
"technical": "bugs, errors, login problems"
}
},
"angry": {
"type": "noul",
"instructions": "Is the customer angry?"
},
"urgency": {
"type": "score",
"instructions": "How urgent is this?",
"criteria": [
"can wait",
"this week",
"today",
"right now"
]
}
}
}'
Notice something important.
We didn't ask the model to write an answer.
We gave it:
state
+
questions
+
possible decisions
and asked it to score those decisions.
Three Types of Decisions
The current System One API supports three useful question types.
1. Choice
Use choice when you want exactly one option from a list.
For example:
{
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "payments, charges, refunds, invoices",
"shipping": "delivery, tracking, lost parcels",
"technical": "bugs, errors, login problems"
}
}
Possible result:
{
"choice": "billing"
}
The response also includes probabilities for the available choices.
2. Yes / No
For binary decisions, use noul.
For example:
{
"type": "noul",
"instructions": "Is the customer angry?"
}
The model returns the probability that the answer is yes.
Example:
{
"noul": 0.8208
}
That means the model estimates an approximately 82% probability for "yes."
3. Score
Sometimes a classification label isn't enough.
You may want a scale.
For example:
0 → Can wait
1 → This week
2 → Today
3 → Right now
You can represent that using score:
{
"type": "score",
"instructions": "How urgent is this?",
"criteria": [
"can wait",
"this week",
"today",
"right now"
]
}
The result contains an expected score plus the probability distribution across the levels.
This is useful for:
- Priority
- Risk
- Severity
- Confidence
- Customer sentiment
- Agent evaluation
The current server documentation describes all three question types and their response formats. GitHub
Example Response
A response might look like this:
{
"model": "ggml-org/Kev-4B-GGUF",
"answers": {
"route": {
"type": "choice",
"choice": "billing",
"probabilities": {
"billing": 0.9049,
"shipping": 0.0275,
"technical": 0.0676
},
"confidence": 0.8574
},
"angry": {
"type": "noul",
"noul": 0.8208
},
"urgency": {
"type": "score",
"score": 2.2821,
"legend": {
"0": "can wait",
"1": "this week",
"2": "today",
"3": "right now"
}
}
},
"usage": {
"input_tokens": 130,
"output_tokens": 0
}
}
Look at the interesting part:
output_tokens: 0
The model isn't producing a conversational response.
It's making decisions about the supplied state.
Why This Is Interesting for AI Agents
This is where I think decision models become particularly useful.
Imagine an AI agent working through a task:
User Request
|
v
AI Agent
|
v
Execute Action
|
v
Decision Model
|
+---- Success ----> Continue
|
+---- Uncertain --> Human Review
|
+---- Failed -----> Retry
Instead of asking a large LLM:
"Did my previous action succeed?"
and parsing a natural-language response, a decision model can evaluate predefined outcomes.
For example:
{
"type": "choice",
"criteria": {
"success": "The requested action completed successfully",
"failure": "The action failed",
"uncertain": "There is not enough evidence"
}
}
Now your application can implement deterministic business logic around the result.
Confidence Thresholds Matter
One of the most important things to understand is that a probability is not automatically a guarantee of correctness.
For example:
billing: 0.96
might be considered sufficiently confident for automatic routing.
But:
billing: 0.51
shipping: 0.47
is much less convincing.
A practical architecture could therefore be:
Decision Model
|
v
Confidence Check
|
+---- >= 0.90 ---> Automatic Action
|
+---- < 0.90 ----> Human Review
The exact threshold should be determined using your own evaluation data.
The llama.cpp decision-model documentation specifically recommends testing confidence cutoffs for the model and task instead of assuming one universal threshold. Hugging Face
Don't Use Bare Labels
Here's another subtle but important point.
Instead of giving the model:
{
"billing": null,
"shipping": null,
"technical": null
}
provide descriptions:
{
"billing": "payments, charges, refunds, invoices",
"shipping": "delivery, tracking, lost or late parcels",
"technical": "bugs, errors, login problems"
}
Why?
Because the descriptions give the model semantic information about what each option means.
This can significantly improve routing quality compared with ambiguous labels alone. Hugging Face
Multiple Questions in One Request
Another nice feature is that you can ask multiple questions about the same state.
For example:
Customer Message
|
+---- Route?
|
+---- Angry?
|
+---- Urgency?
|
+---- Refund required?
|
+---- Escalate?
Instead of making separate model calls for every question, the API allows you to batch questions.
This is particularly useful for agent pipelines and high-volume classification systems. Hugging Face
What About Images?
This gets even more interesting.
Some decision models support vision.
For example, OpenJev can be used with images.
That means you can send a screenshot or document and ask questions such as:
Is this an invoice?
Does this document contain a table?
Is the uploaded document a receipt?
Does this screenshot show an error?
The image can be supplied as a data URL, and the decision API can return the corresponding result. Hugging Face
This opens up interesting local AI workflows:
Image
|
v
Local Vision Decision Model
|
+---- Invoice
|
+---- Receipt
|
+---- Contract
|
+---- Other
No external API is required for the inference itself.
Decision Models vs Traditional LLMs
Here's how I think about the difference:
| Task | Better Fit |
|---|---|
| Write an email | Generative LLM |
| Explain a concept | Generative LLM |
| Write code | Generative LLM |
| Chat with a user | Generative LLM |
| Route a support ticket | Decision model |
| Detect intent | Decision model |
| Yes/no validation | Decision model |
| Assign priority | Decision model |
| Moderate content | Decision model |
| Evaluate an agent step | Decision model |
| Classify a document | Decision model |
| Generate a report | Generative LLM |
It's not about replacing LLMs.
It's about using the right model for the job.
Where GBNF Still Fits
This doesn't mean grammar-constrained generation has become irrelevant.
llama.cpp still supports grammars and structured generation.
You can constrain model output using:
--grammar
or:
--grammar-file
and llama.cpp also supports JSON-schema-based structured output. GitHub
For example:
llama-cli \
-m model.gguf \
--grammar-file grammar.gbnf \
-p "Classify this request"
Grammar-constrained generation is useful when you need a generated structure.
Decision models are different.
They are designed specifically for choosing between predefined outcomes.
That's an important distinction.
Local AI Changes the Architecture
Traditional architecture:
Application
|
v
Cloud API
|
v
Large LLM
|
v
Generated Response
Decision-model architecture:
Application
|
v
Local llama.cpp
|
v
Decision Model
|
v
Structured Decision
This can reduce:
- API dependency
- Network latency
- API costs
- Data leaving your environment
- Parsing complexity
And it can make high-frequency decision workloads much easier to run locally.
Of course, actual latency and accuracy depend on your model, hardware, workload, and evaluation setup.
When Should You NOT Use a Decision Model?
Decision models aren't a replacement for general-purpose LLMs.
Don't use one when you need:
- Long-form generation
- Creative writing
- Complex explanations
- Open-ended conversations
- Code generation
- Multi-step reasoning where the output isn't predefined
For those cases, a capable instruction-tuned LLM is still the better choice.
Decision models shine when the question is:
"Which decision should I make?"
rather than:
"What should I write?"
The Bigger Picture
I think this is one of the more interesting directions for local AI.
We're moving from:
One giant LLM for everything
towards:
┌── Generative LLM
│
User/Application ───┼── Decision Model
│
├── Embedding Model
│
├── Vision Model
│
└── Small Specialized Model
Instead of asking one model to perform every task, we can choose specialized models based on the actual job.
For agentic systems, that can become particularly powerful:
┌───────────────┐
│ User Input │
└───────┬───────┘
│
v
┌──────────────────┐
│ Decision Model │
│ Route / Validate │
└────────┬─────────┘
│
┌──────────┴──────────┐
│ │
v v
Generative LLM Local Tool
Complex reasoning Deterministic action
│ │
└──────────┬──────────┘
v
┌──────────────────┐
│ Decision Model │
│ Validate result │
└──────────────────┘
That architecture can be more efficient than throwing every step at a large conversational model.
Final Takeaway
The interesting part of llama.cpp isn't just that it can run LLMs locally.
It's becoming a broader local inference engine for specialized AI workloads.
Decision models add another tool to the toolbox:
Generate when you need generation.
Decide when you need a decision.
And for applications performing thousands or millions of small decisions, that distinction can matter.
Resources
- llama.cpp: https://github.com/ggml-org/llama.cpp
- llama.cpp server documentation: https://github.com/ggml-org/llama.cpp/tree/master/tools/server
- llama.cpp grammars: https://github.com/ggml-org/llama.cpp/tree/master/grammars
- Decision Models: https://huggingface.co/blog/ggml-org/decision-models-in-llamacpp
- llama.cpp Decision Models collection: https://huggingface.co/collections/ggml-org/decision-models
- GGUF models: https://huggingface.co/models?library=gguf
What do you think?
Would you use a small local decision model inside an AI agent instead of calling a large LLM for every routing, validation, and classification step?
I'd be especially interested in real-world use cases for:
AI Agents → Routing → Validation → Human-in-the-loop
Top comments (0)