A support email lands in an agent's inbox.
Before the agent writes a single word back, it makes five small calls. Is this a prompt injection? How urgent is it? What does the customer actually want? Which tool runs first? Is the draft backed by the docs it pulled?
Every one of those is a multiple-choice question. In most stacks I have seen, every one goes to the biggest model on hand, which writes a paragraph that a regex then picks apart.
A model that only picks
Jeff is a 0.8B open model built for exactly those calls. You give it a situation and a list of labeled options. It gives back a probability for each option from one forward pass. It writes no text, so there is nothing to parse.
The author calls it a "System 1" model: the fast, reflexive half of a stack, with a big model as the slow, deliberate half. It was built at home as an open take on a hosted decision API, and the HN thread hit 574 points this week.
The base model handles any options you give it, zero-shot. On top sit nine adapters, small LoRA add-ons of about 41 MB each, one per job: a prompt-injection guard, ticket triage, tool choice, grounding checks, spam, and a few more. Per the README, each one trained in one epoch on one GPU, in half an hour to four hours.
The numbers that made me read the code
The headline setup is the one I care about. Jeff answers first. Only when it is unsure does the question go on to Qwen3.8-27B. Across eight adapters, against the 27B deciding everything alone:
| 27B decides alone | Jeff first, 27B when unsure | |
|---|---|---|
| Accuracy | 86.6% | 95.3% |
| Time per decision | 8.1 s | 0.25 s |
| Extra memory | 28.6 GB | +1.96 GB |
Those are the project's numbers, from an M4 Max with both models on MLX and the 27B's step-by-step reasoning turned off. The detail that earned my trust is how the "unsure" line was set. Each adapter's confidence threshold was picked on a separate set of calibration rows, and fixed before the test rows were scored. Plenty of model READMEs skip that step.
They also show the task where the small model does not win. On grounding (is this answer supported by its sources?), the 27B alone scores 96.7% and Jeff ends at 96.3%. That is one question in 300, at 20 times the speed. A results table with a near-loss in it reads like a measurement, not a pitch.
How it decides without writing
The server is a small FastAPI app. The route is /v1/systemone. Inside, the model scores every option and a softmax turns the scores into probabilities. The usage block it returns always says output_tokens: 0.
One flag I liked: orders: 2. It asks the same question a second time with the options reversed, then averages the two. Models lean toward options near the top of a list, and this cancels some of that out at the cost of a second pass.
The Python client, adapted from the README's own example (I did not run this):
from jeff import Client
from jeff.client import choice_question
tools = Client("http://localhost:8765", model="tools") # the tool-choice adapter
answers = tools.ask("User: main is red again, can you see why?", {
"tool": choice_question({
"ci_logs": "Read the latest CI run logs",
"git_log": "List the recent commits",
"web_search": "Search the web",
}, "Which tool should the agent call first?"),
})
pick = answers.choice("tool")
if pick.confidence < THRESHOLD:
... # hand this one to the big model
Every answer carries key, probability and confidence, where confidence runs from 0 (no better than a guess) to 1 (certain). That last field is the whole trick. The server also serves a playground page at its root URL, for trying questions by hand.
A decision is not a generation
Here is the claim I would build on, with or without Jeff. When your agent asks a big model "which tool?", you pay for a prompt, wait seconds, get prose back, and parse it. You also get no honest signal of how sure it was. A classifier hands you a calibrated probability, and the probability is the feature. It tells you exactly when to escalate.
That flips the usual design. The big model stops being the default for every small call and becomes the fallback for the hard ones. Your agent loop gets faster, your bill shrinks, and you gain something most stacks lack: a number that says "I am not sure, ask someone smarter."
The guard adapter is where I would start. A prompt-injection check in front of every tool output, at about a tenth of a second per check in their table, is cheap enough to run on everything.
Where it bites
The edges are real, and the issue tracker is honest about them.
- One decision at a time. The server takes its lock without waiting. A request that overlaps another gets HTTP 529 and a one-second retry hint. In issue #7, an eval at a concurrency of just 2 lost 100 of 287 rows to it. The Python client retries nothing, so you own the queue.
- Laptop numbers are not headline numbers. The roughly 30 ms figures come from big hardware. Issue #8 measured 144 to 274 ms medians on a laptop RTX 3070 Ti, using about 2 GB of its 8 GB.
- Zero-shot is uneven. In the HN thread, one user saw 70% on their own task where the hosted API they compared got 94%. Another called the 0.8B useless for job-ad labels and found the 2B better. The README's big numbers come from the adapters. On its own, the base scores 46.9% on the guard task. With the guard adapter, it scores 98.4%.
- It is a preview. The README calls v1.2 a community preview and says a long-term base, v1.3, is close. Adapters will not carry over between base versions, though your training data will.
The maintainer does move fast. Issue #1 found that no option past the 26th was ever chosen. That was fixed and released as v1.1 within days, with 32,000 extra long-list questions in the training data.
What I did not run
I read the README, the server and client code, three issues, and the HN thread. I did not run it. The install is a full Python ML environment plus a 1.7 GB model download, and that is more than I install for a post. Every number above is the project's own or comes from a named issue.
If you try it, start the server, open the playground, and throw your own five inbox questions at it.
Which decision in your agent would you hand to a 0.8B model first, and which one would you never let it make?
Top comments (0)