DEV Community

Cover image for AI on Edge: One forward pass, one letter: a decision model that runs on a laptop
chinmay garg
chinmay garg

Posted on AI-assisted

AI on Edge: One forward pass, one letter: a decision model that runs on a laptop

How I get a confidence score for every option without generating text, and why that score should only ever make a gate stricter.

A voice bot does not need a paragraph about what the caller wants. It needs a label, and a number that says how sure the model is. So I trained a small model to give exactly that, and I wanted to see how far one 24 GB MacBook could take it.

TypeSafe AI has described this style publicly as "decision models". This is my own version, built in the open-source forge fine-tuning tool. It is not their model, and I make no comparison to it.

The setup

The model is Gemma 4 E2B (4-bit) with a LoRA adapter, rank 16. The adapter is 52 MB. Training peaked at about 5.7 GB of memory and took about an hour. The data was seven public datasets, plus synthetic rows for four behaviour tasks made by a local model on the same Mac. No paid API and no cloud GPU.

Every example has the same shape: the conversation, a question, and lettered options.

Customer: I need to move my appointment, it's urgent.

Question: Which request is the customer making?
Options:
A) reschedule an appointment
B) cancel an appointment
C) ask about pricing
Enter fullscreen mode Exit fullscreen mode

The model is trained to answer with one letter. At test time I do not let it write anything. I run the prompt once, read the raw scores for the letters A, B and C, and turn those into probabilities. A temperature fitted on a validation set makes them well calibrated. There is nothing to parse, and the answer can never be an invalid label.

Three details mattered:

  • Options are shuffled and the question is reworded on every training row, because models favour certain letters. Evaluation keeps a fixed order.
  • Multi-select questions become many yes/no questions, so one format covers everything.
  • I read the scores in-process. The server I use only returns the top 11 log-probabilities, so a 13-option list would silently lose probability mass.

What good calibration gives you

On about 68,000 test rows, the calibration error was between 0.006 and 0.04 on most tasks. Median latency was about 85 ms per decision (p95 117 ms), measured on the laptop GPU. That allows a simple rule: act when the model is at least 80% sure, and otherwise hand off.

Test Handled at 0.8 or higher Accuracy on those
CLINC150 (held out) 90% 98.3%
MASSIVE intent 79% 95.4%
BANKING77 (held out) 74% 92.3%

The part that is not about the model

A confident 0.92 says the label is probably right. It does not say the agent is allowed to act. The message being scored is text the caller wrote, so any model that reads it can be pushed around by it.

This is how I would combine the two in an agent gateway:

check identity and scoped permission first
if denied: stop

score = decision_model(message)
if score is below the threshold: send to a human
else: go ahead
Enter fullscreen mode Exit fullscreen mode

The score can turn a "go ahead" into a "send to a human". It can never turn a "denied" into a "go ahead". Identity and scoped permission decide. The model only narrows. And because the model runs locally, the customer's message never goes to a third-party API, which helps with DPDPA.

What I am not claiming

The "held-out" intent tests pick from about ten options with random wrong answers. That is easier than the published 77-way and 151-way benchmarks.
The behaviour-task tests are small, and the test rows were made by the same kind of model that made the training rows.
I have not tested on real recorded calls. Everything above is on public data.

Where I am stuck

Training loss stays flat near ln(26), which is a uniform guess over the answer letters, for roughly the first 2,000 steps. Then it drops sharply. A lower learning rate moved the plateau but did not remove it. I do not know why.

Open question:

where should the hand-off threshold live? Per task, per data type, or per permission? I lean towards per data type. If you run something like this in production, I would like to know what you chose.

Top comments (0)