DEV Community

PerryLink
PerryLink

Posted on Originally published at github.com

Tiny model. Big decisions. — How I built a 144.3M-parameter typed decision model that routes 55.0% of agent decisions off LLMs

版本归属(2026-10-09):本文数字为 v1.1 / 2026-10-09 刷新版 口径:
τ=0.6 时 LLM 调用 −55.0%(45.0% 升级),τ=0.5 档为 −79.6%;
保留集准确率 0.906 → 0.9936。本文 URL 与早期分享卡片中的 "82%" 为历史读数,对应 τ=0.5 档。
以 仓库 BENCHMARKS.md 为准。


Tiny model. Big decisions. — How I built a 144.3M-parameter typed decision model that routes 55.0% of agent decisions off LLMs

Or: why your agent's "should I run this command?" does not need a 70B chat model.


Every agentic workflow I build hits the same wall: the interesting logic is 10 lines, and the other 90% is a language model deciding "allow or deny?", "which tool?", "pass or escalate?" — thousands of times a day, at chat-model prices and chat-model latency.

So I built the opposite of a chatbot: Phocinae-Largha-150M-v1, a 144.3M-parameter typed decision model. It cannot generate text. It takes a state plus a list of typed questions (yes/no, pick-one, 2-10 score) and returns, in one forward pass, a verdict per question with calibrated confidence. GPU: 21.0 ms p50 per decision (RTX 5090). CPU-only: ~1.64 s per case, no GPU at all. Open-source, Apache-2.0.

Latency race: local model finishes at 21.0ms (RTX 5090) while the API request is still in flight

The idea: decisions are not text

A decision is a typed output over a closed option set:

state: "agent wants to run: rm -rf /var/log/app"
question: { type: noul, qid: allow, options: [false, true] }
answer:  { allow: { label: false, prob: 0.96, confidence: 0.96 } }
Enter fullscreen mode Exit fullscreen mode

There is no sentence to generate, no chain-of-thought to emit, no format to parse back. So why rent a generative model for it? A 150M-class encoder (mmBERT-small base, 256k vocab, Gemma tokenizer) does one forward pass and outputs label logits per question — that is the entire inference. Deterministic: same input, same output. No sampling, no parsing failures.

The evaluation protocol is typed-decisions (the format used by Laya and others), which makes scores directly comparable across models on the same rows.

What it measures up to

English typed-decisions: 0.906 (above the 0.735 teacher self-agreement reference — the card flags scores far above it as label-specific overfitting, so we report it as a label-agreement reading, with JevBench 0.5455 (126/231) as the generalization boundary; 400 cases / 2,000 decisions). Chinese (machine-translated eval set): 0.848. Same-protocol published scores: Laya 0.766 (self-measured, native) · TypeSafe JEV-27B 0.727 · meraGPT 0.768.

The metrics most model cards skip, we publish:

  • Option-order flip rate: shuffle the options, does the answer move? reversed 0.0217 / random-mean 0.0144 / any-of-3 0.0283. Roughly one changed answer per ~46 reorders.
  • Calibration: shipped-column ECE 0.2519 (0.0168 with the bundled calib/ column) — disclosed as-is, not hidden.
  • JevBench public-231: 0.5455 (126/231) — below the 58.4% acceptance gate, published anyway. Never trained on eval rows.

The economics: an escalate gate, not a replacement

A small model does not need to be perfect — it needs to know when it is not. With a τ=0.6 confidence gate, confident decisions stay local and the rest escalate to a bigger model. Result on the en route: LLM calls cut 55.0% at τ=0.6 (100% → 45.0%; 79.6% at τ=0.5), while kept-subset accuracy went 0.906 → 0.9936 — routing the hard 45.0% upward made the kept local decisions slightly better, not just cheaper.

Decision ledger: 54 green local cells, 46 grey escalate cells

That is the framing I want to leave with you: System 1 in BERT. The two-system picture for agents is not "small model vs big model" — it is typed, deterministic, milliseconds, free for the repetitive 55.0%, and generative, expensive only for the ambiguous 45.0%.

Using it (three commands)

pip install phocinae-server huggingface_hub
huggingface-cli download Phocinae/Phocinae-Largha-150M-v1 --local-dir ./model
PHOC_MODEL_DIR=./model python -m phocinae.main
Enter fullscreen mode Exit fullscreen mode
curl -s http://127.0.0.1:8155/v1/systemone \
  -H 'Content-Type: application/json' \
  -d '{"state":"The agent wants to run: rm -rf /var/log/app",
       "questions":[{"type":"noul","qid":"allow",
                     "question":"Allow this command?","options":["false","true"]}]}'
# → {"allow": {"label": "false", "prob": 0.96, "confidence": 0.96}}
Enter fullscreen mode Exit fullscreen mode

The server is local-only (127.0.0.1), pure PyTorch at runtime, and the protocol (/v1/systemone: noul / choice / score questions, calibrated probabilities) is fully specified in the repo. There is also a DeepSeek Harness bundle (dsh-phocinae, npm) with a PreToolUse approval gate that fails closed.

Honest limitations (the section everyone skips — please don't)

  • It is not a chatbot, generator, or long-document reasoner. World-knowledge QA is not the job.
  • Chinese rows are machine-translated English cases.
  • Long inputs degrade: 16k/32k probes score 0.453 / 0.387.
  • JevBench gate not passed (that is why the number is on the card).
  • No demographic/fairness evaluation yet.

Try, verify, reproduce

  • Repo (Apache-2.0, full docs + 27 scenario demos): Phocinae/Phocinae-Largha-150M-v1
  • Weights: Hugging Face · ModelScope
  • Reproduction: seeds, row-set hashes, and the environment file ship in the repo; the frozen eval-harness scripts (typed scoring, τ sweep, latency bench) follow in a packaged release. Row sets and protocols are public today, so every published number is checkable.

If you have a use case that is mostly repetitive decisions, I'd love to hear what breaks first.

Top comments (0)