<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: PerryLink</title>
    <description>The latest articles on DEV Community by PerryLink (@perrylink).</description>
    <link>https://dev.to/perrylink</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4169985%2F46e80c90-b59d-4d04-94eb-7fd34b4c3945.png</url>
      <title>DEV Community: PerryLink</title>
      <link>https://dev.to/perrylink</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/perrylink"/>
    <language>en</language>
    <item>
      <title>Tiny model. Big decisions. — How I built a 144.3M-parameter typed decision model that routes 55.0% of agent decisions off LLMs</title>
      <dc:creator>PerryLink</dc:creator>
      <pubDate>Thu, 08 Oct 2026 03:41:32 +0000</pubDate>
      <link>https://dev.to/perrylink/tiny-model-big-decisions-how-i-built-a-144m-parameter-typed-decision-model-that-routes-82-of-23fi</link>
      <guid>https://dev.to/perrylink/tiny-model-big-decisions-how-i-built-a-144m-parameter-typed-decision-model-that-routes-82-of-23fi</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;版本归属（2026-10-09）&lt;/strong&gt;：本文数字为 &lt;strong&gt;v1.1 / 2026-10-09 刷新版&lt;/strong&gt; 口径：&lt;br&gt;
&lt;strong&gt;τ=0.6 时 LLM 调用 −55.0%&lt;/strong&gt;（45.0% 升级），&lt;strong&gt;τ=0.5 档为 −79.6%&lt;/strong&gt;；&lt;br&gt;
保留集准确率 0.906 → 0.9936。本文 URL 与早期分享卡片中的 "82%" 为历史读数，对应 τ=0.5 档。&lt;br&gt;
以 &lt;a href="https://github.com/Phocinae/Phocinae-Largha-150M-v1/blob/main/BENCHMARKS.md" rel="noopener noreferrer"&gt;仓库 BENCHMARKS.md&lt;/a&gt; 为准。&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  Tiny model. Big decisions. — How I built a 144.3M-parameter typed decision model that routes 55.0% of agent decisions off LLMs
&lt;/h1&gt;

&lt;p&gt;&lt;em&gt;Or: why your agent's "should I run this command?" does not need a 70B chat model.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;Every agentic workflow I build hits the same wall: the interesting logic is 10 lines, and the other 90% is a language model deciding &lt;em&gt;"allow or deny?", "which tool?", "pass or escalate?"&lt;/em&gt; — thousands of times a day, at chat-model prices and chat-model latency.&lt;/p&gt;

&lt;p&gt;So I built the opposite of a chatbot: &lt;strong&gt;Phocinae-Largha-150M-v1&lt;/strong&gt;, a 144.3M-parameter typed decision model. It cannot generate text. It takes a state plus a list of typed questions (yes/no, pick-one, 2-10 score) and returns, in one forward pass, a verdict per question with calibrated confidence. GPU: &lt;strong&gt;21.0 ms&lt;/strong&gt; p50 per decision (RTX 5090). CPU-only: ~1.64 s per case, no GPU at all. Open-source, Apache-2.0.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Funnk1hu0jnxfvy8z3lr0.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Funnk1hu0jnxfvy8z3lr0.gif" alt="Latency race: local model finishes at 21.0ms (RTX 5090) while the API request is still in flight" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The idea: decisions are not text
&lt;/h2&gt;

&lt;p&gt;A decision is a &lt;strong&gt;typed output over a closed option set&lt;/strong&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="err"&gt;state:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"agent wants to run: rm -rf /var/log/app"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;question:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;type:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;noul&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;qid:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;allow&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;options:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;answer:&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;allow:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;label:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;prob:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.96&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;confidence:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.96&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There is no sentence to generate, no chain-of-thought to emit, no format to parse back. So why rent a generative model for it? A 150M-class encoder (mmBERT-small base, 256k vocab, Gemma tokenizer) does one forward pass and outputs label logits per question — that is the entire inference. Deterministic: same input, same output. No sampling, no parsing failures.&lt;/p&gt;

&lt;p&gt;The evaluation protocol is typed-decisions (the format used by Laya and others), which makes scores directly comparable across models on the same rows.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it measures up to
&lt;/h2&gt;

&lt;p&gt;English typed-decisions: &lt;strong&gt;0.906&lt;/strong&gt; (above the 0.735 teacher self-agreement reference — the card flags scores far above it as label-specific overfitting, so we report it as a label-agreement reading, with JevBench 0.5455 (126/231) as the generalization boundary; 400 cases / 2,000 decisions). Chinese (machine-translated eval set): &lt;strong&gt;0.848&lt;/strong&gt;. Same-protocol published scores: Laya 0.766 (self-measured, native) · TypeSafe JEV-27B 0.727 · meraGPT 0.768.&lt;/p&gt;

&lt;p&gt;The metrics most model cards skip, we publish:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Option-order flip rate&lt;/strong&gt;: shuffle the options, does the answer move? reversed 0.0217 / random-mean 0.0144 / any-of-3 0.0283. Roughly one changed answer per ~46 reorders.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Calibration&lt;/strong&gt;: shipped-column ECE 0.2519 (0.0168 with the bundled &lt;code&gt;calib/&lt;/code&gt; column) — disclosed as-is, not hidden.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;JevBench public-231: 0.5455 (126/231)&lt;/strong&gt; — below the 58.4% acceptance gate, published anyway. Never trained on eval rows.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The economics: an escalate gate, not a replacement
&lt;/h2&gt;

&lt;p&gt;A small model does not need to be perfect — it needs to know when it is not. With a τ=0.6 confidence gate, confident decisions stay local and the rest escalate to a bigger model. Result on the en route: LLM calls &lt;strong&gt;cut 55.0% at τ=0.6&lt;/strong&gt; (100% → 45.0%; 79.6% at τ=0.5), while kept-subset accuracy went &lt;strong&gt;0.906 → 0.9936&lt;/strong&gt; — routing the hard 45.0% upward made the kept local decisions slightly &lt;em&gt;better&lt;/em&gt;, not just cheaper.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbs5lh4blf6e7vubr3o2a.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbs5lh4blf6e7vubr3o2a.gif" alt="Decision ledger: 54 green local cells, 46 grey escalate cells" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That is the framing I want to leave with you: &lt;strong&gt;System 1 in BERT&lt;/strong&gt;. The two-system picture for agents is not "small model vs big model" — it is &lt;em&gt;typed, deterministic, milliseconds, free&lt;/em&gt; for the repetitive 55.0%, and &lt;em&gt;generative, expensive&lt;/em&gt; only for the ambiguous 45.0%.&lt;/p&gt;

&lt;h2&gt;
  
  
  Using it (three commands)
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;phocinae-server huggingface_hub
huggingface-cli download Phocinae/Phocinae-Largha-150M-v1 &lt;span class="nt"&gt;--local-dir&lt;/span&gt; ./model
&lt;span class="nv"&gt;PHOC_MODEL_DIR&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;./model python &lt;span class="nt"&gt;-m&lt;/span&gt; phocinae.main
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; http://127.0.0.1:8155/v1/systemone &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'Content-Type: application/json'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"state":"The agent wants to run: rm -rf /var/log/app",
       "questions":[{"type":"noul","qid":"allow",
                     "question":"Allow this command?","options":["false","true"]}]}'&lt;/span&gt;
&lt;span class="c"&gt;# → {"allow": {"label": "false", "prob": 0.96, "confidence": 0.96}}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The server is local-only (127.0.0.1), pure PyTorch at runtime, and the protocol (&lt;code&gt;/v1/systemone&lt;/code&gt;: noul / choice / score questions, calibrated probabilities) is fully specified in the repo. There is also a DeepSeek Harness bundle (dsh-phocinae, npm) with a PreToolUse approval gate that fails closed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Honest limitations (the section everyone skips — please don't)
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;It is &lt;strong&gt;not&lt;/strong&gt; a chatbot, generator, or long-document reasoner. World-knowledge QA is not the job.&lt;/li&gt;
&lt;li&gt;Chinese rows are machine-translated English cases.&lt;/li&gt;
&lt;li&gt;Long inputs degrade: 16k/32k probes score 0.453 / 0.387.&lt;/li&gt;
&lt;li&gt;JevBench gate not passed (that is why the number is on the card).&lt;/li&gt;
&lt;li&gt;No demographic/fairness evaluation yet.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Try, verify, reproduce
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Repo (Apache-2.0, full docs + 27 scenario demos): &lt;a href="https://github.com/Phocinae/Phocinae-Largha-150M-v1" rel="noopener noreferrer"&gt;Phocinae/Phocinae-Largha-150M-v1&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Weights: &lt;a href="https://huggingface.co/Phocinae/Phocinae-Largha-150M-v1" rel="noopener noreferrer"&gt;Hugging Face&lt;/a&gt; · &lt;a href="https://modelscope.cn/models/PerryLink/Phocinae-Largha-150M-v1" rel="noopener noreferrer"&gt;ModelScope&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Reproduction: seeds, row-set hashes, and the environment file ship in the repo; the frozen eval-harness scripts (typed scoring, τ sweep, latency bench) follow in a packaged release. Row sets and protocols are public today, so every published number is checkable.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you have a use case that is mostly repetitive decisions, I'd love to hear what breaks first.&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>ai</category>
      <category>python</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
