<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: gj0xv</title>
    <description>The latest articles on DEV Community by gj0xv (@gjusev).</description>
    <link>https://dev.to/gjusev</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4152041%2F96d00e75-5c00-44f2-bfca-52e87bd888ca.jpg</url>
      <title>DEV Community: gj0xv</title>
      <link>https://dev.to/gjusev</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/gjusev"/>
    <language>en</language>
    <item>
      <title>AUC 0.953 at $0 per 1,000 emails: explainable phishing detection that runs locally (open source and measured)</title>
      <dc:creator>gj0xv</dc:creator>
      <pubDate>Thu, 01 Oct 2026 11:10:01 +0000</pubDate>
      <link>https://dev.to/gjusev/auc-0953-at-0-per-1000-emails-explainable-phishing-detection-that-runs-locally-145o</link>
      <guid>https://dev.to/gjusev/auc-0953-at-0-per-1000-emails-explainable-phishing-detection-that-runs-locally-145o</guid>
      <description>&lt;p&gt;When a security analyst flags an email as phishing, the first question is always "why". A black-box classifier that says "97% phishing" does not answer it. You need the reasons, the specific signals, the things a human can verify.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/Gjusev/laya-phishield" rel="noopener noreferrer"&gt;laya-phishield&lt;/a&gt; decomposes the phishing decision into eight narrow yes/no signals, checks deterministic email headers first, and combines everything into a risk score with visible feature contributions. It runs on your machine, for free.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AUC: 0.953 (composite)
Recall at 0.5 threshold: 0.808
Cost per 1,000 emails: $0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  How it decides
&lt;/h2&gt;

&lt;p&gt;Two layers, one pass:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layer 1: deterministic pre-pass.&lt;/strong&gt; Before any model runs, the tool checks email headers, reply-to mismatches, punycode and IP-literal links, homoglyphs in domains, and SPF/DKIM failures. These are hard signals: no ML needed, no ambiguity.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layer 2: eight yes/no signals.&lt;/strong&gt; The laya decision engine (local, CPU) scores each email on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;asks_credentials&lt;/code&gt; (does it request login or banking details)&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;urgency_pressure&lt;/code&gt; (artificial time pressure)&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;brand_impersonation&lt;/code&gt; (mimics a known brand)&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;external_link_risk&lt;/code&gt; (suspicious outbound links)&lt;/li&gt;
&lt;li&gt;and four more&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A logistic regression head combines them into a final score. The output is not "phishing: yes". It is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;risk_score&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;0.87&lt;/span&gt;
&lt;span class="na"&gt;top_contributors&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;brand_impersonation&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;+0.34&lt;/span&gt;
  &lt;span class="na"&gt;urgency_pressure&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;+0.21&lt;/span&gt;
  &lt;span class="na"&gt;external_link_risk&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;+0.18&lt;/span&gt;
  &lt;span class="na"&gt;asks_credentials&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;+0.14&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;An analyst can verify each of those in seconds. That is what explainable means in practice.&lt;/p&gt;

&lt;h2&gt;
  
  
  The benchmarks
&lt;/h2&gt;

&lt;p&gt;183 emails (105 legitimate, 78 phishing), temporal split, from the Nazario phishing corpus and Enron, deduplicated with MinHash/LSH:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Method&lt;/th&gt;
&lt;th&gt;AUC&lt;/th&gt;
&lt;th&gt;Precision @ 0.5&lt;/th&gt;
&lt;th&gt;Recall @ 0.5&lt;/th&gt;
&lt;th&gt;Cost / 1,000&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Keyword rules&lt;/td&gt;
&lt;td&gt;0.594&lt;/td&gt;
&lt;td&gt;1.000&lt;/td&gt;
&lt;td&gt;0.051&lt;/td&gt;
&lt;td&gt;$0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Forced choice (laya)&lt;/td&gt;
&lt;td&gt;0.944&lt;/td&gt;
&lt;td&gt;0.902&lt;/td&gt;
&lt;td&gt;0.590&lt;/td&gt;
&lt;td&gt;$0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Composite (laya + pre-pass)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.953&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;0.913&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.808&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$0&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The keyword baseline catches almost nothing. The laya model alone is strong. The composite (deterministic checks + ML scoring) is the best, because header failures and homoglyphs are signals a language model should not have to relearn.&lt;/p&gt;

&lt;p&gt;Baselines were generated and judged with Z.ai's &lt;strong&gt;GLM-5.3-flashX&lt;/strong&gt;. Local decisions cost $0; only the benchmark pipeline calls a hosted model.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The honest trade-off:&lt;/strong&gt; at 95% recall, the forced-choice variant has a lower false-positive rate (0.124 vs 0.210 for composite). If you need fewer false alarms and can accept more misses, use forced choice. The README shows the full curve.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;laya-phishield
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Scan an email from the CLI:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;laya-scan suspicious.eml
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Or serve it as an API:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;laya-serve &lt;span class="nt"&gt;--port&lt;/span&gt; 8080
&lt;span class="c"&gt;# POST /scan with raw .eml body&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Streamlit demo included for visual inspection.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Disclaimer:&lt;/strong&gt; this is not a replacement for a production secure email gateway. It is an explainable, local signal layer for analysts and researchers. The README says this too.&lt;/p&gt;

&lt;h2&gt;
  
  
  The rest of the series
&lt;/h2&gt;

&lt;p&gt;This is the last of four open-source tools built on the &lt;a href="https://github.com/NandhaKishorM/laya" rel="noopener noreferrer"&gt;laya decision engine&lt;/a&gt;, the same idea that OpenAI shipped as the Decisions API on Luna and TypeSafe shipped as Jev this month. The difference is where it runs: your machine, your data, $0 per decision.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;laya-router&lt;/strong&gt;: route prompts between cheap and frontier models, 54.9% cost reduction&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;laya-compactor&lt;/strong&gt;: cut 70% of RAG context tokens, same answer quality&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;laya-triage&lt;/strong&gt;: support ticket triage, fine-tuned from 51% to 90.5% intent accuracy&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Every benchmark is committed with the code. Including the numbers that hurt.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>security</category>
      <category>python</category>
      <category>opensource</category>
    </item>
    <item>
      <title>From 51% to 90.5%: fine-tuning a local triage model on 10,003 support tickets (open source and measured)</title>
      <dc:creator>gj0xv</dc:creator>
      <pubDate>Thu, 01 Oct 2026 11:09:49 +0000</pubDate>
      <link>https://dev.to/gjusev/from-51-to-905-fine-tuning-a-local-triage-model-on-10003-support-tickets-14a</link>
      <guid>https://dev.to/gjusev/from-51-to-905-fine-tuning-a-local-triage-model-on-10003-support-tickets-14a</guid>
      <description>&lt;p&gt;When OpenAI launched the Decisions API on Luna this week, and TypeSafe launched Jev two weeks before that, they were both making the same bet: the future of AI in production is not bigger models, it is smaller, faster decision models embedded in your workflow.&lt;/p&gt;

&lt;p&gt;I agree. And I have the numbers to prove it works.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/Gjusev/laya-triage" rel="noopener noreferrer"&gt;laya-triage&lt;/a&gt; routes support tickets through a local decision model: intent (77 categories), department (8), urgency, frustration, churn risk, refund flag, and a confidence-gated escalation to a human. All in one forward pass, on CPU, for free.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Hierarchical intent accuracy: 51.0% -&amp;gt; 90.5% (after fine-tuning)
Coarse-cluster accuracy: 68.0% -&amp;gt; 96.0%
Urgency MAE: 0.81 -&amp;gt; 0.70
Cost per ticket: $0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The architecture
&lt;/h2&gt;

&lt;p&gt;Flat 77-way classification is hard for a small model. So I broke it into two stages:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Ticket -&amp;gt; [coarse stage: 12 clusters] -&amp;gt; [fine stage: 3-10 candidate intents] -&amp;gt; decision
              |
              +-&amp;gt; urgency, frustration, churn, refund (same forward pass)
              |
              +-&amp;gt; confidence check -&amp;gt; escalate to human if below 0.84
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The hierarchy alone added 14.5 percentage points over flat routing. The fine-tune added another 39.5.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I fine-tuned on
&lt;/h2&gt;

&lt;p&gt;10,003 BANKING77 tickets, split into two sequences per ticket, trained on Kaggle's free 2x T4 GPUs. The checkpoint is published on HuggingFace: &lt;a href="https://huggingface.co/Gjusev/laya-triage-banking77" rel="noopener noreferrer"&gt;Gjusev/laya-triage-banking77&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The eval harness that tracked the improvement ran on Z.ai's GLM models, with &lt;strong&gt;GLM-5.3-flashX&lt;/strong&gt; as the workhorse. When the judge model changes, the numbers move a little. The fine-tune gain does not.&lt;/p&gt;

&lt;h2&gt;
  
  
  The results, before and after
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Measurement&lt;/th&gt;
&lt;th&gt;Zero-shot&lt;/th&gt;
&lt;th&gt;Fine-tuned&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Hierarchical intent accuracy&lt;/td&gt;
&lt;td&gt;51.0%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;90.5%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Macro-F1&lt;/td&gt;
&lt;td&gt;not measured&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.848&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Coarse-cluster accuracy&lt;/td&gt;
&lt;td&gt;68.0%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;96.0%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Urgency MAE&lt;/td&gt;
&lt;td&gt;0.81&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.70&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Frustration MAE&lt;/td&gt;
&lt;td&gt;1.07&lt;/td&gt;
&lt;td&gt;1.07 (untouched heads)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;At the 0.84 confidence threshold (zero-shot): 75.2% accuracy with 60.5% of tickets auto-handled. The rest escalate to a human with a recorded reason.&lt;/p&gt;

&lt;h2&gt;
  
  
  The escalation policy (and why it matters)
&lt;/h2&gt;

&lt;p&gt;The threshold gates only on cluster and intent confidence. Signal confidence (urgency, frustration, churn) is deliberately not used, because score-type questions have structurally lower confidence and would trigger near-universal handoffs.&lt;/p&gt;

&lt;p&gt;On an 8-ticket out-of-domain smoke test (IT-ops tickets against banking intents): 7 escalated correctly. The one that did not, an SSO lockout, mapped to "unable to verify identity". Not wrong, exactly, but the reason it matters is that the behavior is visible and fixable, not hidden behind an opaque binary.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;

&lt;p&gt;The demo is live on Streamlit, in 6 languages (English, Spanish, French, German, Hindi, Arabic):&lt;/p&gt;

&lt;p&gt;&lt;a href="https://laya-triage.streamlit.app" rel="noopener noreferrer"&gt;&lt;strong&gt;Live demo&lt;/strong&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Or run it yourself:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;laya-triage
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Fine-tuning notebook included for Kaggle (2x T4, ~4 hours).&lt;/p&gt;

&lt;h2&gt;
  
  
  The rest of the series
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;laya-router&lt;/strong&gt;: route prompts between cheap and frontier models, 54.9% cost reduction&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;laya-compactor&lt;/strong&gt;: cut 70% of RAG context tokens, same answer quality&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;laya-phishield&lt;/strong&gt;: explainable phishing detection, $0 per 1,000 emails&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;All four tools use the same &lt;a href="https://github.com/NandhaKishorM/laya" rel="noopener noreferrer"&gt;laya decision engine&lt;/a&gt;, Apache 2.0, local, with published benchmarks. If OpenAI's Decisions API and TypeSafe's Jev are the proprietary versions of this idea, this series is the open one. Run it on your machine, keep your data, pay nothing per decision.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>machinelearning</category>
      <category>python</category>
    </item>
    <item>
      <title>Cutting 70% of RAG context tokens and keeping the answers identical (open source and measured)</title>
      <dc:creator>gj0xv</dc:creator>
      <pubDate>Thu, 01 Oct 2026 09:30:00 +0000</pubDate>
      <link>https://dev.to/gjusev/cutting-70-of-rag-context-tokens-and-keeping-the-answers-identical-measured-5bdd</link>
      <guid>https://dev.to/gjusev/cutting-70-of-rag-context-tokens-and-keeping-the-answers-identical-measured-5bdd</guid>
      <description>&lt;p&gt;Your RAG pipeline retrieves 12 chunks because the retrieval score said "maybe". Your LLM reads all of them. You pay for all of them. And the answer quality was decided by chunks 2 and 7 anyway.&lt;/p&gt;

&lt;p&gt;On September 29, OpenAI launched the Decisions API built on Luna, and on September 15, TypeSafe launched Jev. Both are fast decision models for exactly this kind of problem: deciding what matters before the expensive model runs. But they are APIs. Your documents leave your machine, and you pay per decision.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/Gjusev/laya-compactor" rel="noopener noreferrer"&gt;laya-compactor&lt;/a&gt; does the same thing locally, for free. It scores your entire retrieval batch in a single forward pass and decides which documents stay in context and which get cut.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;70% fewer input tokens
Zero answer-quality loss on SQuAD (exact match: 0.345 vs 0.345)
94.5% of gold documents retained on HotpotQA
$0 per compaction decision
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  How it works
&lt;/h2&gt;

&lt;p&gt;One principle drives the design: &lt;strong&gt;delete, do not rewrite.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The tool never summarizes or paraphrases your documents. It keeps the ones that matter verbatim and removes the rest, with an auditable reason (low score or exhausted budget). The evidence stays intact for the LLM to reason over.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;+--------------------------------------------------+
|  Retrieval batch (12 docs, ~3,200 tokens)         |
|                                                   |
|  laya decision engine (one forward pass, local)   |
|  scores each doc: 0=irrelevant .. 3=essential     |
|                                                   |
|  Budget: keep top-scoring docs until limit        |
|                                                   |
|  Output: 4 docs, ~970 tokens, same answer         |
+--------------------------------------------------+
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Python API and CLI:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;laya_compactor&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;compact&lt;/span&gt;

&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;compact&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;What is the refund policy for annual plans?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;documents&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;retrieved_chunks&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;token_budget&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="c1"&gt;# result.kept: the documents that survived
# result.dropped: [{doc, score, reason}] for the ones cut
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The benchmarks (all of them)
&lt;/h2&gt;

&lt;p&gt;200 questions per dataset, BM25 retrieval, generator and blind judge powered by Z.ai's &lt;strong&gt;GLM-5.3-flashX&lt;/strong&gt;:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dataset&lt;/th&gt;
&lt;th&gt;Full context&lt;/th&gt;
&lt;th&gt;Compacted&lt;/th&gt;
&lt;th&gt;Tokens saved&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;SQuAD&lt;/td&gt;
&lt;td&gt;EM 0.345, 3,214 avg tokens&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;EM 0.345&lt;/strong&gt;, 973 avg tokens&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;69.7%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HotpotQA&lt;/td&gt;
&lt;td&gt;EM 0.230, 3,193 avg tokens&lt;/td&gt;
&lt;td&gt;EM 0.200, 947 avg tokens&lt;/td&gt;
&lt;td&gt;70.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;On SQuAD the exact match is identical: the compactor removed noise without touching signal. On HotpotQA it dropped 3 points but kept 94.5% of the gold documents, meaning the loss came from multi-hop reasoning, not from cutting the wrong chunks. Truncation baselines (head-only, tail-only) scored worse on both.&lt;/p&gt;

&lt;p&gt;Caveats, because they exist: CPU latency is significant (6.3 to 10 seconds p50 for a batch). Multi-hop questions remain hard. The rubric sensitivity study is committed with a hand-labeled 100-row dataset so you can see exactly where the scoring is fragile.&lt;/p&gt;

&lt;h2&gt;
  
  
  Integrations
&lt;/h2&gt;

&lt;p&gt;Drop-in for the two frameworks you already use:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# LangChain
&lt;/span&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;langchain.retrievers&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;ContextualCompressionRetriever&lt;/span&gt;
&lt;span class="n"&gt;retriever&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;ContextualCompressionRetriever&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;base_compressor&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nc"&gt;LayaCompactor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;token_budget&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="n"&gt;base_retriever&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;your_retriever&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# LlamaIndex
&lt;/span&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;laya_compactor.integrations&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;LayaNodePostprocessor&lt;/span&gt;
&lt;span class="n"&gt;query_engine&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;index&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;as_query_engine&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;node_postprocessors&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nc"&gt;LayaNodePostprocessor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;token_budget&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="p"&gt;)],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;laya-compactor
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;CLI included: &lt;code&gt;laya-compact --question "..." --docs docs/ --budget 1000&lt;/code&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The rest of the series
&lt;/h2&gt;

&lt;p&gt;This is the second of four tools built on the same idea: local, measured decisions in LLM pipelines.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;laya-router&lt;/strong&gt;: route prompts between cheap and frontier models, 54.9% cost reduction&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;laya-triage&lt;/strong&gt;: support ticket triage, fine-tuned from 51% to 90.5% intent accuracy&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;laya-phishield&lt;/strong&gt;: explainable phishing detection, $0 per 1,000 emails&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each one ships with its benchmarks in the README, including the numbers that hurt.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>rag</category>
      <category>python</category>
    </item>
    <item>
      <title>How I cut my LLM bill by 54.9%: a local router that costs $0 per decision (open source and measured)</title>
      <dc:creator>gj0xv</dc:creator>
      <pubDate>Thu, 01 Oct 2026 06:30:00 +0000</pubDate>
      <link>https://dev.to/gjusev/how-i-cut-my-llm-bill-by-549-a-local-router-that-costs-0-per-decision-1fmk</link>
      <guid>https://dev.to/gjusev/how-i-cut-my-llm-bill-by-549-a-local-router-that-costs-0-per-decision-1fmk</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9rqqozcncg7eo2w1iumc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9rqqozcncg7eo2w1iumc.png" alt=" " width="800" height="420"&gt;&lt;/a&gt;&lt;br&gt;
This week, two of the biggest names in AI shipped the same idea within days of each other. On September 15, TypeSafe launched Jev, the first "System One model": an AI that outputs typed decisions instead of text. On September 29, OpenAI answered at DevDay with the Decisions API built on Luna, their fast, cheap model for real-time decision-making.&lt;/p&gt;

&lt;p&gt;The premise is right. The pricing usually isn't.&lt;/p&gt;

&lt;p&gt;That is where open source matters. When the decision layer runs on your machine, with your data, at $0 per call, the economics flip: every prompt gets the right-sized model instead of the one that maximises someone else's margin. You can audit it, fork it, and turn it off when it does something stupid. No vendor slide can promise you that.&lt;/p&gt;

&lt;p&gt;So here is mine: &lt;a href="https://github.com/Gjusev/laya-router" rel="noopener noreferrer"&gt;laya-router&lt;/a&gt;, the first of four open-source tools that put local, measured decisions between your app and your LLMs, built on the &lt;a href="https://github.com/NandhaKishorM/laya" rel="noopener noreferrer"&gt;laya decision engine&lt;/a&gt;. Benchmarked with Z.ai's new &lt;strong&gt;GLM-5.3-flashX&lt;/strong&gt; as the cheap tier, which is itself a fitting story: a fast, cheap model deciding which prompts it can handle itself.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;54.9% estimated cost reduction
80.6% of prompts routed to the cheap tier
$0 per routing decision
460 ms warm latency (p50)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The idea
&lt;/h2&gt;

&lt;p&gt;You keep your code exactly as it is. You only change the &lt;code&gt;base_url&lt;/code&gt;. The proxy ignores the model you asked for and picks one itself:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;http://localhost:8000/v1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="c1"&gt;# from here on, everything is standard
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three layers decide:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Regex fast path&lt;/strong&gt;: trivial prompts ("hi", "translate: merci") skip the model entirely&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A local decision model&lt;/strong&gt; (the &lt;a href="https://github.com/NandhaKishorM/laya" rel="noopener noreferrer"&gt;laya&lt;/a&gt; engine, Apache 2.0, runs on CPU) classifies the prompt as simple, standard or complex, with calibrated confidence&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A confidence gate&lt;/strong&gt;: uncertain prompts escalate to the frontier tier. The system is allowed to say "I don't know" and pay more&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Every response carries inspection headers, so you can audit what happened and why:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;X-Laya-Route: cheap
X-Laya-Model: glm-4.6-flash
X-Laya-Confidence: 0.91
X-Laya-Reason: classified simple, above threshold
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The benchmarks (all of them)
&lt;/h2&gt;

&lt;p&gt;180 prompts, one cheap and one frontier model, blind LLM judge. I publish the wins and the caveats:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Measurement&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Routed to cheap tier&lt;/td&gt;
&lt;td&gt;80.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Estimated cost reduction&lt;/td&gt;
&lt;td&gt;54.9%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cheap win/tie rate vs frontier&lt;/td&gt;
&lt;td&gt;79.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Routing latency p50 / p95 / p99&lt;/td&gt;
&lt;td&gt;460 ms / 1.4 s / 2.7 s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Caveats, because they exist: one judge, one model pair, and the judge was noisy on trivial prompts. The backtest pipeline is committed (&lt;code&gt;make backtest&lt;/code&gt;) with the dataset, the answers and the judgements, so you can re-run it on your own traffic. If your prompts are all hard, you will save nothing. Measure first.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it deliberately does not do
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;No silent failover.&lt;/strong&gt; If routing fails, you get a structured 503, not a surprise GPT-4 call on your bill&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No rewriting.&lt;/strong&gt; Your prompt arrives at the upstream model byte-identical. The proxy decides, it does not mutate&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No vendor lock.&lt;/strong&gt; Works with OpenAI, vLLM, Ollama, OpenRouter, Z.ai, anything OpenAI-shaped&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;laya-router
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Define your tiers in YAML (models + prices), point your &lt;code&gt;base_url&lt;/code&gt; at it, done. Docker image included, &lt;code&gt;/metrics&lt;/code&gt; endpoint for Prometheus, JSONL decision log if you want to analyze your routing later.&lt;/p&gt;

&lt;h2&gt;
  
  
  The rest of the series
&lt;/h2&gt;

&lt;p&gt;laya-router is one of four tools I'm building around the same idea: local, measured decisions in LLM pipelines.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;laya-compactor&lt;/strong&gt;: cut 70% of RAG context tokens with zero answer-quality loss (measured)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;laya-triage&lt;/strong&gt;: support ticket triage, fine-tuned from 51% to 90.5% intent accuracy&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;laya-phishield&lt;/strong&gt;: explainable phishing detection, $0 per 1,000 emails&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each one ships with its benchmarks in the README, including the numbers that hurt.&lt;/p&gt;

&lt;p&gt;If your LLM app ever died in production after a perfect demo, we have something to talk about :)&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>python</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Judge cheap, audit confidence: a CI gate for LLM evals (open source and measured)</title>
      <dc:creator>gj0xv</dc:creator>
      <pubDate>Thu, 01 Oct 2026 06:00:00 +0000</pubDate>
      <link>https://dev.to/gjusev/judge-cheap-audit-confidence-a-ci-gate-for-llm-evals-open-source-269l</link>
      <guid>https://dev.to/gjusev/judge-cheap-audit-confidence-a-ci-gate-for-llm-evals-open-source-269l</guid>
      <description>&lt;p&gt;Your LLM passes the demo every time. Then a customer sends a prompt you didn't test, and the answer is garbage. The gap between "it worked in the eval" and "it works in production" is almost always the same: you measured accuracy, but you never measured whether the model's confidence was trustworthy.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/Gjusev/laya-evals" rel="noopener noreferrer"&gt;laya-evals&lt;/a&gt; closes that gap. It is the fifth tool in my open-source series on the laya decision engine, and the one that keeps the other four honest.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="err"&gt;Judge&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;accuracy:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.8968&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;(Cohen's&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;kappa&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.7938&lt;/span&gt;&lt;span class="err"&gt;)&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;Calibration:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;ECE&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.0253&lt;/span&gt;&lt;span class="err"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;Brier&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.0779&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;Cost&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;per&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="err"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;000&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;judgments:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;$&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;(local&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;inference)&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;CI&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;gate:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;fails&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;the&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;build&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;when&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;accuracy&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;or&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;calibration&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;regresses&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  What it does
&lt;/h2&gt;

&lt;p&gt;Three things, in one pipeline:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Rubric-based LLM-as-judge.&lt;/strong&gt; Score your eval sets with a local laya decision model instead of paying for GPT-4 as a judge. Choice questions, score questions, or no-ground-truth mode. 6.4 decisions per second on CPU.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Calibration audit.&lt;/strong&gt; Accuracy tells you how often the model is right. Calibration tells you whether you can trust its confidence. laya-evals computes Expected Calibration Error, Brier score, and reliability bins, then tells you which confidence threshold to use for each question shape.&lt;/p&gt;

&lt;p&gt;Why per shape? Because a 2-option question needs a different gate than a 20-option question. The SST-2 benchmark (2 options) needs a threshold of 0.865 for 95% accuracy. The MASSIVE intent benchmark (20 options) needs 0.994. One global &lt;code&gt;min_confidence&lt;/code&gt; is not a policy, and laya-evals proves it with your own numbers.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. CI regression gate.&lt;/strong&gt; Wire it into GitHub Actions. Exit code 0 means pass, exit code 1 means a measured regression, exit code 2 means the run itself was invalid. Your eval suite becomes a build gate, not a suggestion.&lt;/p&gt;

&lt;h2&gt;
  
  
  The numbers
&lt;/h2&gt;

&lt;p&gt;872 SST-2 sentences, 2-option rubric, local laya judge:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Measurement&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Judge accuracy vs gold&lt;/td&gt;
&lt;td&gt;0.8968&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cohen's kappa&lt;/td&gt;
&lt;td&gt;0.7938&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ECE (15 bins)&lt;/td&gt;
&lt;td&gt;0.0253&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Brier score&lt;/td&gt;
&lt;td&gt;0.0779&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Throughput&lt;/td&gt;
&lt;td&gt;6.4 decisions/s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;API cost&lt;/td&gt;
&lt;td&gt;$0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Cross-checked against Z.ai's GLM-5.3-flash as reference judge: 95.1% accuracy, 90.7% agreement, kappa 0.814. The local judge is cheaper and nearly as reliable.&lt;/p&gt;

&lt;p&gt;I also re-measured four public laya benchmark claims with laya-evals and got values within 0.0003 of the published numbers. The reproduction pack is committed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;laya-evals
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;laya_evals&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Judge&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ece&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;brier_score&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;advise_thresholds&lt;/span&gt;

&lt;span class="n"&gt;judge&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Judge&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;results&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;judge&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;eval_set&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;rubric&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;ece&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;results&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;confidences&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;results&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;correct&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;advise_thresholds&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;results&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The series
&lt;/h2&gt;

&lt;p&gt;This is the fifth and last tool in the laya series, built on the &lt;a href="https://github.com/NandhaKishorM/laya" rel="noopener noreferrer"&gt;laya decision engine&lt;/a&gt; by NandhaKishorM. The same engine that TypeSafe shipped as Jev and OpenAI shipped as the Decisions API on Luna this week. The difference: this one runs on your machine, for free, and the benchmarks are committed.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;laya-router&lt;/strong&gt;: 54.9% cost reduction, $0 per routing decision&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;laya-compactor&lt;/strong&gt;: 70% fewer RAG tokens, same answer quality&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;laya-triage&lt;/strong&gt;: support triage from 51% to 90.5% after fine-tuning&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;laya-phishield&lt;/strong&gt;: explainable phishing detection, $0 per 1,000 emails&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Every tool ships with its benchmarks, including the numbers that hurt.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>testing</category>
      <category>python</category>
    </item>
  </channel>
</rss>
