<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Prince </title>
    <description>The latest articles on DEV Community by Prince  (@pkprajapati7402).</description>
    <link>https://dev.to/pkprajapati7402</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4153502%2Fc453270d-a69c-4e46-82ed-6005aefe9f5d.png</url>
      <title>DEV Community: Prince </title>
      <link>https://dev.to/pkprajapati7402</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/pkprajapati7402"/>
    <language>en</language>
    <item>
      <title>OpenTriage: A Read-Only Triage Assistant for Maintainers, and Why a No-LLM Baseline Beat Gemma on Precision</title>
      <dc:creator>Prince </dc:creator>
      <pubDate>Fri, 02 Oct 2026 08:31:53 +0000</pubDate>
      <link>https://dev.to/pkprajapati7402/opentriage-a-read-only-triage-assistant-for-maintainers-and-why-a-no-llm-baseline-beat-gemma-on-2e25</link>
      <guid>https://dev.to/pkprajapati7402/opentriage-a-read-only-triage-assistant-for-maintainers-and-why-a-no-llm-baseline-beat-gemma-on-2e25</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/hacktoberfest-weekend-2026-10-01"&gt;Hacktoberfest Weekend Challenge: Build for a Friend&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;OpenTriage&lt;/strong&gt; is a local-first triage assistant for open-source maintainers. Point it at a public GitHub repo and, for every issue and PR, it suggests:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;🏷️ &lt;strong&gt;Labels&lt;/strong&gt;, chosen only from the repo's existing labels, with a one-line reason (it abstains when unsure)&lt;/li&gt;
&lt;li&gt;🔁 &lt;strong&gt;Possible duplicates&lt;/strong&gt;: the top 3 similar past issues with similarity scores&lt;/li&gt;
&lt;li&gt;🚩 &lt;strong&gt;Low-effort PR signals&lt;/strong&gt;: transparent rules like an empty description, a whitespace-only diff, or a trivial README edit&lt;/li&gt;
&lt;li&gt;✉️ &lt;strong&gt;Draft first replies&lt;/strong&gt; that politely ask for missing details, for the maintainer to edit&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It is strictly &lt;strong&gt;read-only&lt;/strong&gt;. It never labels, comments, or closes anything on GitHub. A human always decides.&lt;/p&gt;

&lt;h3&gt;
  
  
  Who I built it for
&lt;/h3&gt;

&lt;p&gt;I built OpenTriage for &lt;strong&gt;open-source maintainers and contributors&lt;/strong&gt; who spend significant time reviewing issues and pull requests, especially in repositories with active communities and growing backlogs.&lt;/p&gt;

&lt;p&gt;Every October the backlog gets worse for maintainers everywhere: duplicate reports, issues with no reproduction steps, and PRs that change a single space in a README. The first pass is repetitive, and it eats weekends.&lt;/p&gt;

&lt;p&gt;To go beyond one repo, I also tested OpenTriage against the real history of public repos, where maintainers' actual labeling and duplicate decisions are visible, and I report where it fails below.&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;🌐 &lt;strong&gt;Live demo:&lt;/strong&gt; &lt;a href="https://opentriage-0xzg.onrender.com/" rel="noopener noreferrer"&gt;https://opentriage-0xzg.onrender.com/&lt;/a&gt; (Render free tier, so the first load can take a minute to wake up)&lt;/li&gt;
&lt;li&gt;Start with the precomputed samples (&lt;code&gt;expressjs/express&lt;/code&gt;, &lt;code&gt;colinhacks/zod&lt;/code&gt;). Live runs on other public repos are capped. Hosted mode sends public issue text to a hosted model API, so use local mode for private repos.&lt;/li&gt;
&lt;li&gt;💻 &lt;strong&gt;Run it yourself.&lt;/strong&gt; It's built to run on a laptop:
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/pkprajapati7402/OpenTriage.git
&lt;span class="nb"&gt;cd &lt;/span&gt;OpenTriage
ollama pull gemma3:1b &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; ollama pull nomic-embed-text
&lt;span class="nb"&gt;cp&lt;/span&gt; .env.example .env.local   &lt;span class="c"&gt;# add a read-only GITHUB_TOKEN&lt;/span&gt;
npm &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; npm run dev
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/pkprajapati7402" rel="noopener noreferrer"&gt;
        pkprajapati7402
      &lt;/a&gt; / &lt;a href="https://github.com/pkprajapati7402/OpenTriage" rel="noopener noreferrer"&gt;
        OpenTriage
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      A local-first, open-model triage assistant for open-source maintainers. It automates issue labeling, duplicate detection, and PR reviews using small open-weight models while keeping data private and the human in the loop.
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;div&gt;
&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;🧭 OpenTriage&lt;/h1&gt;
&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;A local-first, open-model triage assistant for open-source maintainers.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Suggests labels, spots duplicates, flags low-effort PRs, and drafts kind first replies. It runs on a modest laptop with small open-weight models, and a human always stays in charge.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/pkprajapati7402/OpenTriage#-models" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/a38391767687340beffc71f5fe9b2ae0ff1accd56c25e205fbdbc67be909768c/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f6275696c74253230776974682d47656d6d612d343238354634" alt="Built with Gemma"&gt;&lt;/a&gt;
&lt;a href="https://github.com/pkprajapati7402/OpenTriage#-quick-start-local-mode" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/817194f8f19616925438bc0e436b1fe1053d37471943e7819fd0bb7c8153de19/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f72756e732d6c6f63616c6c792d326561343466" alt="Local first"&gt;&lt;/a&gt;
&lt;a href="https://github.com/pkprajapati7402/OpenTriage#-privacy--safety" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/f73dd2976b79efd16c82240380ae67d6657c50136e6d690c482dc371752a234c/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f4769744875622532306163636573732d726561642d2d6f6e6c792d626c7565" alt="Read-only"&gt;&lt;/a&gt;
&lt;a href="https://github.com/pkprajapati7402/OpenTriage/LICENSE" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/07a7d0169027aac6d7a0bfa8964dfef5fbc40d5a2075cabb3d8bc67e17be3451/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f6c6963656e73652d4d49542d79656c6c6f772e737667" alt="License: MIT"&gt;&lt;/a&gt;
&lt;a href="https://hacktoberfest.com" rel="nofollow noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/68d3ff08fe8694066e494996c5fa4fe966624b54818667407f5d82f2d39ac626/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f4861636b746f626572666573742d323032362d666636396234" alt="Hacktoberfest 2026"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://github.com/pkprajapati7402/OpenTriage#" rel="noopener noreferrer"&gt;&lt;strong&gt;Live demo&lt;/strong&gt;&lt;/a&gt; · &lt;a href="https://github.com/pkprajapati7402/OpenTriage#" rel="noopener noreferrer"&gt;&lt;strong&gt;Demo video&lt;/strong&gt;&lt;/a&gt; · &lt;a href="https://github.com/pkprajapati7402/OpenTriage#" rel="noopener noreferrer"&gt;&lt;strong&gt;DEV post&lt;/strong&gt;&lt;/a&gt; · &lt;a href="https://github.com/pkprajapati7402/OpenTriage/project-details.md" rel="noopener noreferrer"&gt;&lt;strong&gt;Project details&lt;/strong&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Built for the DEV Hacktoberfest 2026 Weekend Challenge: &lt;em&gt;Build for a Friend&lt;/em&gt;.&lt;/p&gt;
&lt;/div&gt;

&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;📖 The story&lt;/h2&gt;
&lt;/div&gt;
&lt;p&gt;Maintaining an open-source project means a steady stream of issues and pull requests, usually handled in spare time. During Hacktoberfest the stream turns into a flood: duplicate reports, issues with no reproduction steps, and pull requests that change a single space in a README.&lt;/p&gt;
&lt;p&gt;OpenTriage was built for &lt;strong&gt;open-source maintainers and developer friends&lt;/strong&gt; juggling full-time engineering work while maintaining active open-source projects. They shared the core frustration: &lt;em&gt;"I spend hours every weekend filtering through one-character typo PRs, detecting duplicate bug reports, and asking submitters&lt;/em&gt;…&lt;/p&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/pkprajapati7402/OpenTriage" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Stack:&lt;/strong&gt; Next.js + TypeScript + Tailwind, &lt;code&gt;zod&lt;/code&gt; for schema validation, &lt;code&gt;vitest&lt;/code&gt; for the deterministic parts.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Open-source AI:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Gemma&lt;/strong&gt; (&lt;code&gt;gemma3:1b&lt;/code&gt; via Ollama) for label re-ranking and reply drafts, running locally. I develop on a Ryzen 5 5500U laptop with 8 GB RAM, so everything is sized for that.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Local embeddings&lt;/strong&gt; (&lt;code&gt;nomic-embed-text&lt;/code&gt;) plus a BM25 lexical score for duplicate detection. It still works without embeddings.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hosted open-weight models&lt;/strong&gt; (Gemma through OpenRouter and a Qwen model through Groq) behind the same provider interface, used for the demo and for comparison. Swapping models is one config change.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Design decisions that mattered:&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Retrieval first.&lt;/strong&gt; Suggestions start from similar past items and the labels maintainers really used, because small models are weak at open-ended judgment.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rules decide risk.&lt;/strong&gt; The low-effort PR band comes from deterministic rules. The model can explain it but can't escalate it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Constrained, validated output.&lt;/strong&gt; The model picks from the repo's own labels and returns JSON that must pass a schema. Unknown labels are dropped.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Abstaining is a feature.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Issue text is untrusted.&lt;/strong&gt; The model has no tools, and item text is delimited and treated as data.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  What the numbers say
&lt;/h3&gt;

&lt;p&gt;I evaluated with a temporal split: the oldest 70% of items form the history, and the newest 30% are the unseen test set. I tried it on [N] public repos while developing. The tables report the repos I fully evaluated, &lt;code&gt;express&lt;/code&gt; and &lt;code&gt;zod&lt;/code&gt;. &lt;strong&gt;Samples are small&lt;/strong&gt; (&lt;code&gt;express&lt;/code&gt;: n=[fill] labeled test items; &lt;code&gt;zod&lt;/code&gt;: n=[fill]), so treat these as indicative.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Label suggestions (&lt;code&gt;express&lt;/code&gt;)&lt;/th&gt;
&lt;th&gt;Precision&lt;/th&gt;
&lt;th&gt;Recall&lt;/th&gt;
&lt;th&gt;F1&lt;/th&gt;
&lt;th&gt;Abstain&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;k-NN baseline (no LLM)&lt;/td&gt;
&lt;td&gt;71.4%&lt;/td&gt;
&lt;td&gt;45.5%&lt;/td&gt;
&lt;td&gt;55.6%&lt;/td&gt;
&lt;td&gt;36.4%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hosted Gemma + retrieval&lt;/td&gt;
&lt;td&gt;14.3%&lt;/td&gt;
&lt;td&gt;36.4%&lt;/td&gt;
&lt;td&gt;20.5%&lt;/td&gt;
&lt;td&gt;0%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hosted Qwen + retrieval&lt;/td&gt;
&lt;td&gt;50.0%&lt;/td&gt;
&lt;td&gt;63.6%&lt;/td&gt;
&lt;td&gt;56.0%&lt;/td&gt;
&lt;td&gt;27.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The no-LLM baseline beat Gemma on precision.&lt;/strong&gt; For repo-specific labels, history mattered more than model size. Gemma never abstained, so it was wrong more often. I did not measure local Gemma 1B accuracy.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Duplicates:&lt;/strong&gt; Recall@3 was 100% on both repos (small samples).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Low-effort PRs:&lt;/strong&gt; precision 100%; recall 57.1% (&lt;code&gt;express&lt;/code&gt;) and 85% (&lt;code&gt;zod&lt;/code&gt;). The ground truth is a proxy (closed-unmerged PRs with &lt;code&gt;invalid&lt;/code&gt;/&lt;code&gt;spam&lt;/code&gt; labels), so it is noisy.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Where it fails:&lt;/strong&gt; &lt;code&gt;express&lt;/code&gt; maintainers label PRs by release branch (&lt;code&gt;4.x&lt;/code&gt;/&lt;code&gt;5.x&lt;/code&gt;), which a content-based model can't guess. Another PR was tagged &lt;code&gt;duplicate&lt;/code&gt; while the model predicted &lt;code&gt;pr&lt;/code&gt; and &lt;code&gt;javascript&lt;/code&gt; from the diff.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Speed and memory:&lt;/strong&gt; [include only if measured: ~X tokens/sec and ~Y GB peak RAM for Gemma 1B on the 8 GB laptop].&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Does Open Innovation Matter?
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Privacy:&lt;/strong&gt; a maintainer's private or pre-release repo shouldn't have to leave their machine. With Ollama it doesn't.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost:&lt;/strong&gt; a volunteer shouldn't pay per token. Local inference is free to run.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Low-end hardware:&lt;/strong&gt; it was built and run on an 8 GB laptop, which depends on small open weights.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Swappable and auditable:&lt;/strong&gt; I compared three models by changing config, and the risk signals are readable rules, not a black box. If a model misreads a repo, the maintainer can change the model or the label descriptions.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;To be fair, the hosted demo does send public issue text to a third-party API. The privacy claim applies to local mode.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prize Categories
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Best Use of Gemma&lt;/strong&gt;: Gemma runs locally via Ollama for labels and drafts, and hosted Gemma powers the demo and model comparison.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Best Use of Render&lt;/strong&gt;: the hosted demo runs on Render, with limited session credentials. high throughput can be achieved locally.
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>devchallenge</category>
      <category>weekendchallenge</category>
      <category>hf26challenge</category>
    </item>
  </channel>
</rss>
