This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
OpenTriage is a local-first triage assistant for open-source maintainers. Point it at a public GitHub repo and, for every issue and PR, it suggests:
- ๐ท๏ธ Labels, chosen only from the repo's existing labels, with a one-line reason (it abstains when unsure)
- ๐ Possible duplicates: the top 3 similar past issues with similarity scores
- ๐ฉ Low-effort PR signals: transparent rules like an empty description, a whitespace-only diff, or a trivial README edit
- โ๏ธ Draft first replies that politely ask for missing details, for the maintainer to edit
It is strictly read-only. It never labels, comments, or closes anything on GitHub. A human always decides.
Who I built it for
I built OpenTriage for open-source maintainers and contributors who spend significant time reviewing issues and pull requests, especially in repositories with active communities and growing backlogs.
Every October the backlog gets worse for maintainers everywhere: duplicate reports, issues with no reproduction steps, and PRs that change a single space in a README. The first pass is repetitive, and it eats weekends.
To go beyond one repo, I also tested OpenTriage against the real history of public repos, where maintainers' actual labeling and duplicate decisions are visible, and I report where it fails below.
Demo
- ๐ Live demo: https://opentriage-0xzg.onrender.com/ (Render free tier, so the first load can take a minute to wake up)
- Start with the precomputed samples (
expressjs/express,colinhacks/zod). Live runs on other public repos are capped. Hosted mode sends public issue text to a hosted model API, so use local mode for private repos. - ๐ป Run it yourself. It's built to run on a laptop:
git clone https://github.com/pkprajapati7402/OpenTriage.git
cd OpenTriage
ollama pull gemma3:1b && ollama pull nomic-embed-text
cp .env.example .env.local # add a read-only GITHUB_TOKEN
npm install && npm run dev
Code
pkprajapati7402
/
OpenTriage
A local-first, open-model triage assistant for open-source maintainers. It automates issue labeling, duplicate detection, and PR reviews using small open-weight models while keeping data private and the human in the loop.
๐งญ OpenTriage
A local-first, open-model triage assistant for open-source maintainers.
Suggests labels, spots duplicates, flags low-effort PRs, and drafts kind first replies. It runs on a modest laptop with small open-weight models, and a human always stays in charge.
Live demo ยท Demo video ยท DEV post ยท Project details
Built for the DEV Hacktoberfest 2026 Weekend Challenge: Build for a Friend.
๐ The story
Maintaining an open-source project means a steady stream of issues and pull requests, usually handled in spare time. During Hacktoberfest the stream turns into a flood: duplicate reports, issues with no reproduction steps, and pull requests that change a single space in a README.
OpenTriage was built for open-source maintainers and developer friends juggling full-time engineering work while maintaining active open-source projects. They shared the core frustration: "I spend hours every weekend filtering through one-character typo PRs, detecting duplicate bug reports, and asking submittersโฆ
How I Built It
Stack: Next.js + TypeScript + Tailwind, zod for schema validation, vitest for the deterministic parts.
Open-source AI:
-
Gemma (
gemma3:1bvia Ollama) for label re-ranking and reply drafts, running locally. I develop on a Ryzen 5 5500U laptop with 8 GB RAM, so everything is sized for that. -
Local embeddings (
nomic-embed-text) plus a BM25 lexical score for duplicate detection. It still works without embeddings. - Hosted open-weight models (Gemma through OpenRouter and a Qwen model through Groq) behind the same provider interface, used for the demo and for comparison. Swapping models is one config change.
Design decisions that mattered:
- Retrieval first. Suggestions start from similar past items and the labels maintainers really used, because small models are weak at open-ended judgment.
- Rules decide risk. The low-effort PR band comes from deterministic rules. The model can explain it but can't escalate it.
- Constrained, validated output. The model picks from the repo's own labels and returns JSON that must pass a schema. Unknown labels are dropped.
- Abstaining is a feature.
- Issue text is untrusted. The model has no tools, and item text is delimited and treated as data.
What the numbers say
I evaluated with a temporal split: the oldest 70% of items form the history, and the newest 30% are the unseen test set. I tried it on [N] public repos while developing. The tables report the repos I fully evaluated, express and zod. Samples are small (express: n=[fill] labeled test items; zod: n=[fill]), so treat these as indicative.
Label suggestions (express) |
Precision | Recall | F1 | Abstain |
|---|---|---|---|---|
| k-NN baseline (no LLM) | 71.4% | 45.5% | 55.6% | 36.4% |
| Hosted Gemma + retrieval | 14.3% | 36.4% | 20.5% | 0% |
| Hosted Qwen + retrieval | 50.0% | 63.6% | 56.0% | 27.3% |
- The no-LLM baseline beat Gemma on precision. For repo-specific labels, history mattered more than model size. Gemma never abstained, so it was wrong more often. I did not measure local Gemma 1B accuracy.
- Duplicates: Recall@3 was 100% on both repos (small samples).
-
Low-effort PRs: precision 100%; recall 57.1% (
express) and 85% (zod). The ground truth is a proxy (closed-unmerged PRs withinvalid/spamlabels), so it is noisy.
Where it fails: express maintainers label PRs by release branch (4.x/5.x), which a content-based model can't guess. Another PR was tagged duplicate while the model predicted pr and javascript from the diff.
Speed and memory: [include only if measured: ~X tokens/sec and ~Y GB peak RAM for Gemma 1B on the 8 GB laptop].
Why Does Open Innovation Matter?
- Privacy: a maintainer's private or pre-release repo shouldn't have to leave their machine. With Ollama it doesn't.
- Cost: a volunteer shouldn't pay per token. Local inference is free to run.
- Low-end hardware: it was built and run on an 8 GB laptop, which depends on small open weights.
- Swappable and auditable: I compared three models by changing config, and the risk signals are readable rules, not a black box. If a model misreads a repo, the maintainer can change the model or the label descriptions.
To be fair, the hosted demo does send public issue text to a third-party API. The privacy claim applies to local mode.
Prize Categories
- Best Use of Gemma: Gemma runs locally via Ollama for labels and drafts, and hosted Gemma powers the demo and model comparison.
- Best Use of Render: the hosted demo runs on Render, with limited session credentials. high throughput can be achieved locally.
Top comments (0)