In the first post I shared laya-triage, a GitHub Action that labels new issues as bug, feature, question or docs with a fine-tuned Laya model that runs inside the Actions runner. No API key, and issue text never leaves GitHub.
v1.0 scored 88.8% on the NLBSE'23 benchmark, and I was happy with that. Then I tested it on issues people are writing today, and it dropped to 79.8%. This post is about closing part of that gap, what I added in v1.2, and what happened when I measured it against OpenAI's new Decisions API.
I'm a systems engineering student still learning this field, so take it as notes from someone figuring things out, with every number measured and published in the repo.
Benchmarks age
NLBSE'23 is a great benchmark, but most of its issues are a few years old. To check how the model handles today's repositories, I collected 10,026 closed issues from 2025 and 2026, across 288 active repositories with more than 1,000 stars. I kept only issues labeled by a maintainer, not by the author or a template, and none of those repositories were used for training.
On that set, v1.0 got 79.8%. Same model, nine points lower.
What worked, and what didn't
My first idea was to fine-tune v1.0 on 127,140 recent issues from other repositories. It improved recent issues to 81.9%, but it lost almost a point on NLBSE'23. The model was forgetting what it already knew.
So I trained v1.1 from scratch on both: the original million NLBSE'23 issues plus the recent ones, together. It kept 88.8% on NLBSE'23 and reached 82.3% on recent issues.
The biggest change was in questions. v1.0 recognized 43.7% of them; v1.1 recognizes 58.3%. Questions are the hardest class, because many read like bug reports ("it doesn't work when I...").
v1.2: things maintainers asked for
v1.2 keeps the v1.1 models and adds:
- Backlog mode. Run it once by hand and it labels the open issues you already have. Only the ones it's confident about, and it never comments on old issues.
- Run summaries. Every run writes a table with its decisions to the workflow summary, so you can review them without reading logs.
- Repo memory (experimental). When the model hesitates between two types, the 20 most similar closed issues of your repository vote. It helps a little (82.3% to 82.7%), so it's off by default until it's tested on more real repositories.
-
More label names, contributed by someone during Hacktoberfest, so
defectorusagemap to the right type.
Then OpenAI released a Decisions API
On September 29 OpenAI announced a Decisions API: you give it a question and a fixed list of answers, and it returns one answer with probabilities. It's the same kind of tool as Laya or TypeSafe's Jev, so I measured it the same way:
- the same question, word for word, that laya-triage was trained with;
- the exact 5,000 NLBSE'23 issues Jev answered;
- the recent-issues set, used once, with nothing tuned on it;
- the 2026 set and 14 languages.
| laya-triage v1.2 | Jev | OpenAI Decisions | |
|---|---|---|---|
| NLBSE'23 | 88.8% | 84.4% | 83.7% |
| Recent issues | 82.3% | 78.1% | 78.1% |
| 2026 set, macro F1 | 0.758 | 0.660 | 0.602 |
| 14 languages, avg. without English | 85.8% | 85.1% | 82.7% |
| Cost of the recent-issues run | $0 | $0.30 | $0.46 |
The Decisions API is still in beta, so these numbers may change. I measured it on October 6, 2026.
The number I find most useful isn't accuracy, though. It's what happens out of every 100 new issues:
All three get about 77 right on recent issues. The difference is the rest: the hosted APIs label almost everything, so 20 or 21 issues get a wrong label. laya-triage only labels when its confidence is at least 0.60, so it puts a wrong label on 12 and leaves 11 for a maintainer with needs-triage. For a bot that writes in your repository, being wrong less often matters more than labeling everything.
What I learned
- A benchmark score is a starting point. 88.8% on a known benchmark became 79.8% on today's repositories. Measuring on fresh data changed what I worked on.
- A small, specialized model can hold its own. OpenAI's model is general: it can decide almost anything. laya-triage was trained for one question with more than a million examples. On that one question, the specialist does better.
- Knowing when to stop matters. Calibrated confidence lets the action say "I'm not sure" instead of guessing.
-
Measure once and publish everything. I decide on a validation split and use the test set only once. The scripts and raw results are all in
bench/.
Try it
Prefer to watch? There's a one-minute setup video in the README. The short version:
name: Triage
on:
issues:
types: [opened]
permissions:
contents: read
issues: write
jobs:
triage:
runs-on: ubuntu-latest
steps:
- uses: elnachto/laya-triage@v1
It starts in dry-run mode, so it only shows what it would do until you switch it on.
Next, I want to look at duplicate issues and pull requests. If you try laya-triage on your repository, I'd really like to hear how it does, and especially where it gets things wrong.


Top comments (0)