DEV Community

Cover image for Building laya-triage, part 2: today's issues, a new model, and OpenAI's Decisions API
Ignacio Joaquin Sanga Olmos
Ignacio Joaquin Sanga Olmos

Posted on

Building laya-triage, part 2: today's issues, a new model, and OpenAI's Decisions API

In the first post I shared laya-triage, a GitHub Action that labels new issues as bug, feature, question or docs with a fine-tuned Laya model that runs inside the Actions runner. No API key, and issue text never leaves GitHub.

v1.0 scored 88.8% on the NLBSE'23 benchmark, and I was happy with that. Then I tested it on issues people are writing today, and it dropped to 79.8%. This post is about closing part of that gap, what I added in v1.2, and what happened when I measured it against OpenAI's new Decisions API.

I'm a systems engineering student still learning this field, so take it as notes from someone figuring things out, with every number measured and published in the repo.

Benchmarks age

NLBSE'23 is a great benchmark, but most of its issues are a few years old. To check how the model handles today's repositories, I collected 10,026 closed issues from 2025 and 2026, across 288 active repositories with more than 1,000 stars. I kept only issues labeled by a maintainer, not by the author or a template, and none of those repositories were used for training.

On that set, v1.0 got 79.8%. Same model, nine points lower.

What worked, and what didn't

My first idea was to fine-tune v1.0 on 127,140 recent issues from other repositories. It improved recent issues to 81.9%, but it lost almost a point on NLBSE'23. The model was forgetting what it already knew.

So I trained v1.1 from scratch on both: the original million NLBSE'23 issues plus the recent ones, together. It kept 88.8% on NLBSE'23 and reached 82.3% on recent issues.

Accuracy on recent issues by training data

The biggest change was in questions. v1.0 recognized 43.7% of them; v1.1 recognizes 58.3%. Questions are the hardest class, because many read like bug reports ("it doesn't work when I...").

v1.2: things maintainers asked for

v1.2 keeps the v1.1 models and adds:

  • Backlog mode. Run it once by hand and it labels the open issues you already have. Only the ones it's confident about, and it never comments on old issues.
  • Run summaries. Every run writes a table with its decisions to the workflow summary, so you can review them without reading logs.
  • Repo memory (experimental). When the model hesitates between two types, the 20 most similar closed issues of your repository vote. It helps a little (82.3% to 82.7%), so it's off by default until it's tested on more real repositories.
  • More label names, contributed by someone during Hacktoberfest, so defect or usage map to the right type.

Then OpenAI released a Decisions API

On September 29 OpenAI announced a Decisions API: you give it a question and a fixed list of answers, and it returns one answer with probabilities. It's the same kind of tool as Laya or TypeSafe's Jev, so I measured it the same way:

  • the same question, word for word, that laya-triage was trained with;
  • the exact 5,000 NLBSE'23 issues Jev answered;
  • the recent-issues set, used once, with nothing tuned on it;
  • the 2026 set and 14 languages.
laya-triage v1.2 Jev OpenAI Decisions
NLBSE'23 88.8% 84.4% 83.7%
Recent issues 82.3% 78.1% 78.1%
2026 set, macro F1 0.758 0.660 0.602
14 languages, avg. without English 85.8% 85.1% 82.7%
Cost of the recent-issues run $0 $0.30 $0.46

The Decisions API is still in beta, so these numbers may change. I measured it on October 6, 2026.

The number I find most useful isn't accuracy, though. It's what happens out of every 100 new issues:

Out of every 100 new issues

All three get about 77 right on recent issues. The difference is the rest: the hosted APIs label almost everything, so 20 or 21 issues get a wrong label. laya-triage only labels when its confidence is at least 0.60, so it puts a wrong label on 12 and leaves 11 for a maintainer with needs-triage. For a bot that writes in your repository, being wrong less often matters more than labeling everything.

What I learned

  1. A benchmark score is a starting point. 88.8% on a known benchmark became 79.8% on today's repositories. Measuring on fresh data changed what I worked on.
  2. A small, specialized model can hold its own. OpenAI's model is general: it can decide almost anything. laya-triage was trained for one question with more than a million examples. On that one question, the specialist does better.
  3. Knowing when to stop matters. Calibrated confidence lets the action say "I'm not sure" instead of guessing.
  4. Measure once and publish everything. I decide on a validation split and use the test set only once. The scripts and raw results are all in bench/.

Try it

Prefer to watch? There's a one-minute setup video in the README. The short version:

name: Triage
on:
  issues:
    types: [opened]
permissions:
  contents: read
  issues: write
jobs:
  triage:
    runs-on: ubuntu-latest
    steps:
      - uses: elnachto/laya-triage@v1
Enter fullscreen mode Exit fullscreen mode

It starts in dry-run mode, so it only shows what it would do until you switch it on.

Next, I want to look at duplicate issues and pull requests. If you try laya-triage on your repository, I'd really like to hear how it does, and especially where it gets things wrong.

Repo: github.com/elnachto/laya-triage

Top comments (0)