DEV Community

Cover image for From 51% to 90.5%: fine-tuning a local triage model on 10,003 support tickets (open source and measured)
gj0xv
gj0xv

Posted on

From 51% to 90.5%: fine-tuning a local triage model on 10,003 support tickets (open source and measured)

When OpenAI launched the Decisions API on Luna this week, and TypeSafe launched Jev two weeks before that, they were both making the same bet: the future of AI in production is not bigger models, it is smaller, faster decision models embedded in your workflow.

I agree. And I have the numbers to prove it works.

laya-triage routes support tickets through a local decision model: intent (77 categories), department (8), urgency, frustration, churn risk, refund flag, and a confidence-gated escalation to a human. All in one forward pass, on CPU, for free.

Hierarchical intent accuracy: 51.0% -> 90.5% (after fine-tuning)
Coarse-cluster accuracy: 68.0% -> 96.0%
Urgency MAE: 0.81 -> 0.70
Cost per ticket: $0
Enter fullscreen mode Exit fullscreen mode

The architecture

Flat 77-way classification is hard for a small model. So I broke it into two stages:

Ticket -> [coarse stage: 12 clusters] -> [fine stage: 3-10 candidate intents] -> decision
              |
              +-> urgency, frustration, churn, refund (same forward pass)
              |
              +-> confidence check -> escalate to human if below 0.84
Enter fullscreen mode Exit fullscreen mode

The hierarchy alone added 14.5 percentage points over flat routing. The fine-tune added another 39.5.

What I fine-tuned on

10,003 BANKING77 tickets, split into two sequences per ticket, trained on Kaggle's free 2x T4 GPUs. The checkpoint is published on HuggingFace: Gjusev/laya-triage-banking77.

The eval harness that tracked the improvement ran on Z.ai's GLM models, with GLM-5.3-flashX as the workhorse. When the judge model changes, the numbers move a little. The fine-tune gain does not.

The results, before and after

Measurement Zero-shot Fine-tuned
Hierarchical intent accuracy 51.0% 90.5%
Macro-F1 not measured 0.848
Coarse-cluster accuracy 68.0% 96.0%
Urgency MAE 0.81 0.70
Frustration MAE 1.07 1.07 (untouched heads)

At the 0.84 confidence threshold (zero-shot): 75.2% accuracy with 60.5% of tickets auto-handled. The rest escalate to a human with a recorded reason.

The escalation policy (and why it matters)

The threshold gates only on cluster and intent confidence. Signal confidence (urgency, frustration, churn) is deliberately not used, because score-type questions have structurally lower confidence and would trigger near-universal handoffs.

On an 8-ticket out-of-domain smoke test (IT-ops tickets against banking intents): 7 escalated correctly. The one that did not, an SSO lockout, mapped to "unable to verify identity". Not wrong, exactly, but the reason it matters is that the behavior is visible and fixable, not hidden behind an opaque binary.

Try it

The demo is live on Streamlit, in 6 languages (English, Spanish, French, German, Hindi, Arabic):

Live demo

Or run it yourself:

pip install laya-triage
Enter fullscreen mode Exit fullscreen mode

Fine-tuning notebook included for Kaggle (2x T4, ~4 hours).

The rest of the series

  • laya-router: route prompts between cheap and frontier models, 54.9% cost reduction
  • laya-compactor: cut 70% of RAG context tokens, same answer quality
  • laya-phishield: explainable phishing detection, $0 per 1,000 emails

All four tools use the same laya decision engine, Apache 2.0, local, with published benchmarks. If OpenAI's Decisions API and TypeSafe's Jev are the proprietary versions of this idea, this series is the open one. Run it on your machine, keep your data, pay nothing per decision.

Top comments (0)