DEV Community

Danil Galeev for MadeBy.Expert

Posted on

I built an L1 ticket deflector that knows when to stop: LangGraph + human-in-the-loop

I built an L1 ticket deflector that knows when to stop: LangGraph + human-in-the-loop

The same tickets arrive at every IT service desk: password resets, VPN that won't connect, "how do I install 7-Zip?", a printer that's offline again. A widely cited industry range puts routine requests at 50–70% of L1 volume — work a script could handle, except nobody trusts a script with the tickets that actually matter.

Building an agent that answers everything is easy. Building one that knows when to hand a ticket to a human is not. I built the second kind and open-sourced it.

Repo: github.com/zedxter/l1-ticket-deflector (MIT)

The design decision that shapes everything

Most "AI helpdesk" demos optimize for deflection rate. That's the wrong target. A high deflection rate is trivial if you let the model guess.

The constraint I optimized for instead: never let the agent act on a sensitive request. Anything touching access, privileges, security, or money goes to a human with prepared context. The agent resolves the boring 30% and escalates the rest — with the relevant KB article already attached, so the human doesn't start from zero.

I called it human-in-the-loop, but the honest framing is narrower: the agent has a hard boundary, and that boundary is the part worth testing.

The graph

The production version is a LangGraph state machine. Three outcomes, no cleverness:

classify ──> no match / low confidence ──> queue (general L1 queue)
   │
   ├── sensitivity = low        ──> auto_resolve ──> notify_user
   └── sensitivity = high/crit  ──> human_review ──> escalate ──> notify_user
Enter fullscreen mode Exit fullscreen mode

The classifier returns a KB article ID and a confidence score. Two gates decide the path:

  1. Confidence gate — below 0.55, the ticket goes to the general queue. No guessing.
  2. Sensitivity gate — high/critical tickets go to service_desk_l2 or security_ops, never to auto-resolve.

Each KB article carries a sensitivity label. That's the trick: the model doesn't decide what's dangerous. The knowledge base does, and it's a human-authored field. The LLM only picks the article.

def route_after_classify(state):
    if not state.get("article_id") or state.get("confidence", 0) < CONFIDENCE_THRESHOLD:
        return "queue"
    if state.get("sensitivity") in AUTO_RESOLVE_SENSITIVITY:
        return "auto_resolve"
    return "human_review"
Enter fullscreen mode Exit fullscreen mode

That's the entire safety story in six lines. Sensitive categories (access grants, license purchases, phishing reports, ransomware) can never reach auto_resolve, no matter what the model says.

Two versions, on purpose

The repo ships two implementations:

  • demo/ — a zero-dependency Python script. Keyword matching plus stemming, no LLM. It runs on a Raspberry Pi. Its job is to make the mechanics visible and the numbers checkable.
  • graph/ — the LangGraph agent. RAG over the KB, structured output via Pydantic, tool-calls into the ITSM, MOCK_MODE=1 so you can run the logic without an API key.

I kept the offline version because a demo you can't run is just a claim. Anyone can clone the repo and reproduce the numbers in ten seconds.

The numbers (and why I'm suspicious of them)

On 45 labeled tickets, the offline classifier scored:

Metric Value
Auto-resolved 25
Escalated to human 18
No match → queue 2
Decision accuracy 100%
Routing accuracy 100%
Sample deflection 55.6%

100% accuracy is a red flag, not a trophy. It means the test set is too small and too clean. 45 tickets, hand-labeled, written by the same person who wrote the KB — of course it scores well. I'm publishing it because it's honest about what it is: a sanity check, not a benchmark. The real test is a client's messy backlog, and that number doesn't exist yet.

The deflection rate is the one I'd defend. 55.6% on this sample, but the ROI model uses a conservative 30%, because market benchmarks for first-year deflection sit at 20–40% and I'd rather be wrong in the client's favor.

The ROI model

Conservative assumptions for a ~500-person company in DACH:

Metric Value
L1 tickets / month ~1,200
Deflection 30%
Avg handling time 19 min
IT rate (fully loaded) €45 / hour
Time saved ~114 h/month (~0.7 FTE)
Savings ~€5,130 / month
Retainer €4,000 / month
ROI 1.3× / month (15× / year)

The retainer is deliberately below the savings. A model where the vendor captures all the value is a model nobody signs. 1.3×/month is not a spectacular number — it's a defensible one, and defensible is what gets a pilot approved.

What's fake and what's real

I want to be precise about this, because "proof of work" usually means "screenshot of a demo."

Real: the classifier, the graph, the routing logic, the KB structure, the ROI math. Clone it, run it, read every line.

Stubbed: the tool-calls. itsm_create_ticket, directory_reset_password, and mdm_push_install return strings, not real ITSM actions. Wiring them to Jira Service Management or Zammad over MCP is a pilot task, not a demo task.

Not built: the Slack/Teams approval UI, telemetry, multilingual KB. They're on the roadmap and I'm not pretending otherwise.

What I'd tell anyone building this

Three things I'd carry into the next version:

  1. Put the safety boundary in data, not in the prompt. A sensitivity field on each KB article is auditable. "The model is instructed to be careful" is not.
  2. Ship the runnable offline version. It costs an afternoon and buys all the credibility.
  3. Publish the suspicious number. 100% accuracy invites the right question — "on how many tickets?" — and answering it honestly is worth more than a flattering metric.

The repo is MIT-licensed and the whole point is that you can pull it apart. If you run an IT desk and want to compare deflection on your own tickets, that's the interesting experiment.

Repo: github.com/zedxter/l1-ticket-deflector

I build AI agents for reporting, CRM, and support workflows — with human review where it counts. Based in Potsdam, Germany.

Top comments (0)