DEV Community

Cover image for The boring part of an AI agent is the part that works
Kennedy Njoroge
Kennedy Njoroge

Posted on

The boring part of an AI agent is the part that works

Most of the time an incident spends "being worked on" isn't work. It's waiting to reach someone who can act on it. An alert fires, an email lands in a shared inbox, somebody skims it, guesses which team owns it, forwards it, and that team decides it isn't theirs. On a bad night that loop costs you an hour before anyone has opened a terminal.

I built an agent to collapse that loop for a BSS platform, and it took mean time to resolution on routed incidents from around two hours to about ten minutes. What surprised me is how little of that came from the language model. The model does one small job in the middle. Everything that makes the system trustworthy is plumbing.

This post is about the plumbing.

The shape of the problem

Incoming incident email is semi-structured in the worst way: there's a template, three systems ignore it, two append their own stack traces, and the one that matters most sends plain prose from a human. Regex gets you maybe 60% of the way and then rots every time an upstream team edits their alert body.

That 40% tail is exactly where an LLM earns its place. But "classify this email" is not the system. The system is:

poll inbox → normalise → classify → resolve owner → create ticket → notify → observe
Enter fullscreen mode Exit fullscreen mode

The model is step three. Steps one, two, four, five, six and seven decide whether anyone trusts it.

Normalise before you classify

The single highest-leverage thing I did was aggressively strip the email before the model ever saw it. Signatures, legal footers, quoted reply chains, base64 inline images, and the forty lines of Received: headers are pure token cost and pure distraction.

import re
from email import policy
from email.parser import BytesParser

QUOTE_MARKERS = (
    "-----Original Message-----",
    "From:",
    "On ... wrote:",
)

def extract_body(raw: bytes) -> str:
    msg = BytesParser(policy=policy.default).parsebytes(raw)
    part = msg.get_body(preferencelist=("plain", "html"))
    text = part.get_content() if part else ""

    # cut the reply chain at the first quote marker
    for marker in QUOTE_MARKERS:
        idx = text.find(marker)
        if idx > 0:
            text = text[:idx]

    text = re.sub(r"\n{3,}", "\n\n", text)
    return text.strip()[:4000]
Enter fullscreen mode Exit fullscreen mode

That [:4000] matters. Truncation is a feature. An incident email that needs more than 4000 characters of context to classify is one a human should be reading anyway.

Make the model return a shape, not a sentence

The classification call is deliberately small and deliberately constrained. I ask for structured output with a fixed enum of teams, a confidence score, and a one-line rationale.

from pydantic import BaseModel, Field
from typing import Literal

Team = Literal[
    "billing", "provisioning", "network", "identity",
    "integrations", "data-platform", "unknown",
]

class Triage(BaseModel):
    team: Team
    severity: Literal["sev1", "sev2", "sev3", "sev4"]
    confidence: float = Field(ge=0.0, le=1.0)
    rationale: str = Field(max_length=200)
Enter fullscreen mode Exit fullscreen mode

Three details that are easy to skip and expensive to skip:

unknown is a valid answer. If the enum has no escape hatch, the model will pick the nearest plausible team and sound certain about it. Giving it somewhere honest to go is the cheapest accuracy win available.

The rationale is capped. Not because anyone reads it in the happy path, but because when triage goes wrong at 3 a.m. the rationale is the only artefact that tells you whether the model misread the email or the owner map was stale. A 200-character cap keeps it a signal rather than an essay.

Confidence is the model's, and it's directional at best. Treat it as a tiebreaker, not a probability. I threshold it, but the threshold is tuned against observed outcomes, not against what the number claims to mean.

The owner map is not the model's problem

Here's the part that no amount of prompting fixes: the model can tell you a ticket is about billing. It cannot know that billing's on-call rotated yesterday, or that the billing team split into two squads last quarter.

So the model resolves to a team, never to a person. A separate lookup against the directory — in my case Microsoft Graph — resolves the team to whoever is actually on call right now.

async def resolve_owner(team: str, graph) -> Owner:
    group = OWNER_GROUPS[team]            # static, version-controlled mapping
    members = await graph.group_members(group)
    on_call = await graph.on_call_for(group)   # calendar-driven
    return on_call or members[0]
Enter fullscreen mode Exit fullscreen mode

This split is what makes the thing survive org changes. When a team reorganises, somebody edits a YAML file. Nobody retrains anything, nobody rewrites a prompt. The LLM's view of the world is "what is this email about", which is stable. The volatile knowledge lives in systems that are already the source of truth for it.

If you take one thing from this post: never let the model hold facts that another system owns.

Route low confidence to a human, loudly

async def handle(mail: Mail) -> None:
    body = extract_body(mail.raw)
    triage = await classify(body)

    if triage.team == "unknown" or triage.confidence < 0.72:
        await post_to_triage_channel(mail, triage, reason="low_confidence")
        return

    owner = await resolve_owner(triage.team, graph)
    ticket = await create_ticket(mail, triage, owner)
    await notify(owner, ticket, triage)
Enter fullscreen mode Exit fullscreen mode

The fallback is not a failure path, it's the default path for anything ambiguous, and it lands somewhere visible. Early on, roughly one in six emails went to the human channel. That felt like a bad number until I looked at what was in it: most were genuinely ambiguous, and the ones that weren't told me exactly which alert template to fix upstream. The fallback queue is your best source of improvement work.

Idempotency, because email is not a queue

Mail APIs will hand you the same message twice. Your poller will crash between "created ticket" and "marked as read". Somebody will replay a batch to test something.

Every message gets a deterministic key, and ticket creation is conditional on it:

def dedupe_key(mail: Mail) -> str:
    # message-id is stable across redelivery; subject+sender is the fallback
    raw = mail.message_id or f"{mail.sender}|{mail.subject}|{mail.sent_at}"
    return hashlib.sha256(raw.encode()).hexdigest()
Enter fullscreen mode Exit fullscreen mode

Then a unique index on that key in whatever store backs you. Not a cache — a unique index. A cache will happily evict the thing under memory pressure and you'll find out when duplicate tickets land during your busiest hour. I have opinions about cache keys and silent data loss that I'll save for another post.

What I'd measure from day one

I started with "did it classify correctly", which turns out to be the least useful metric, because you can only compute it by hand on a sample.

What actually drove decisions:

  • Reroute rate — tickets reassigned within 30 minutes of creation. This is your real accuracy signal and it's free, because humans generate it just by doing their jobs.
  • Fallback rate — share going to the human channel. Trending up means an upstream alert format changed.
  • Time to first human action — the number the whole project exists to move. Not time to resolution; you don't control that.
  • Cost per message — unglamorous, but it's what lets you argue for the system's continued existence.

Reroute rate is the one I'd instrument first in any triage system. It's ground truth that costs nothing to collect.

The honest summary

The model is a classifier with good manners about unstructured text. That's genuinely useful and regex could not have done it. But the two-hour-to-ten-minute number came from: removing a human hop, resolving on-call automatically, and making the ambiguous cases land somewhere visible instead of somewhere polite.

If you're building something similar, budget your time accordingly. The prompt will take an afternoon. The owner map, the idempotency, the fallback path and the metrics will take the rest of the month, and they're the reason it still works six months later.


I write about backend engineering, data pipelines and reliability. If you've built something in this space and solved the owner-mapping problem differently, I'd like to hear it.

Top comments (0)