DEV Community

Cover image for Routing support tickets with a promise on the errors (and what happens when a new kind of ticket shows up)
Maxim Kuznetsov
Maxim Kuznetsov

Posted on AI-assisted

Routing support tickets with a promise on the errors (and what happens when a new kind of ticket shows up)

Most support routers I have seen are a classifier plus a threshold someone picked by eye. It works until it doesn't, and then nobody can say how often it was wrong, or since when. This post builds a router that states its error up front, keeps it, hands what it is unsure of to a slower reader and then to a person, and records every decision. Then it breaks the stream on purpose: a third of the tickets become a kind nobody ever labelled.

I wrote solvi, the open-source library (Apache-2.0) used here. Everything below runs on pip install solvi==1.0.0, with no model and no network. This is part 1 of a five-part series; the overview of solvi 1.0 is here.

The setup

The pieces are deliberately ordinary:

  • System 1, the fast router: a small head over a few facts that plain functions compute from the ticket body.
  • The slow path: something that reads more (in production, an LLM; here, a stand-in function that also reads the subject line).
  • A person, for whatever neither can answer within the promise.
  • The promise: max_error=0.05 — at most 5% of the answers given without a person are wrong, for tickets like the labelled ones.

solvi.build takes the question, labelled examples and the promise, and wires the rest: it fits System 1, calibrates its threshold on examples it did not see, decides which slice of what System 1 hands over the slow path may answer, and stores every decision in a hash-chained file.

The code

The facts are plain Python. Each function's argument names are the facts it reads, its name is the fact it sets:

TEAMS = ["billing", "shipping", "technical", "account"]
WORDS = {"billing": ["charged", "invoice", "refund", "payment", "price"],
         "shipping": ["parcel", "delivery", "courier", "tracking", "arrived"],
         "technical": ["error", "crash", "login", "app", "bug"],
         "account": ["password", "email", "profile", "delete", "username"],
         "fraud": ["stolen", "unauthorized", "scam", "hacked", "suspicious"]}
FILLER = ["please", "help", "today", "order", "again", "thanks", "urgent", "still", "my", "the"]

cat = Catalog()


def words_of(team):                                  # one computed fact per team: how many of its words the text has
    def hits(text):
        low = text.lower()
        return sum(w in low for w in WORDS[team])
    hits.__name__ = f"{team}_words"
    return hits


for team in TEAMS:
    cat.fn(words_of(team))
Enter fullscreen mode Exit fullscreen mode

The tickets are synthetic: one or two words of their team, sometimes a word of another team, and filler. The subject line usually names the team, and System 1 does not read it; the slow path does:

def ticket(rng, kind):
    own = rng.sample(WORDS[kind], rng.choice([1, 1, 2]))
    noise = rng.sample(WORDS[rng.choice(TEAMS)], 1) if rng.random() < 0.35 else []
    words = own + noise + rng.sample(FILLER, 4)
    rng.shuffle(words)
    hint = kind if rng.random() < 0.97 else rng.choice(TEAMS)          # the subject line, which System 1 does not read
    return {"text": " ".join(words), "subject": f"Re: {rng.choice(WORDS[hint])} (ticket {rng.randint(10000, 99999)})"}


def reader(state):                                   # the slow path: a stand-in that also reads the subject line
    subject = state["subject"].lower()
    for t in TEAMS:
        if any(w in subject for w in WORDS[t]):
            return t
    low = state["text"].lower()                      # a subject no team's words match: the body decides
    return max(TEAMS, key=lambda t: sum(w in low for w in WORDS[t]))
Enter fullscreen mode Exit fullscreen mode

Two thousand labelled tickets, and one call:

def run(novel):
    store = Path(tempfile.mkdtemp()) / "decisions.jsonl"
    s = solvi.build(Question("team", "Which team handles this ticket?", Answer.choice(TEAMS)), examples,
                    catalog=cat, max_error=0.05, slow=reader, storage=store, novel=novel)
Enter fullscreen mode Exit fullscreen mode

With a real LLM the slow path is slow=llm(URL, "openai/gpt-oss-120b") plus price= for the dollars; the rest does not change.

What it chose, in its own words

s.explain() prints every choice build made and the numbers behind it. From the run:

Question 'team': 2000 labelled examples, shuffled with seed 0: 1500 to fit System 1, 250 to calibrate System 1's guarantee, 250 to calibrate the dispatcher.
System 1: a ridge head (System.fit) fitted on 1500 examples, reading 4 facts: account_words, technical_words, shipping_words, billing_words.
…
Its promise: error among the answers given alone ≤ 0.05 with probability ≥ 0.9, for inputs like the calibration examples. Calibrated on 250 examples: threshold 0.5895, answered alone 82.0%, error among them 0.00%, P(alone and wrong) 0.00%; AUROC of the signal 0.95.
Slow path: the function reader over text, subject.
Is a slice open to it? The probe on the dispatcher's 250 examples: System 1 held back 39 on the 'guarantee' slice.
  slice 'guarantee' (39 calibration examples): the slow path's answer when its confidence ≥ 1 (right on this slice: System 1 56%, slow path 95%, slow path agreeing with System 1 95%)
…
Not covered: inputs unlike the examples; the promise holds for inputs like the examples, so calibrate again when the inputs change.
Enter fullscreen mode Exit fullscreen mode

Read the slice line: on the tickets System 1 holds back, System 1 itself would be right 56% of the time and the slow path 95%, so build gives that slice to the slow path. If the slow path were not good enough there, the slice would go to a person instead — build measures that, it does not assume it. The last line is the most important one, and the rest of the post is about it.

Breaking the stream

Now 1,200 tickets. From ticket 600 on, a third are fraud reports — "stolen", "unauthorized", "hacked" — a team that does not exist in the labels. The same stream goes through two setups: novel=False, the plain promise, and novel="auto", which adds an open-set gate (more on it below). The tally, verbatim:

novel=False
  before the shift: by {'s1': 452, 's2': 148, 'human': 0}, wrong among answered alone 0.5% of 600
  after  the shift: by {'s1': 395, 's2': 205, 'human': 0}, wrong among answered alone 37.2% of 600
  replay failures: 0
novel='auto'
  before the shift: by {'s1': 165, 's2': 0, 'human': 435}, wrong among answered alone 0.0% of 165
  after  the shift: by {'s1': 92, 's2': 0, 'human': 508}, wrong among answered alone 0.0% of 92
  replay failures: 0
  open-set gate flag: 680 | signals below 0.59: CUSUM 10.6 ≥ 10.3 (tuned to a share of 0.1; since decision 177)
Enter fullscreen mode Exit fullscreen mode

The plain setup is excellent until the shift (0.5% wrong, nothing for a person) and then quietly wrong on 37.2% of what it answers alone. Nothing crashed and nothing escalated: both System 1 and the slow path put every fraud report on a known team, because a known team is all they can answer. The promise was never meant to cover this — the explain text said "for inputs like the calibration examples" — but a router in production does not get to choose its inputs.

The gated setup keeps 0.0% wrong among what it answers alone, before and after, and flags the change at decision 680, 80 tickets after it began. It pays for that every day, not only on the bad day: before any shift it answers 165 of 600 alone where the plain setup answered all 600 (452 + 148), and the slow path gets nothing, because input the gate holds back goes to a person, not to the slow path. A slow path calibrated on known kinds of ticket says nothing about a new kind.

These are the toy's numbers, not a benchmark. The toy is built to make the effect visible.

The same thing on real data

On solvi's public task stand, Banking77 runs as a stream of 2,000 bank requests where, from request 1,000 on, 40% are about 20 intents the classifier never saw (measured with solvi 0.8.0; the classifier is the baseline's own, with an act head). The promise was "at most 5% wrong among the answers given alone":

before the shift: answered alone, wrong among them after the shift
plain threshold (baseline) 83.7%, 3.9% 65.0%, 22.5% (broken)
solvi's plain System.guarantee 80.8%, 2.5% 58.5%, 16.4% (broken)
OpenSetGate 57.6%, 0.7% 13.7%, 0.7% (kept)

The gate flagged the shift 68 requests after it started. Same shape as the toy: the plain promise breaks without a sound, the gate keeps it by answering much less. The script is benchmarks/tasks/banking77/solution.py and its numbers are rechecked weekly from the stored replies. The recipe for this post repeats the two solvi rows on solvi 1.0.0 with the same numbers.

Bar charts for the Banking77 stream. Answered alone before and after the shift: plain threshold 83.7% then 65.0%, solvi's plain promise 80.8% then 58.5%, open-set gate 57.6% then 13.7%. Wrong among answered alone, against a 5% promise: plain threshold 3.9% then 22.5% (broken), plain promise 2.5% then 16.4%, open-set gate 0.7% then 0.7% (kept). The gate flagged the shift 68 requests after it began.

How the gate works, briefly

It cannot see the future, so it simulates it: build refits System 1 three times, each time without a third of the teams, and uses System 1's signal on those left-out tickets as a stand-in for "a ticket no team fits". It then sizes a threshold for a stream with some share of such tickets, follows that share from the recent signals, and raises a CUSUM flag when it jumps. That only works if the signal tells "outside" from "known" — when it cannot, explain() says so and no gate is built.

What it does not do

  • It does not make anything more accurate. The router is as good as its facts and its slow path. build composes what the library has and writes down every choice.
  • The promise does not hold between an abrupt shift and its detection. The gate shrinks that window; it does not close it. When the flag goes up (s.gate.flag_at), stop answering alone until people have looked and the thresholds are calibrated again.
  • It does not learn the new kind. Labelling fraud reports, adding the option and recalibrating are your job.
  • The gate costs answers in calm times. Pay for it only where new kinds of input are expected. If they come gradually, the guide's track=(200,) costs less; if nothing new can appear, the plain promise is the right one.
  • The knowledge store does not make the router better over time. On streams of one kind of decision, facts added from a person's answers to escalations gave no measurable gain in solvi's own tests; there the store gives accountability (sources, retraction), not growth.

Two more things you get for free

Every decision is a stored, hash-chained record that replays without the slow path: replay failures: 0 above means each of the 1,200 decisions per setup was re-checked from what it was given. And s.report() reads the store back — who answered how often, what it cost, the promise in force. The one-entry-point example in the repo (examples/24_one_entry_point.py) shows the same machinery on refund requests, where the slow path reads free-text notes; on 500 new requests it printed {'s1': 406, 's2': 94, 'human': 0} replay failures: 0.

Where next

Two questions for you, if you route tickets today: would you rather answer fewer alone every day to be safe on the day a new kind of request arrives, or answer more and rely on someone noticing? And what would you want the router to do in the minutes between the shift and the flag?

Top comments (0)