DEV Community

Cover image for How I used Hindsight to trigger strict support escalations
ALUVALA REVANTH
ALUVALA REVANTH

Posted on

How I used Hindsight to trigger strict support escalations

How I used Hindsight to trigger strict support escalations

Every support team has a version of the same argument. One person says we escalate too late and customers are furious. Another says we escalate too often and the senior queue is drowning. Both are right, because "escalate when it seems bad" is not a rule, it is a mood.

I wanted a rule I could state in one sentence, defend in a review, and test in CI. This is how I built one with Hindsight and FastAPI.

The rule

A case escalates when a customer has contacted us three or more times about the same unresolved issue.

That sentence hides three separate questions, and each one needs a different kind of answer:

  1. Who is the customer? Identity, across chat, email, and phone.
  2. Which contacts are about the same issue? A judgement call.
  3. Is it still unresolved, and how many times? Arithmetic.

Most of the trouble I had came from mixing those together. The rest of this post is about pulling them apart.

What the system looks like

The backend is a small FastAPI service. hindsight_client.py wraps the memory layer, llm_client.py wraps Groq running qwen/qwen3-32b, memory_schema.py holds the Pydantic models, and app.py exposes endpoints such as POST /escalate/{email} and POST /summarise/{email}.

Customers are keyed on their email address, because it is the one identifier that all three channels reliably carry. Every interaction is stored against that email in Hindsight, and every escalation check starts by recalling them.

Why "strict" matters

A model asked "should we escalate this customer?" will answer fluently and mostly sensibly. Mostly is the problem. Run the same history twice and the answer can move. Change the wording of one message and it can move again. I cannot write a test for a decision that drifts, and I cannot explain to a support lead why customer A was escalated and customer B, with a nearly identical history, was not.

Strict, for me, means three properties:

  • Deterministic. The same history produces the same decision, every time.
  • Explainable. I can point at the exact records that triggered it.
  • Non-overridable by the model. The language model can describe a decision. It cannot make one.

Memory is what makes those properties possible, because a rule can only be strict about facts that were recorded. That is the job agent memory does here: it turns "what happened" from something a model reconstructs into something the system retrieves.

The facts I store

Each interaction is recorded with just enough structure for the rule to use:

from pydantic import BaseModel
from typing import Literal

class Interaction(BaseModel):
    email: str
    channel: Literal["chat", "email", "phone"]
    issue_id: str        # which issue this contact belongs to
    summary: str
    resolved: bool
Enter fullscreen mode Exit fullscreen mode

The issue_id field is where the judgement call from question two lives. I decide it once, when the interaction is stored, and after that it is a fact like any other. Deciding "same issue or not" at write time means I am not re-litigating it on every read.

The rule as code

With facts in hand, the check is short enough to read in one glance:

from collections import Counter

def needs_escalation(interactions: list[Interaction], threshold: int = 3) -> bool:
    open_counts = Counter(i.issue_id for i in interactions if not i.resolved)
    return any(n >= threshold for n in open_counts.values())
Enter fullscreen mode Exit fullscreen mode

Two details carry a lot of weight:

  • Resolved contacts do not count. A customer who had a problem, got it fixed, and later reports something new should not inherit the first problem's history.
  • The count is per issue, not per customer. Three contacts about three different things is a busy customer, not a failing case.

The endpoint then keeps the model on the communication side of the line:

@app.post("/escalate/{email}")
def escalate(email: str):
    interactions = recall_interactions(email)
    decision = needs_escalation(interactions)
    note = llm.write_handoff_note(escalate=decision, history=interactions)
    return {"escalate": decision, "note": note}
Enter fullscreen mode Exit fullscreen mode

The decision is computed before the model is called, and the model receives it as input. There is no path where the model's output changes decision.

How it behaves

My test data has five synthetic customers. The stress case has four interactions across chat, email, and phone, all about a billing problem that never got resolved. It escalates, and the handoff note summarises what was tried on each channel. A customer whose bug was reported and resolved with a workaround does not escalate, because nothing is open. A customer with a smooth plan upgrade does not escalate either.

The cases I care about most are the ones that should stay quiet. A strict rule is only useful if it has a false-positive story as good as its true-positive one.

Lessons learned

Write the rule as a sentence first. If you cannot say it in one sentence, you cannot enforce it. Mine took three tries.

Separate judgement from arithmetic. Let the model do the fuzzy part once, store the result, and count in code.

Choose your error deliberately. A threshold of three means a customer can be frustrated twice before anyone senior sees the case. That was a conscious trade against flooding the senior queue, and I would rather state it than pretend the threshold is neutral.

Make the model unable to override the decision. Not "unlikely to". Unable.

Know where the fuzziness lives. Assigning issue_id is still a model call, and it can still be wrong. A rephrased complaint that gets a new issue_id will slip past the threshold. Tightening that is the next thing I want to fix.

If you want a memory layer to hold the facts while you build rules like this, the Hindsight docs are the place to start.

Top comments (0)