DEV Community

subani mohammad
subani mohammad

Posted on

A support bot that remembers: cross-channel history and escalation with Hindsight

Picture a customer with a billing problem. They start in chat and get a suggestion. A few days later they email, because the suggestion didn't work. Then they phone, and the agent asks them to explain the whole thing from the beginning. Four contacts in, nobody has fixed anything, and nobody has noticed the pattern.
That is the situation I built this bot for. It has two jobs: make sure a customer never has to repeat themselves when they switch channels, and flag the moment a case has been open too long to keep handling the usual way. Memory comes from Hindsight, the API from FastAPI, and the language model from Groq.
One customer, three channels
Chat, email, and phone all land in different systems, so the first question is how to recognise that they belong to the same person. I keyed everything on the customer's email address. It's the one identifier all three channels reliably carry, and choosing it early meant the rest of the design fell out naturally. A chat session ID, for example, would have quietly split each customer's history by channel.
The service stays small:
GET /customers and GET /customers/{email} return customers and their interactions.
POST /summarise/{email} produces a short brief of the full history.
POST /escalate/{email} decides whether a senior agent should take over.
Under the hood there are four files, one job each: hindsight_client.py for memory, llm_client.py for Groq (qwen/qwen3-32b), memory_schema.py for the Pydantic models, and app.py to connect them.
Reading the history
Hindsight is the reason a summary can be written at all. Each interaction is stored against the customer's email, and when an agent needs context the bot recalls it and asks the model for a brief. The agent picking up the phone sees what was tried in chat and email before saying hello.
I wrapped the memory layer in two tiny functions so the rest of the app doesn't care how retrieval works:
def recall_interactions(email: str) -> list[Interaction]:
raw = hindsight.recall(subject=email) # everything remembered for this customer
return [Interaction(**item) for item in raw]

def retain_interaction(interaction: Interaction) -> None:
hindsight.retain(subject=interaction.email, content=interaction.model_dump())
For background on why this kind of memory is different from replaying chat logs, Vectorize's write-up on agent memory is a good read.
Knowing when to stop
Summaries are useful. Escalation is where the bot earns its keep, because it has to decide that another attempt from the usual playbook will make things worse.
The policy is: three or more contacts about the same unresolved issue means escalate.
I was tempted to hand the model the entire history and ask whether to escalate. It would have produced a confident answer, but I couldn't have tested that answer or explained it afterwards. So I split the work three ways:
Layer
Responsibility
Hindsight
Holds the facts: who contacted us, when, about what
Python
Applies the threshold
The LLM
Writes the handoff note for a human
The stored record carries only what the rule needs:
from pydantic import BaseModel
from typing import Literal

class Interaction(BaseModel):
email: str
channel: Literal["chat", "email", "phone"]
issue_id: str # which issue this contact belongs to
summary: str
resolved: bool
The check itself is a handful of lines:
from collections import Counter

def needs_escalation(interactions: list[Interaction], threshold: int = 3) -> bool:
open_counts = Counter(i.issue_id for i in interactions if not i.resolved)
return any(n >= threshold for n in open_counts.values())
And the endpoint gives the model a decision to explain rather than a decision to make:
@app.post("/escalate/{email}")
def escalate(email: str):
interactions = recall_interactions(email)
decision = needs_escalation(interactions)
note = llm.write_handoff_note(escalate=decision, history=interactions)
return {"escalate": decision, "note": note}
Trying it on awkward customers
I tested with a handful of customers chosen to stress different parts of the design:
The billing case. Four interactions across chat, email, and phone, all about one problem that never got fixed. The bot returns escalate: true, with a note summarising what was tried on each channel.
The resolved bug. Reported, then closed with a workaround. Nothing is open, so the bot stays quiet.
The smooth journey. No trouble, no escalation.
The API integration customer. Several related problems compounding over time. Whether this escalates depends on whether those contacts count as one issue or several, and that's the place the design is weakest.
The best property of the result is that it can be explained. When the bot escalates, I can point at the exact records that caused it.
What I'd tell someone building the same thing
Start from the join key. Once email was the identity, cross-channel continuity was almost free.
Treat escalation as a memory problem. The signal is "how many times, and about what," and that only exists if earlier contacts were remembered.
Keep thresholds out of prompts. Let the model write and summarise; let code decide.
Test the messy customers. Easy cases prove the bot is quiet. Messy ones show you where it's wrong.
Be honest about the fuzzy part. Deciding whether two contacts are the same issue is still a judgement call. I've confined it to one place I can inspect, and tightening it is my next task.
If you're building something that needs to behave differently because of what happened before, start with the Hindsight docs.

Top comments (0)