A fixed escalation threshold is a one-time estimate that isn't revisited. "Three contacts about the same unresolved issue" sounds reasonable until you watch a billing dispute turn ugly on the second email. At the same time, a flaky integration quietly resolves itself on the fourth.
I want the agent to determine where the line should be based on what actually happened to past customers. This is how I used Hindsight to give it that history, and what I deliberately did not let it learn.
**
What the system does**
The project is a customer support memory agent. Customers reach us over chat, email, or phone, and everything is keyed on their email address because it is the one identifier all three channels share.
The backend is a small FastAPI service:
-
hindsight_client.pywraps the memory layer. -
llm_client.pywraps Groq, runningqwen/qwen3-32b. -
memory_schema.pyholds the Pydantic models. -
app.pyexposes endpoints such asPOST /summarise/{email}andPOST /escalate/{email}.
The starting rule is simple: three or more contacts about the same unresolved issue means escalate. That rule is the floor of the system, not the whole of it.
The problem with a constant
A threshold of three treats every kind of issue the same. It isn't. A payment problem carries urgency and emotion that a cosmetic bug does not, so the third contact on a billing issue is often very late, while the third contact on a known workaround is routine.
I could have hard-coded a threshold per category. That just moves the guess into a lookup table I would then have to maintain by hand. What I actually wanted was for the agent's own history to say where each category tips over.
That requires something a stateless agent doesn't have: a durable record of what happened after each decision.
What I log
Every time the agent makes an escalation decision, I write a record to Hindsight. The interesting part is that I write it twice: once at the decision, and again when the outcome is known.
from pydantic import BaseModel
from typing import Literal, Optional
class EscalationRecord(BaseModel):
email: str
issue_category: str # e.g. "billing", "api_integration"
contact_count: int # contacts at the moment of decision
channels: list[str]
decision: Literal["escalated", "held"]
outcome: Optional[Literal["resolved", "recontacted", "churn_risk"]] = None
The outcome is what turns a log into a lesson. A record that says "held at two contacts" is only useful once I know the customer came back a third time, angrier. A record that says "escalated at three" is only useful once I know whether that escalation resolved it.
Writing and updating these goes through the same thin wrapper as everything else:
def retain_escalation(record: EscalationRecord) -> None:
hindsight.retain(subject=f"escalations:{record.issue_category}",
content=record.model_dump())
def recall_escalations(issue_category: str) -> list[EscalationRecord]:
raw = hindsight.recall(subject=f"escalations:{issue_category}")
return [EscalationRecord(**r) for r in raw]
Note the subject. Customer interactions are keyed on the customer's email. Escalation records are keyed on the issue category, because the lesson I want belongs to the category, not to any one person.
How the agent uses it
Before deciding, the agent recalls the past records for the issue's category and asks one narrow question: of the cases we held at N contacts, how many came back?
def learned_threshold(category: str, default: int = 3,
floor: int = 2, ceiling: int = 4) -> int:
history = recall_escalations(category)
held_then_recontacted = [
r for r in history
if r.decision == "held" and r.outcome == "recontacted"
]
if len(held_then_recontacted) >= 3:
lowest = min(r.contact_count for r in held_then_recontacted)
return max(floor, min(ceiling, lowest))
return default
Three things about this function are deliberate.
It's arithmetic, not a model call. The language model never sees the records and never proposes a threshold. Hindsight supplies the history, Python does the counting, and the result is a number I can log and test.
It's bounded. The threshold can move between two and four, never outside. If the history is strange or thin, the worst case is a slightly different number, not an agent that escalates everyone or no one.
It needs evidence. With fewer than three qualifying records the function returns the default of three. A category with no history behaves exactly like the original rule, so the system degrades to the fixed threshold instead of guessing.
The escalate endpoint combines the two ideas:
@app.post("/escalate/{email}")
def escalate(email: str, issue_category: str):
interactions = recall_interactions(email)
threshold = learned_threshold(issue_category)
decision = needs_escalation(interactions, threshold=threshold)
note = llm.write_handoff_note(escalate=decision, history=interactions)
retain_escalation(EscalationRecord(
email=email, issue_category=issue_category,
contact_count=len(interactions), channels=[i.channel for i in interactions],
decision="escalated" if decision else "held",
))
return {"escalate": decision, "threshold_used": threshold, "note": note}
The response includes threshold_used. When the bot escalates at two contacts, I can see that it did so because the billing category's history moved the line, not because a model felt uneasy.
What it looks like in behaviour
Take the seed customer with four interactions across chat, email, and phone about an unresolved billing problem. Under the default rule she escalates on the third contact, and the fourth is already overdue.
Now suppose the log holds several billing cases that were held at two contacts and came back. The learned threshold for billing drops to two, and the next customer with that pattern is escalated a contact earlier. A customer with a resolved bug and a workaround is unaffected, because nothing about their category has moved.
I am describing designed behaviour here, not measured improvement. I have not benchmarked whether earlier escalation improves resolution, and a claim like that would need real cases over real time.
Lessons learned
Log the outcome, not just the decision. A decision without its result teaches nothing. Most of the value came from the second write.
Key the lesson to the thing it is about. Customer history belongs to a customer. Escalation lessons belong to a category. Using the wrong subject would have mixed them up.
Keep learning out of the model. Let memory supply history and let code compute the adjustment. It stays testable and explainable.
Bound everything that adapts. A floor, a ceiling, and a minimum evidence count are what make a learning system safe to leave running.
Expect the cold start to be honest and boring. With no history, the system uses the fixed rule. It only gets more opinionated as evidence accumulates, and that is the right order.
The limitation I care about most is delay. Outcomes arrive days after decisions, so the threshold reflects the past, and a category whose behaviour changes will lag. Handling that, probably by weighting recent records more heavily, is next on my list.
If you are building an agent that should act differently because of what happened before, Vectorize's explanation of agent memory is a good place to understand why this needs memory rather than a longer prompt.**
Top comments (0)