How Hindsight Let My Incident Copilot Recall a 91% Match
When INC-017 fired, a Payment API database timeout, the workspace surfaced a resolved incident from months earlier as the top match before anyone had opened a log. It scored 91% similarity and came with a written explanation of why.
Incident response is a race against missing context. An on-call engineer usually has plenty of live data, such as dashboards, logs and error rates, but almost no fast way to ask "has this happened before, and how was it fixed?" Closing that gap is what we built RecallOps for. I worked on the frontend and the investigation experience, so this article covers how Hindsight, an open-source agent memory system, and the interface on top of it turn a wall of alerts into a workspace where past incidents appear exactly when they're useful.
What the system does
RecallOps is organized around seven areas: Dashboard, Incidents, Incident Workspace, Copilot, Memory, Learning and Services. The Incident Workspace is where investigation happens. From that one screen an engineer can:
read the current incident
trigger AI analysis
pull up historical memory
see why a past incident was judged similar
compare current and historical evidence side by side
work through structured investigation paths
resolve the incident
Underneath, every resolved incident is stored as memory, and every new incident queries that memory. The interface's job is to make the recalled history easy to evaluate rather than easy to blindly trust.
[Screenshot 1: the full Incident Workspace for INC-017, showing the incident summary, AI investigation console and operational memory panel together.]
The Incident Workspace lays out live evidence, AI analysis and recalled memory in one screen.
The core problem: an assistant with no history
The instinct when something breaks is to dig through metrics and logs from scratch, even if a teammate solved the same problem months ago. An AI assistant without agent memory hits the same wall. It can read the current symptoms but has no sense of your system's history, so every incident looks like the first one.
The interface had to solve for one specific moment: the few minutes after an incident opens, when an engineer needs relevant history without hunting for it.
How Hindsight fits in
Hindsight sits between the incident data and the workspace. We use two operations: retain when an incident is resolved, and recall when a new one opens. The Hindsight documentation covers both in detail. Here is a simplified version of our integration:
from hindsight_client import Hindsight
client = Hindsight(base_url=HINDSIGHT_URL)
# Retain: store a resolved incident with its root cause and resolution
client.retain(
bank_id="recallops",
content=(
f"{incident.id}: {incident.service} {incident.summary}. "
f"Root cause: {incident.root_cause}. "
f"Resolution: {incident.resolution}."
),
)
# Recall: look up similar past incidents when a new one opens
memories = client.recall(
bank_id="recallops",
query=f"{incident.service} {incident.symptoms}",
)
INC-001 was a Payment API database timeout caused by connection-pool exhaustion from a connection leak. It was resolved by fixing the leak and increasing pool capacity, then retained. When INC-017 arrived, another Payment API database timeout, recall returned INC-001 as the top historical match at 91% similarity.
Retention matters as much as recall. Because each incident is stored with its root cause and resolution, not just its symptoms, the recalled memory carries the answer to "how was it fixed?" and not only "did this happen before?"
Making the match explainable
The 91% is not the important part. Next to it, the workspace shows a "Why this memory?" list: same service, same service family, similar symptoms, similar database behavior, similar timing. I cared more about that list than the percentage. A bare score says "trust this." The reasons let the engineer evaluate the match instead of taking it on faith.
[Screenshot 2: the Operational Memory card for INC-017, showing the 91% score and the "Why this memory?" checklist.]
INC-001 recalled at 91% similarity, with the specific reasons behind the match.
Keeping current and historical evidence separate
The decision that shaped the most frontend work was refusing to blend live and historical data into one summary. The workspace keeps four things visibly distinct: current evidence, historical evidence, the AI recommendation and uncertainty.
For INC-017, current evidence was:
98% connection utilization
a 14.2% timeout rate
a deployment 23 minutes earlier
a connection-timeout error family
Historical evidence from INC-001 was:
96% connection utilization
the same service and error family
a confirmed connection leak as root cause
Showing these in two labeled panels, instead of one merged narrative, makes a recommendation much easier to sanity-check. You can see which numbers came from now and which came from before.
[Screenshot 3: the side-by-side Current Evidence and Historical Evidence panels for INC-017.]
Live metrics on the left, INC-001's data on the right. Nothing is merged into a single blob.
Recommendation and uncertainty
The overlap between INC-017 and INC-001 is useful, but it isn't proof of a shared root cause. RecallOps treats the match as evidence to investigate, not a conclusion to accept, and the UI keeps uncertainty visible right next to the AI's recommendation. Structured investigation paths let the engineer test the hypothesis directly, for example by checking whether the recent deployment introduced a new leak.
This was the hardest part to get right. It's tempting to make a 91% match feel like an answer. But an interface that oversells confidence teaches engineers to stop checking their own work, which is the opposite of what you want during an incident.
No automatic changes to production
RecallOps never touches production systems on its own. It surfaces evidence, history, recommendations, investigation paths and uncertainty, and the engineer makes the final call. That constraint pushed the design toward decision support rather than automation, which is why so much of the layout is built around comparison instead of a single "do this" button.
Retain and reflect
Once an engineer resolves INC-017 and retains it, it becomes memory for the next incident. The Learning area then looks across retained incidents for recurring patterns. On our sample incidents it surfaces a recurring Payment API pattern across five related incidents, with the evidence trail behind each synthesized lesson, so an engineer can trace any lesson back to the incidents that produced it.
[Screenshot 4: the Learning area showing the recurring Payment API pattern across five incidents.]
The Learning view traces a recurring pattern back to the incidents that produced it.
What didn't work
My first version showed only the recalled incident and its score. Testers, and I, started treating the recall as the diagnosis. In INC-017's case that would have meant assuming an old connection leak without checking whether the new deployment was the trigger. Adding the reasons list, the evidence panels and the visible uncertainty was a fix for that mistake, not a polish pass.
Lessons learned
Separating evidence sources is a UI problem, not just a data problem. The backend keeping current and historical data distinct isn't enough. The interface has to keep reinforcing that separation, or engineers will treat a recalled memory as confirmed fact.
A similarity score needs a reason attached. A bare percentage invites blind trust. A "why this memory?" breakdown invites verification.
Uncertainty is a feature, not a gap. Showing what the system doesn't know is what makes the parts it does show trustworthy.
Store the resolution, not just the symptoms. Memory is only useful in an incident if it carries what fixed the problem.
Keep the human in the loop. Recall makes an engineer faster, but it shouldn't make the decision.
The core interaction is the part I'm most interested in continuing to refine: recall a similar past incident, show why it matched, compare it against what's happening now, and let the engineer decide.




Top comments (0)