The alert looks familiar. The Payment API is timing out against its database, and I have a nagging feeling someone fixed this before. But that fix is in a closed ticket, a Slack thread, or one colleague's head, so the on-call engineer starts from scratch.
That gap is why our team built RecallOps, an AI incident-response copilot. This article covers the investigation workflow: how it uses past incidents, how it treats uncertainty, and where the engineer stays in charge. I'm one of five people who built it, so when I describe a component, I'm describing the system our team built, not claiming I wrote every part.
The incident: INC-017
Here is the scenario from our demo data. INC-017 is a Payment API database timeout. The current evidence is:
- 98% connection utilization
- a 14.2% timeout rate
- a deployment 23 minutes earlier
- a connection timeout error
Any one of these could have several explanations. The deployment might be a coincidence. High utilization might just be heavy traffic. The useful question isn't "what is the root cause?" but "have we seen this shape before?"
Screenshot 1: The incident workspace for INC-017 during investigation. It shows the starting point: the current evidence next to the recalled historical memory, so you can see what an engineer sees when the investigation begins.
Why historical context matters
Experienced engineers investigate faster because they pattern-match against incidents they've lived through. That knowledge usually stays personal. We wanted a memory layer with a clear contract, not a pile of old tickets stuffed into a prompt.
Memory goes through Hindsight, which gives us two operations: retain, which stores what we learned from an incident, and recall, which brings back what's relevant to a new one.
Recalling INC-001
When INC-017 is investigated, the backend exposes a recall endpoint. This is a simplified, representative example:
@router.get("/incidents/{incident_id}/recall")
def recall_memory(incident_id: str, db: Session = Depends(get_db)):
incident = get_incident_or_404(db, incident_id)
try:
matches = memory.recall(incident)
degraded = False
except MemoryUnavailable:
matches = db_fallback_recall(db, incident)
degraded = True
return {
"matches": matches,
"degraded": degraded,
}
For INC-017, recall returns INC-001 as the top match. INC-001 was also a Payment API database timeout. Its root cause was connection-pool exhaustion from a connection leak, and the resolution was to fix the leak and increase pool capacity. It had been resolved and retained, which is why it was available to recall.
What "91% similar" means
The match comes back at 91% similarity, and the response carries the reasons with it: same service, same service family, similar symptoms, similar database behavior, and similar timing.
I care more about the reasons than the number. A bare score asks the engineer to trust it. A score with reasons lets them check whether the match makes sense. We haven't benchmarked the system, so I'm not making claims about how reliable similarity scores are in general. I'm describing what the demo shows and how it's presented.
Screenshot 2: INC-001 as historical memory at 91% similarity. It shows the match reasons and the historical root cause displayed together, demonstrating that the score is explained rather than shown alone.
Current evidence vs. historical evidence
The system keeps four things separate: current evidence, historical evidence, the AI recommendation, and uncertainty. For INC-017 the two evidence sets look like this:
| Current (INC-017) | Historical (INC-001) |
|---|---|
| 98% connection utilization | High connection utilization |
| 14.2% timeout rate | Same service and error family |
| Deployment 23 minutes ago | Confirmed connection leak |
| Connection timeout |
The overlap is real: high utilization, the same service, the same error family. But look at what's on each side. INC-001's leak was confirmed. For INC-017, nothing has confirmed a leak yet.
Keeping the sources separate means a recommendation can be traced back to specific past incidents, and each side can be reviewed and debugged on its own. Blending them into one blob would make the output harder to trust.
Screenshot 3: Current evidence and historical evidence side by side. It shows the separation of sources and lets you see exactly which observations overlap and which don't.
How the old root cause guides the investigation
INC-001's connection leak becomes a lead, not a conclusion. The workspace offers structured investigation paths, and one of them fits the deployment 23 minutes before the alert: check whether the recent deployment introduced a new leak.
The history narrows where to look first. It doesn't say what the answer is.
Recommendation vs. certainty
A 91% match is worth investigating, but a similar past incident is not proof of the same root cause. Two incidents can look alike and fail for different reasons.
So the UI labels INC-001 as evidence for investigation, not confirmation of the current root cause. The system also surfaces uncertainty as its own element instead of folding it into a confident-sounding recommendation. If a copilot pretends to know, engineers either over-trust it or learn to ignore it, and neither helps at 3 a.m.
When memory is unavailable
We couldn't assume a remote memory service would always be reachable. In the recall snippet above, if Hindsight fails, the endpoint catches MemoryUnavailable and recalls from the system's own persisted incident data through db_fallback_recall.
The important detail is the degraded flag. The application keeps working, but the response says a fallback was used, so the UI can show a degraded state instead of silently passing off the fallback as the primary memory provider. Designing this early forced us to decide what "degraded" should look like to the person investigating.
Closing the loop, and the engineer's role
Once the engineer resolves INC-017, it can be retained so future incidents can recall it:
if incident.status != "resolved":
raise HTTPException(400, "Resolve the incident before retaining it")
memory.retain(incident)
incident.retained = True
The check keeps unfinished investigations out of memory.
RecallOps never changes production systems automatically. It provides evidence, history, recommendations, investigation paths, and uncertainty. The engineer makes the final call.
After retention, a Learning area looks across retained incidents. On the demo data it identified a recurring Payment API pattern across five related incidents, with provenance showing which incidents each lesson came from.
Screenshot 4: The Learning page. It shows the payoff of retention: a recurring pattern across incidents, each lesson traced back to its source incidents.
What was and wasn't verified
Verified: backend startup, API health, backend tests, SQLite persistence, memory recall, retention, learning/reflection, and the browser end-to-end workflow (Launch Demo through INC-001, INC-017, recall, resolve, retain, Learning, and demo reset). The verified demo runs on SQLite with the deterministic AI fallback.
Not live-verified: a production PostgreSQL deployment, a remote Hindsight service, and live Groq/OpenAI inference. They're configured integrations in the architecture, but we haven't tested them live, so I make no claims about them.
What this taught me
- Memory should explain itself. Match reasons made the recall result something I could evaluate, not just accept.
- A similar incident is a lead, not a verdict. Separating evidence from recommendation from uncertainty keeps the tool honest.
- Degraded modes are part of the design. Deciding early how failure looks made the whole system more trustworthy.
- The engineer stays the decision-maker. The tool's value is shortening the path to the right question, not answering it for you.
Our next step is validating the remote integrations so the configured architecture gets tested too. If you're interested in agent memory, the Hindsight repository is a good place to start.




Top comments (0)