I Gave an Incident Response Agent a Memory With Hindsight
RecallOps is an AI-powered incident response assistant I built to help engineers investigate production incidents using experience from previous incidents.
The frustrating part of incident response isn't always finding a fix. Sometimes, it's realizing that your team has already solved almost the same problem before—and that experience is effectively gone when the next incident starts.
That is the problem RecallOps is designed to address.
When a new incident is reported, RecallOps uses Hindsight to recall relevant historical incidents before generating an analysis. Once an incident is resolved, the engineer can save the incident, root cause, resolution, and outcome back into Hindsight. That experience can then be recalled when a similar problem happens again.
Instead of treating every incident as a completely new problem, RecallOps creates a continuous memory loop:
New Incident
↓
Recall Previous Incidents
↓
AI Analysis
↓
Engineer Investigates
↓
Incident Resolved
↓
Store What We Learned
↓
Future Incident
The rest of the system is built around making that loop useful and transparent to the engineer.
The Problem: Every Incident Starts From Zero
Production incidents tend to repeat themselves.
A database connection pool gets exhausted. A deployment introduces a configuration problem. A service becomes unstable after a change. The exact symptoms may differ, but the underlying causes and solutions often have similarities.
An LLM can analyze the incident in front of it, but that doesn't automatically give it access to the engineering experience accumulated from previous incidents.
That was the problem I wanted RecallOps to address.
The application accepts a production incident from an engineer and sends it to a Node.js/Express backend. Before asking the language model to analyze the incident, the backend searches persistent Hindsight memory for relevant previous incidents.
The retrieved context is then provided to the model alongside the current incident.
The basic flow is:
Engineer
↓
RecallOps
↓
Hindsight Recall
↓
Historical Incident Context
↓
Groq LLM Analysis
↓
Investigation Guidance
↓
Engineer Resolves Incident
↓
Hindsight Retain
↓
Persistent Incident Memory
The important part is the loop at the end.
A resolved incident doesn't disappear after the engineer fixes it. It becomes potential context for the next investigation.
Making Hindsight Part of the Incident Loop
I used Hindsight as the persistent memory layer for RecallOps.
When an engineer resolves an incident, the application collects four pieces of information:
- The incident description
- The root cause
- The resolution
- The outcome
RecallOps combines these into a memory and stores it in Hindsight.
The memory bank is identified in the backend with:
const BANK_ID = "recallops-incidents";
The retain operation then stores the incident experience:
const memory = `
Production incident:
${incident}
Root cause:
${rootCause}
Resolution:
${resolution}
Outcome:
${outcome || "Not specified"}
`;
await hindsight.retain(BANK_ID, memory);
I didn't want to store only the original error message. An error tells us what happened from the application's perspective. A useful engineering memory should also contain what the engineer discovered and what actually fixed the problem.
That makes a resolved incident much more useful when it is retrieved later.
Recall Before Reasoning
The other half of the system happens when a new incident arrives.
RecallOps first sends the incident description to Hindsight:
const memoryResult = await hindsight.recall(
BANK_ID,
incident,
{
budget: "low",
}
);
const memories = memoryResult.results || [];
The returned memories are then converted into context:
const memoryContext = memories.length
? memories
.map(
(memory, index) =>
`Memory ${index + 1}:\n${memory.text}`
)
.join("\n\n")
: "No relevant previous incidents were found.";
Only after that retrieval step does RecallOps ask the LLM to analyze the current incident.
The model receives two distinct pieces of information:
CURRENT INCIDENT:
<what is happening now>
RELEVANT HISTORICAL MEMORIES:
<previous incidents retrieved from Hindsight>
That separation matters.
The current incident is evidence about what is happening now. Historical memories are previous engineering experiences that may or may not be relevant.
A previous incident shouldn't automatically become the answer to a new one. It should give the investigation a useful starting point.
A Concrete Before-and-After
Consider an engineer reporting:
After a deployment, the production API is experiencing intermittent database timeouts and some requests are failing.
Without historical memory, the model has to reason from the current incident alone.
It might consider several possibilities:
Database load
Connection pool exhaustion
Network problems
Configuration changes
Application-level connection handling
Now imagine that the team previously resolved a similar incident.
The stored memory could contain:
Production incident:
Production API experienced intermittent database timeouts.
Root cause:
Database connection pool was exhausted because
connections were not being released correctly.
Resolution:
Connection handling was fixed and the connection pool
size was increased.
Outcome:
API database timeouts stopped.
When the new incident arrives, Hindsight can retrieve that historical experience.
The model can then reason over:
Current incident
+
Relevant historical incident
↓
Context-aware investigation
The historical incident does not prove that the new incident has the same root cause.
It gives the engineer and the model a concrete previous case to investigate against.
Making the Memory Visible
I also wanted the memory layer to be visible rather than completely hidden behind the API.
The RecallOps interface has separate areas for the AI analysis and the Hindsight memories used during that analysis.
After an incident is submitted, the frontend receives the analysis and memory information:
setAnalysis(data.analysis || "No analysis returned.");
setMemories(data.memoriesUsed || []);
setMemoryCount(data.memoryCount || 0);
The memory panel then displays the retrieved memories.
That gives the engineer something concrete to inspect.
Instead of seeing only:
"Here is what you should do."
the interface can also show:
"Here is the historical context that was retrieved."
After the incident is resolved, the interface provides a separate Resolve & Remember workflow where the engineer can enter the root cause, resolution, and outcome.
This creates a simple cycle:
Analyze
↓
Investigate
↓
Resolve
↓
Remember
↓
Future incident
↓
Recall
↓
Investigate again
Where Hindsight Fits
The application is split into a React frontend and a Node.js/Express backend.
The frontend handles the user experience:
- Incident submission
- Analysis display
- Historical memory display
- Root-cause entry
- Resolution entry
- Outcome entry
The backend handles the integrations:
- Hindsight recall
- Hindsight retain
- Groq model requests
- API validation
- Error handling
The architecture is intentionally straightforward:
┌─────────────────┐
│ React Frontend │
└────────┬────────┘
│
▼
┌─────────────────┐
│ Express Backend │
└───────┬─┬───────┘
│ │
┌──────────┘ └──────────┐
▼ ▼
┌──────────────┐ ┌─────────────┐
│ Hindsight │ │ Groq │
│ Recall/Retain│ │ LLM │
└──────────────┘ └─────────────┘
The interesting part isn't a complicated orchestration framework.
It is the feedback loop between incident resolution and future incident analysis.
What Changed Once Memory Became Persistent
Without persistent memory, the interaction is essentially:
Incident → LLM → Response
The next incident starts over.
With persistent incident memory:
Incident
↓
Recall previous experience
↓
LLM analysis with historical context
↓
Engineer investigates and resolves
↓
Store the new experience
↓
Future incident can retrieve it
That changes the role of the application.
RecallOps is no longer only an interface for asking an LLM questions about incidents. It becomes a place where incident experience can accumulate and become available during future investigations.
This is where Hindsight fits into the architecture. Its persistent memory layer provides the retain and recall capabilities that RecallOps uses to turn resolved incidents into reusable context.
The Hindsight documentation describes the underlying memory system, while Vectorize's agent memory overview provides broader context around persistent memory for agents.
What I Learned
1. Store the resolution, not just the error
An incident message tells you what went wrong.
A useful engineering memory should also capture what the team discovered and what actually fixed the problem.
That's why RecallOps stores the incident, root cause, resolution, and outcome together.
2. Memory should provide context, not certainty
A historical incident can be relevant without being identical to the current incident.
RecallOps therefore provides historical memories as context rather than treating them as guaranteed answers.
That distinction is important in incident response.
3. The memory write should happen at a meaningful point
I didn't want every interaction to automatically become permanent memory.
Instead, the engineer explicitly records the root cause, resolution, and outcome after the incident has been investigated.
That gives the stored memory a much clearer meaning: it represents something the team learned from a resolved incident.
4. Retrieved memory should be visible
Showing the retrieved memories alongside the AI analysis makes the system easier to inspect.
An engineer can see the historical context instead of having to trust an unexplained response from the model.
5. The real value of memory appears later
The interesting test isn't whether an agent can remember something immediately.
It is whether information captured today can become useful during a different interaction later.
That's the idea behind RecallOps.
What Comes Next
The current application expects engineers to provide incident information through the interface.
A natural next step would be connecting RecallOps to the systems where incident information already exists: monitoring platforms, logs, traces, deployment infrastructure, or team communication tools.
That would allow the memory loop to begin with real operational signals rather than manually entered incidents.
The core architecture would remain the same:
Observe
↓
Investigate
↓
Resolve
↓
Remember
↓
Recall
↓
Investigate the next incident
The main lesson from building RecallOps is simple:
An incident-response agent doesn't become more useful just because the model gets another prompt.
It becomes more useful when the system can bring relevant engineering experience into the next investigation.
That's the role Hindsight plays in RecallOps: turning resolved incidents into persistent context that can be recalled when similar problems appear again.




Top comments (0)