Production incidents rarely happen in isolation.
A payment API becomes slow. A Redis connection pool gets exhausted. A deployment introduces unexpected errors. Someone investigates the issue, finds the fix, resolves the incident, and moves on.
But when a similar incident happens again, a new investigation often starts from scratch.
That was the problem I wanted to solve with MemoryOps: an AI incident response agent that can investigate an incident, recall what happened in similar incidents, recommend an action based on that experience, and retain the new resolution for future incidents.
The key part is not just the AI model.
It is memory.
The Problem: Incident Knowledge Gets Lost
A typical incident response workflow has access to a lot of information:
- Current incident details
- Application logs
- Metrics
- Deployment information
- Runbooks
- Previous incident reports
The difficult part is connecting the current incident with relevant historical experience.
A language model can reason about the information given to it, but I wanted the agent to do something more useful:
"Have we seen something like this before, and what worked last time?"
That is where Hindsight became the memory layer of the system.
The Architecture
I designed MemoryOps around a simple separation of responsibilities.
text
┌─────────────────────┐
│ React Frontend │
│ Incident Dashboard │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ FastAPI API │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ Incident Response │
│ Agent │
└───────┬─────┬───────┘
│ │
┌─────────┘ └──────────┐
▼ ▼
┌───────────────┐ ┌───────────────┐
│ LLM │ │ Hindsight │
│ Reasoning │ │ Long-Term Mem │
└───────────────┘ └───────────────┘
│
▼
┌─────────────────┐
│ PostgreSQL │
│ Structured Data │
└─────────────────┘
PostgreSQL stores structured application information such as incident IDs, services, severity, status, logs, metrics, and deployment information.
Hindsight has a different responsibility.
It stores and recalls operational experience.
This distinction became important during development. I did not want to use the memory system as another traditional database. I wanted it to answer questions such as:
"What happened in previous incidents that looked similar to this one?"
The Memory Loop
The core workflow is:
Incident
↓
Investigate
↓
Recall previous experience
↓
Combine current evidence + historical memory
↓
Generate recommendation
↓
Human reviews the recommendation
↓
Resolve incident
↓
Retain the outcome
↓
Future incidents can recall it
This creates a learning loop.

The agent does not simply answer a question and forget the interaction.
A resolved incident becomes useful information for a future investigation.
Using Hindsight for Incident Recall
The agent creates a query from the current incident.
For example:
query = f"""Incident: {incident.summary}Service: {incident.service}Logs:{incident.logs}Metrics:{incident.metrics}Deployment:{incident.deployment}"""memories = hindsight.recall_incidents(query)
The memory service sends this information to Hindsight and retrieves relevant previous experiences.
The agent then includes those memories when asking the LLM to analyze the incident.
The important part is that the model is not reasoning only from the current incident.
It is reasoning from:
Current Incident
+
Current System Evidence
+
Relevant Historical Experience
↓
Analysis
That changes the behavior of the system considerably.
An Example
Suppose the current incident is a critical authentication-service failure.
The agent receives information such as:
Service: auth-service
Severity: CRITICAL
Symptoms:
HTTP 500 responses increased after deployment.
Metrics:
Error rate increased significantly.
Logs:
Connection pool exhaustion detected.
The agent can then recall previous incidents.
For example, one historical incident involved a payment service where a Redis connection pool became exhausted. The resolution involved increasing the pool size and verifying the service after the change.
The historical incident does not automatically become the answer.
Instead, it becomes evidence that the agent can consider.
The LLM combines that evidence with the current incident data and generates a structured response containing:
- Root cause
- Confidence
- Recommended actions
- Reasoning
- Retrieved memories
This keeps the decision process understandable instead of simply returning a single opaque answer.
Turning Resolutions Into Memory
The other half of the system is just as important.
After an incident is resolved, the outcome is retained in Hindsight.
memory_text = f"""Incident {incident_id} was resolved.Service: {incident.service}Severity: {incident.severity}Summary:{incident.summary}Logs:{incident.logs}Metrics:{incident.metrics}Deployment:{incident.deployment}Resolution:{resolution}Outcome:Incident resolved successfully."""hindsight.retain( bank_id=bank_id, content=memory_text, context="Production incident response experience", document_id=incident_id,)
Now the resolution becomes part of the agent's future experience.
This is the part I found most interesting.
The system is not just retrieving documentation.
It is building a collection of operational experiences from previous incidents.
A Small Engineering Problem I Ran Into
One of the more useful lessons came from something that initially looked unrelated to memory.
The LLM was generating slightly different representations for confidence.
At one point, the model returned:
{
"confidence": "high"
}
while the API expected a floating-point value.
Later, another response returned:
{
"confidence": 0.9
}
while the schema had temporarily been changed to expect a string.
The API validation failed because the generated output and the Pydantic schema disagreed.
I fixed this by normalizing the model output before returning it from the agent.
For example:
confidence_map = { "very low": 0.2, "low": 0.4, "medium": 0.6, "high": 0.9, "very high": 0.95,}
The final API schema uses:
class InvestigationOut(BaseModel): incident_id: str root_cause: str confidence: float recommendation: list[str] reasoning: str memories: list[dict]
This was a good reminder that LLM applications need strong boundaries between probabilistic model output and deterministic application code.
Why Human Approval Matters
I deliberately did not design the agent to automatically execute arbitrary production commands.
The agent investigates the incident and recommends an action.
A human can review the recommendation before anything operational is changed.
That gives the system this workflow:
AI investigates
↓
AI recommends
↓
Human reviews
↓
Human approves
↓
Resolution
↓
Outcome becomes memory
For an incident-response system, this separation is important because a recommendation and an actual production change are two different things.
What I Learned
Building MemoryOps changed how I think about memory in AI agents.
1. Memory should have a purpose
Adding a memory system does not automatically make an agent intelligent.
The memory needs to answer a useful question.
For this system, the question is:
"What previous incident experience is relevant to this incident?"
2. Structured Data and Agent Memory Are Different
PostgreSQL is useful for structured application state.
Hindsight is useful for remembering and recalling experiences.
Keeping those responsibilities separate made the architecture easier to reason about.
3. Retrieval Is Only Useful When It Changes the Decision
The goal is not to retrieve the largest number of memories.
The goal is to retrieve memories that provide useful context for the current investigation.
4. LLM Output Needs Validation
Even when the model is instructed to return JSON, application code should still validate and normalize the result.
The confidence-field issue was a small example of why this matters.
5. The Interesting Part Is the Learning Loop
The most useful property of the system is not simply that it can investigate an incident.
It is that:
Incident → Investigation → Resolution → Memory
↑
│
Future Recall
Every resolved incident can potentially make future investigations more informed.
What's Next
There are several areas I would explore next:
- Connecting the agent to real observability platforms
- Adding richer incident timelines
- Improving memory retrieval and filtering
- Tracking whether recommended actions actually solved incidents
- Adding stronger evaluation for retrieved memories
- Supporting multiple services and dependency relationships
- Adding more detailed approval and audit workflows
The current implementation is intentionally focused on one core idea: giving an incident-response agent persistent operational memory.
Final Thoughts
AI agents are often described in terms of reasoning, tools, and autonomous actions.
I think memory deserves the same level of attention.
An agent that can investigate today's incident is useful.
An agent that can remember what happened yesterday, understand why a previous solution worked, and use that experience when investigating tomorrow's incident is a different kind of system.
That is what I wanted to explore with MemoryOps.
Hindsight provided the persistent memory layer that made this possible.
The project is available on GitHub:
https://github.com/sriamsatwik2005/memoryops-ai-incident-response
Hindsight:
https://github.com/vectorize-io/hindsight
Hindsight documentation:
https://hindsight.vectorize.io/****
Top comments (0)