I Built an Incident Responder That Remembers What Happened Before
Production incidents are rarely completely new. A database may hit its connection limit again. An API may start timing out after a traffic spike. A deployment may introduce the same kind of failure that the engineering team has already experienced. The problem is that traditional incident-response workflows often force engineers to rediscover this knowledge from scratch.
That was the problem I wanted to solve with HindsightOps: an AI Incident Response Agent that can remember previous incidents, retrieve relevant operational knowledge, and use that experience when investigating the next incident.
The central idea is not simply to give an AI access to more documentation. It is to give the agent persistent operational memory.
HindsightOps is built around a simple learning loop:
New incident → Recall relevant history → Investigate → Recommend actions → Resolve → Retain the outcome → Use that learning in future incidents.
The project uses Hindsight as the persistent memory layer, with Groq providing the LLM layer. The rest of the application is built with React, Vite, Tailwind CSS, FastAPI, SQLAlchemy, Pydantic, and SQLite.
Why incident response needs memory
When an incident happens, an engineer usually needs to answer several questions quickly:
What is happening?
Has this happened before?
What caused it previously?
What was done to fix it?
Which runbook should be followed?
Did the previous solution actually work?
A stateless AI assistant can analyze the current symptoms and generate possible explanations. However, it does not automatically have the operational experience of the particular system it is helping with.
For example, suppose a Payment API starts returning database connection errors. A generic AI might suggest checking database capacity, connection pools, traffic, and application configuration.
Those suggestions may be reasonable, but they are still generic.
Now imagine that the engineering team previously experienced the same problem and discovered that the connection pool was exhausting during traffic surges. They increased the pool size from 50 to 100 using Runbook DB-04, and the incident recovered in 12 minutes.
That previous experience is much more useful than generic advice.
This is where Hindsight becomes important in HindsightOps.
Making Hindsight the operational memory
I integrated Hindsight through its Python SDK and created a dedicated memory bank:
hindsightops-incidents
When an incident is resolved, the application can retain a structured memory unit containing information such as the incident ID, service, severity, error, root cause, resolution, recovery time, runbook, and lessons learned.
For example, this is the actual style of memory retention used in the project:
from hindsight_client import Hindsight
client = Hindsight(
base_url=settings.HINDSIGHT_BASE_URL,
api_key=settings.HINDSIGHT_API_KEY,
timeout=30.0
)
client.retain(
bank_id="hindsightops-incidents",
content="""[INCIDENT MEMORY UNIT]
Incident ID: INC-1042
Service: Payment API
Error: maximum database connections reached
Root Cause: Connection pool exhaustion under surge
Resolution: Increased pool size from 50 to 100 via Runbook DB-04
Recovery Time: 12 minutes
Key Lessons: Connection pool ceiling must dynamically scale with pod replica count""",
metadata={
"incident_id": "INC-1042",
"service": "Payment API",
"severity": "CRITICAL",
"runbook": "DB-04"
},
tags=["incident", "service:payment-api", "runbook:db-04", "postgres"]
)
The important part here is that the application is not just storing an arbitrary text note.
The incident is represented as operational knowledge.
It contains what happened, why it happened, what fixed it, which runbook was involved, and what should be remembered next time.
Recalling previous incidents
Once memories have been retained, the agent can retrieve relevant information when another incident occurs.
For example:
recall_response = client.recall(
bank_id="hindsightops-incidents",
query="Payment API database connection timeout QueuePool limit 50 reached",
tags=["service:payment-api"]
)
This allows the current incident to be connected with previous operational experience.
Instead of treating every incident as an isolated event, the agent can use historical evidence as part of its investigation.
This creates an important behavioral difference.
Without memory:
Current symptoms → AI reasoning → Generic recommendations
With memory:
Current symptoms → Historical evidence → AI reasoning → Context-specific recommendations
The goal is not to blindly follow whatever happened previously. Historical information becomes evidence that can be considered alongside the current incident.
The learning loop
The most important part of the system is not only recalling information. It is learning from the resolution of an incident.
HindsightOps allows an incident to move through an investigation and resolution workflow.
An engineer can investigate the incident, review observed telemetry, examine AI hypotheses, look at historical evidence, apply a mitigation, and then resolve the incident.
During resolution, information such as the confirmed root cause, actual resolution, resolution time, helpfulness feedback, and runbook can be captured.
That information can then be retained in Hindsight.
This means the system has a feedback loop:
- An incident happens.
- The agent investigates it.
- Historical incidents are recalled.
- The engineer takes action.
- The incident is resolved.
- The confirmed outcome is retained.
- Future incidents can recall the new learning.
This is the part that makes the project different from a simple chatbot with a prompt containing some documentation.
The memory can grow as incidents are resolved.
Building the Incident Investigation Studio
I built the interface around the actual incident-response workflow rather than making it a simple chat screen.
The Incident Investigation Studio separates several types of information:
• Observed telemetry and incident facts
• AI-generated root-cause hypotheses
• Historical evidence retrieved from memory
• Recommended actions
• Relevant runbooks
• Agent activity and memory source
This makes it possible to see not only what the AI recommends, but also where the historical context came from.
For example, if the agent recommends checking a database connection pool, the engineer can also see whether similar incidents were previously associated with that issue.
Similar Incidents Engine
Another part of HindsightOps is the Similar Incidents Engine.
It surfaces previously recalled incidents and provides information such as similarity and recovery time.
This is useful because incident response often depends on pattern recognition.
An engineer may recognize that the current error resembles something that happened weeks ago, but that information can easily be buried in old tickets, Slack conversations, or postmortems.
The memory layer provides a way for the application to bring that historical context back into the current workflow.
Hindsight Memory Explorer
I also added a Hindsight Memory Explorer so that the memory layer is not hidden from the user.
The interface can expose memory units and allow semantic searching of the stored operational knowledge.
This was important for development because it provides visibility into what the system actually remembers.
For an agent that is supposed to learn over time, being able to inspect its memory is useful for debugging retrieval quality and understanding why particular historical evidence was surfaced.
Runbooks and postmortems
Incident response is not only about identifying a root cause. Engineers also need actionable procedures.
HindsightOps includes a Runbook Catalog containing operational runbooks such as:
DB-04
REDIS-02
API-07
K8S-03
DEP-05
SEC-01
The application also includes a Postmortem Generator.
After an incident is resolved, the system can generate a structured postmortem and provide a way to save the resulting learning back into Hindsight.
This connects the traditional incident lifecycle with the agent's memory lifecycle.
Incident → Investigation → Resolution → Postmortem → Memory → Future Incident
That connection is one of the main design ideas behind the project.
Stateless vs memory-assisted behavior
I also included a comparison workflow that can demonstrate the difference between stateless analysis and Hindsight-assisted analysis.
The backend supports an analysis mode where Hindsight recall can be enabled or disabled.
With stateless analysis, the agent works without the historical memory retrieval step.
With memory-assisted analysis, the current incident can be enriched with relevant historical evidence.
This makes the effect of memory visible instead of simply claiming that the application has a memory system.
The goal of the demonstration is to show a behavioral difference: the same kind of incident can produce more context-aware investigation when the agent has access to previous operational experience.
Architecture
The frontend is built with React 19, Vite, Tailwind CSS, and Lucide icons.
The backend uses Python 3.11, FastAPI, SQLAlchemy, and Pydantic v2.
Groq is used for the LLM layer, while Hindsight provides persistent agent memory.
SQLite is currently used for application data, with the architecture prepared for PostgreSQL.
The application exposes FastAPI endpoints for incidents, analysis, resolution, postmortems, learning, memory statistics, semantic memory search, memory graphs, and runbooks.
For example, the incident analysis endpoint supports both memory-assisted and stateless analysis, which makes the before-and-after behavior easier to demonstrate.
I also added a health endpoint that reports the state of the backend, LLM, Hindsight, and database without exposing secrets.
Handling the demo environment
One practical problem when building a project around an external service is that a demonstration environment may not always have valid cloud credentials or network access.
For that reason, HindsightOps includes a local semantic memory fallback.
When Hindsight Cloud is configured, the application can use the live Hindsight integration.
When the Hindsight credentials are unavailable, the application can operate in a clearly labeled Demo Mode.
The interface distinguishes between the live Hindsight memory source and the demo memory source instead of pretending that the cloud service is being used.
This was important because I wanted the demo behavior to remain transparent.
What I learned
The first lesson was that memory needs structure.
Simply saving large amounts of incident text is not enough. Information such as service, severity, incident ID, runbook, root cause, resolution, and lessons learned makes the memory much more useful.
The second lesson was that retrieval needs to happen at the right point in the reasoning process.
The agent needs to understand the current incident before historical information can be useful. Dumping an entire memory store into every prompt would create unnecessary context and make reasoning harder.
The third lesson was that retention is just as important as recall.
A system that can remember but never learns from newly resolved incidents does not really improve over time.
The resolution workflow therefore became a critical part of the design.
The fourth lesson was that an agent's memory should be inspectable.
The Memory Explorer helped me understand what information was being stored and how it could be retrieved. It also makes the concept of persistent memory easier to demonstrate.
Finally, I learned that the most useful demonstration of agent memory is behavioral rather than visual.
Showing a screen full of stored memories is less meaningful than showing how those memories change the agent's response to a later incident.
Where I want to take it next
The current version focuses on incident investigation, historical recall, runbooks, resolution learning, and postmortem retention.
The next direction I would explore is deeper integration with real observability and incident-management systems.
For example, production telemetry, deployment events, monitoring alerts, and incident tickets could become additional sources of operational experience.
The long-term goal would be an incident-response agent that develops a useful memory of a specific engineering environment instead of acting like a generic troubleshooting assistant every time an outage occurs.
That is what made Hindsight particularly interesting for this project.
The value is not simply that an AI can remember a piece of text.
The value is that an agent can accumulate operational experience and use that experience when it encounters a similar situation again.
Project:
https://github.com/S-jais/Incident-Response-Agent
Hindsight:
https://github.com/vectorize-io/hindsight
Hindsight Documentation:
https://hindsight.vectorize.io/
Top comments (0)