How I Built an Incident Copilot That Remembers What Production Already Taught Us
Production incidents are rarely new.
The exact combination of service, configuration change, traffic pattern, database behavior, or deployment mistake might be different, but the underlying failure often looks familiar. The difficult part is that the engineer responding to the incident may not know that the organization has already solved something similar.
I built Incident Memory Copilot around a simple idea: before an AI assistant reasons about a new incident, it should first ask what the organization already knows about it.
The system uses Hindsight as the organizational memory layer and Groq as the reasoning layer. The important design decision is that memory is not an optional feature placed beside the chatbot. It sits directly in the incident-response workflow.
The loop is:
Incident
|
v
Hindsight Recall
|
v
Relevant Historical Memory
|
v
Groq Reasoning
|
v
Incident Guidance
|
v
Resolution
|
v
Hindsight Retain
|
v
Future Incident
In other words:
Recall -> Reason -> Resolve -> Retain
The problem I wanted to solve
When a production incident starts, engineers need answers quickly.
Was this database problem seen before? Has the authentication service had a similar outage? Did a previous configuration change cause the same symptom? Was there already a permanent fix?
Without organizational memory, an AI assistant can still produce a reasonable troubleshooting checklist. But that checklist is generic. It does not know what our organization has already experienced.
That distinction matters.
A team can have months or years of useful operational experience stored in incidents and postmortems, while a new engineer still spends the first few minutes rediscovering the same patterns.
I wanted the assistant to behave differently.
Instead of:
Incident -> LLM -> Answer
the workflow becomes:
Incident -> Memory -> LLM -> Answer
That small architectural change is the foundation of the project.
Why I made Hindsight the center of the system
If you're interested in the memory layer behind this architecture, see the Hindsight GitHub repository, the Hindsight documentation, and Vectorize's overview of agent memory.
Hindsight provides the persistent memory layer I needed. Its Node.js/TypeScript client exposes operations such as retain() for storing information and recall() for retrieving relevant memories. Hindsight Node.js / TypeScript quickstart
At the application level, the integration is intentionally simple.
A Hindsight client is created once and used by the server-side workflows:
import { HindsightClient } from "@vectorize-io/hindsight-client";
const hindsight = new HindsightClient({
baseUrl: process.env.HINDSIGHT_API_URL!,
apiKey: process.env.HINDSIGHT_API_KEY!,
});
The exact value I care about is not simply "does the memory database contain records?" It is whether a current incident can retrieve knowledge that is actually relevant to the situation.
The recall operation is the important boundary:
const memories = await hindsight.recall(
process.env.HINDSIGHT_BANK_ID!,
incident
);
That gives the reasoning layer historical context before it has to produce an answer.
What happens during incident analysis
The Analyze page is the main demonstration of the system.
An engineer enters a production symptom such as:
The Payment API is timing out again. Database connections are reaching
the pool limit while a reporting workload is running on the primary database.
The application first sends that incident description through the memory workflow.
Hindsight retrieves relevant historical knowledge. That memory is then passed into the Groq reasoning step.
The important point is that Groq is not being asked to magically remember the organization's previous incidents.
It is being given organizational context that was explicitly retrieved for this incident.
Conceptually:
Current incident
+
Recalled organizational memory
|
v
Groq
|
v
Memory-informed incident reasoning
This separation made the architecture much easier to reason about.
Hindsight answers:
What does the organization already know?
Groq answers:
Given that knowledge and the current incident, what should the engineer consider next?
The before-and-after difference
The clearest way to understand the value is to compare the behavior before memory is involved.
Imagine the Payment API incident above.
Without organizational memory, an assistant might respond with a generic checklist:
1. Check database connectivity.
2. Check application logs.
3. Check CPU and memory.
4. Check connection pool settings.
5. Restart affected services if necessary.
Those steps are not necessarily wrong. They are simply disconnected from the organization's history.
With memory, the system can retrieve a previous Payment API lesson involving a long-running analytics workload exhausting the shared database connection pool.
That changes the investigation.
Instead of treating the incident as an unknown problem, the assistant can point the engineer toward the known pattern:
Current symptom
+
Previous Payment API incident
=
Check long-running analytics work,
database connection usage,
and whether reporting traffic is
using the primary database.
The difference is not just that the answer contains more text.
The difference is that the reasoning is grounded in organizational experience.
Why I separated Memory from Memory Copilot
I also wanted memory to be visible rather than hidden inside one API response.
That is why the application has a dedicated Organizational Memory page.
It shows retained incident knowledge such as:
- Incident ID
- Service
- Severity
- Root cause
- Resolution
- Engineering lesson
The current examples cover Authentication, Redis Cache, and Payment API incidents.
This makes the underlying memory inspectable.
An engineer can see what the organization remembers instead of treating the AI response as a black box.
That also led to the second major workflow: Memory Copilot.
An engineer can ask:
What was the permanent fix for the Redis incident?
The question goes through the same pattern:
Question
|
v
Hindsight Recall
|
v
Relevant Memory
|
v
Groq
|
v
Answer
The result is a conversational interface backed by persistent organizational memory rather than a completely stateless chatbot.
Retaining the lesson is just as important as recalling it
Recall solves only half of the problem.
If an engineer resolves a production incident today but that resolution disappears afterward, the organization has learned nothing that can help tomorrow's incident.
That is why I built the Resolve and Retain workflow.
The engineer records:
Symptoms
Root Cause
Immediate Fix
Permanent Fix
and that resolution can then be written back into Hindsight.
The underlying operation is the reverse of recall:
await hindsight.retain(
process.env.HINDSIGHT_BANK_ID!,
resolution
);
Now the organization has a chance to retrieve that lesson later.
The complete loop becomes:
Today's incident
|
v
Recall yesterday's knowledge
|
v
Reason about today's incident
|
v
Resolve the incident
|
v
Retain today's lesson
|
v
Tomorrow's incident can recall it
That is the part of the project I find most interesting.
The system is not just designed to answer questions. It is designed to make incident experience reusable.
Building the operator experience
Once the core workflow worked, I changed the interface to make the memory loop obvious.
The application now uses an operator-style dashboard with separate areas for:
Overview
Analyze
Memory
Incidents
Memory Copilot
Analytics
Resolve
The homepage deliberately puts Hindsight near the center of the experience.
The message is simple:
REMEMBER BEFORE YOU RESPOND.
The dashboard also visualizes the incident archive, service concentration, memory activity, and the four-stage workflow:
RECALL
REASON
RESOLVE
RETAIN
I wanted the interface to communicate the architecture without forcing the engineer to read documentation before understanding the product.
Analytics are useful, but I kept their scope honest
The analytics page shows the current incident dataset through service concentration, severity distribution, and peak-latency values.
One design decision mattered here: I did not want illustrative values to be mistaken for real observability data.
The latency figures in the current dashboard are explicitly labeled as demo telemetry.
That means the system demonstrates the shape of an incident analytics layer without pretending it is already connected to production monitoring infrastructure.
A future version could connect the same workflow to real observability sources, but the current implementation keeps that boundary clear.
One limitation I ran into
The biggest limitation of a memory-powered incident assistant is not retrieving some memory. It is retrieving the right memory.
A memory system can return historical information that sounds related but is not actually useful for the current incident. That means retrieval quality becomes part of the product experience, not just an infrastructure detail.
I also had to resist the temptation to treat every dashboard number as production telemetry. For a prototype, realistic-looking data can make a product feel convincing, but misleading data is worse than obviously limited data.
That is why the current analytics page labels illustrative latency values honestly.
The next step is not to invent more dashboard numbers.
It is to connect the system to real incident and observability data.
What I learned
1. Memory has to be part of the workflow
Adding a memory page beside a chatbot would not have solved the problem.
The useful pattern is:
Retrieve memory
->
Use memory
->
Produce reasoning
Memory becomes valuable when it changes what the system does.
2. Retention matters as much as recall
A system that can only remember the past eventually becomes stale.
The more interesting loop is:
Recall -> Reason -> Resolve -> Retain
Every resolved incident has the potential to improve future responses.
3. AI reasoning and organizational knowledge are different jobs
I found it useful to keep the roles separate.
Hindsight provides the historical context.
Groq provides the reasoning.
That separation makes the architecture easier to explain, debug, and extend.
4. Good incident data matters
The quality of the system depends heavily on the quality of what gets retained.
An incident record with a clear root cause, resolution, and lesson is much more useful than a vague description of what happened.
5. The best incident assistant should reduce rediscovery
The goal is not to make engineers stop thinking.
The goal is to prevent them from repeatedly rediscovering lessons the organization already paid to learn.
Where I want to take it next
The current system is focused on the memory-driven incident workflow.
The natural next step is connecting it to real operational data:
Monitoring
+
Incident Management
+
Postmortems
+
Runbooks
|
v
Hindsight
|
v
Incident Memory Copilot
From there, the system could support richer incident correlation, real service telemetry, stronger incident similarity, and more automated runbook generation.
The important part is that the memory layer remains central.
Final thought
The most useful thing an incident assistant can remember is not a generic troubleshooting checklist.
It is what this organization learned the last time something went wrong.
That is why I built Incident Memory Copilot around Hindsight.
The goal is simple:



Top comments (0)