An AI assistant can analyze logs. It can summarize an incident. It can even suggest a possible root cause.
But there is a problem with starting every investigation from zero.
If the same class of failure happens again next month, the assistant may have no idea that the team already solved something similar.
That was the part of incident response I found most interesting while working on IncidentMind.
The goal wasn't just to build an AI assistant that could investigate production incidents. The more interesting goal was to give the investigation process a memory layer so that a resolved incident could become useful context for a future one.
Incident → Investigation → Resolution → Memory → Future Investigation
And Hindsight is what makes the last two steps possible.
The Problem With Starting From Zero
Production incidents rarely arrive as clean, isolated problems.
An authentication failure might involve:
- a recent deployment
- configuration changes
- multiple service instances
- logs and traces
- external dependencies
- user-facing symptoms
An engineer has to connect all of those pieces before deciding what is actually happening.
AI can help organize that information, but there is still a missing piece:
What happened the last time we saw something similar?
A traditional application database can tell us that an incident existed.
It can tell us its ID, status, service, severity, and other application data.
But operational experience is different.
The useful information might be:
This type of authentication failure was previously caused by inconsistent signing-key versions across service instances.
That is not just application state.
It is experience.
That distinction became the central design idea behind IncidentMind.
What IncidentMind Actually Remembers
IncidentMind has two different persistence responsibilities.
SQLite handles application state.
Hindsight handles retained incident experience.
That means the application can keep normal records such as incident IDs and statuses in SQLite while using Hindsight to retain information that could help with future investigations.
A resolved incident can contribute information such as:
- Confirmed root cause
- Resolution steps
- Runbook
- Lessons learned
- Prevention steps
- Service and severity context
The important part is that memory is created from a resolved incident rather than treating every piece of incoming information as something worth remembering.
IncidentMind separates current incident state from retained incident experience.
A Production-Style Failure
To test the workflow, we used an authentication incident.
The service was an Authentication API.
A deployment introduced JWT validation middleware and changed the authentication service's signing-key configuration to use a centralized secrets provider.
After the deployment, previously authenticated users were intermittently logged out, while new authentication attempts returned HTTP 401 errors.
The logs contained a particularly useful clue:
ERROR auth-api
JWT validation failed:
InvalidSignatureError: signature verification failed
WARN auth-api
POST /auth/refresh returned status=401
WARN auth-api
Instance=auth-api-7c8d9 signing_key_version=v2
WARN auth-api
Instance=auth-api-5f2a1 signing_key_version=v1
The interesting detail is that the database itself was healthy.
The problem was distributed across the authentication instances.
One instance was using one signing-key version while another was using a different version.
The incident logs contain the clues needed to connect the authentication failures with inconsistent signing-key configuration.
Giving the Agent the Full Incident Context
The next step is submitting the incident to IncidentMind.
Instead of sending the model only an error message, the workflow provides broader context:
- affected service
- incident ID
- severity
- recent deployment or configuration changes
- symptoms
- error logs
That gives the investigation agent more than a single error string to reason about.
IncidentMind collects the surrounding context before triggering the investigation.
The agent then produces a structured investigation.
In this example, the investigation connected the authentication failures with the inconsistent signing-key versions.
The important point is that the model isn't replacing the engineer.
The engineer still needs to verify the diagnosis.
The agent is helping organize the available evidence and surface a likely explanation.
The investigation view combines the incident context with the AI-generated investigation.
The Part That Changes Everything: The Resolution Becomes Memory
Finding a root cause is not the end of the workflow.
Once the cause was confirmed, the affected authentication instances were brought onto the same signing-key version, the affected instances were restarted, and authentication flows were verified again.
We also recorded the operational knowledge around the fix.
The runbook was:
RB-AUTH-KEY-08 — JWT Signing-Key Synchronization Failure
The lesson was that authentication deployments should verify signing-key consistency across instances and include automated configuration checks and controlled key rotation.
This is the information that should survive beyond the individual incident.
The resolved incident is converted into reusable operational experience.
How Hindsight Fits Into the Workflow
The Hindsight integration is deliberately small.
When the incident has been resolved, IncidentMind retains the completed incident knowledge:
response = client.retain(
bank_id=bank,
content=memory_payload,
metadata=metadata,
tags=[service, severity, "incident_resolution"]
)
The important detail here isn't the number of lines of code.
It is what those lines represent.
The application is taking an outcome that was previously trapped inside one incident and making it available to future investigations.
Later, when another incident occurs, IncidentMind can query Hindsight for relevant historical experience:
recall_resp = client.recall(
bank_id=bank,
query=query,
tags=[service] if service else None,
budget="mid"
)
That creates a very different investigation flow.
Without memory:
Current incident → AI investigation
With memory:
Current incident + relevant past experience → AI investigation
The retain and recall operations form the memory layer of the investigation workflow.
Why Application State and Memory Are Different
One design decision I found particularly useful was keeping application state and operational memory separate.
The application database can answer:
Which incidents exist?
Hindsight can help answer:
Have we seen something like this before, and what did we learn?
Those questions sound similar, but they serve different purposes.
SQLite gives the application predictable persistence.
Hindsight gives the agent a mechanism for retaining and recalling experience.
This separation also makes the architecture easier to reason about.
IncidentMind separates application state in SQLite from operational memory in Hindsight.
The important architectural boundary is between application state and operational memory.
Making Memory Visible
There was another lesson that wasn't obvious at first.
A memory operation can succeed in the backend without being obvious to the person using the application.
That makes debugging difficult.
If an engineer clicks a button to retain a resolved incident, they should be able to see that the operation actually contributed to the application's memory state.
We therefore made the retained-memory state visible on the dashboard.
The dashboard after successful retention shows the updated retained-memory count.
This sounds like a small UI detail, but it matters for systems involving memory.
When persistence is part of the product's behavior, it should be observable.
Otherwise, it becomes difficult to tell whether the system actually remembered something or merely appeared to.
Looking at the Code Behind the Memory Layer
The Hindsight client is initialized separately from the rest of the application logic.
The project also has a local fallback path when Hindsight credentials aren't configured, while the configured environment can use Hindsight Cloud.
That made development easier because the rest of the application didn't have to completely stop working when the external memory service wasn't available.
The memory module handles the Hindsight integration and the application's fallback behavior.
The important architectural boundary is that the rest of IncidentMind doesn't need to know every detail about how memory is stored.
It can work with two simple operations:
retain this experience
and
recall relevant experience
That keeps the memory mechanism relatively isolated from the incident workflow.
Testing the Workflow
Another thing I didn't want was a system that only looked convincing during a manual demo.
The project includes automated tests around important application behavior, including incident handling and the retained-memory dashboard counter.
The final test suite passed with:
7 passed
Automated tests cover important application behavior, including the retained-memory flow.
This became particularly useful while changing the retention workflow.
A memory-related change can affect both backend behavior and what the dashboard reports, so having tests around that behavior provides a safety net.
What I Learned About Useful Agent Memory
The biggest lesson for me was that adding memory isn't simply adding a retain() call.
There are at least three separate questions.
What should the system remember?
For incident response, storing every raw log isn't necessarily the most useful form of memory.
A confirmed root cause, the actual resolution, the runbook, and the lesson learned are much more useful operational knowledge.
When should something become memory?
A submitted incident isn't necessarily a lesson.
The incident becomes much more valuable after the root cause has been confirmed and the resolution is known.
That is why the retention step belongs after resolution.
How should memory affect a new investigation?
This is probably the most important question.
A previous incident should provide context, not automatically become the answer.
If an old incident looks similar to a new one, the agent still needs to consider the evidence from the current incident.
Historical memory should help the investigation, not replace it.
That distinction matters if an AI system is going to be useful in an engineering environment.
What I Would Improve Next
There are several things I would explore from here.
First, I would make recalled incidents more visible during the investigation itself.
If a previous incident influenced the agent's reasoning, an engineer should be able to see which historical experience was relevant.
Second, I would build richer relationships between incidents.
For example, incidents could be connected through:
- service
- deployment
- dependency
- configuration
- failure pattern
Third, I would measure memory quality instead of only measuring whether something was successfully stored.
A memory system can retain information perfectly and still retrieve irrelevant information.
So an important future metric is:
Did the recalled experience actually help with this incident?
Finally, I would explore how retained runbooks and prevention steps could move the system beyond diagnosis and toward improving incident-response practices over time.
The Bigger Idea
The interesting part of IncidentMind isn't that an LLM can read an incident.
The more interesting question is what happens when the system can carry useful experience from one incident into another.
A normal incident workflow often looks like:
Detect → Investigate → Fix → Close
The workflow we explored adds another step:
Detect → Investigate → Fix → Learn → Recall
That extra step changes the role of the AI assistant.
Instead of being useful only during the current incident, it can become part of a longer learning loop.
The goal isn't to make the agent automatically correct.
The goal is to make sure that when the team learns something valuable from an incident, that knowledge doesn't disappear when the incident is closed.
That's the idea behind using Hindsight with IncidentMind:
Resolve the incident. Retain the lesson. Use it when the next one arrives.
Resources
Hindsight GitHub: https://github.com/vectorize-io/hindsight
Hindsight Documentation: https://hindsight.vectorize.io/
What Is Agent Memory?: https://vectorize.io/what-is-agent-memory
IncidentMind: https://github.com/MRajeshwariReddy/incidentmind








Top comments (0)