DEV Community

Dikshitha Kasoju
Dikshitha Kasoju

Posted on

What Changes When an Incident Agent Can Remember?

An AI assistant can analyze logs. It can summarize an incident. It can even suggest a possible root cause.

But there is a problem with starting every investigation from zero.

If the same class of failure happens again next month, the assistant may have no idea that the team already solved something similar.

That was the part of incident response I found most interesting while working on IncidentMind.

The goal wasn't just to build an AI assistant that could investigate production incidents. The more interesting goal was to give the investigation process a memory layer so that a resolved incident could become useful context for a future one.

Incident → Investigation → Resolution → Memory → Future Investigation

And Hindsight is what makes the last two steps possible.

The Problem With Starting From Zero

Production incidents rarely arrive as clean, isolated problems.

An authentication failure might involve:

  • a recent deployment
  • configuration changes
  • multiple service instances
  • logs and traces
  • external dependencies
  • user-facing symptoms

An engineer has to connect all of those pieces before deciding what is actually happening.

AI can help organize that information, but there is still a missing piece:

What happened the last time we saw something similar?

A traditional application database can tell us that an incident existed.

It can tell us its ID, status, service, severity, and other application data.

But operational experience is different.

The useful information might be:

This type of authentication failure was previously caused by inconsistent signing-key versions across service instances.

That is not just application state.

It is experience.

That distinction became the central design idea behind IncidentMind.

What IncidentMind Actually Remembers

IncidentMind has two different persistence responsibilities.

SQLite handles application state.

Hindsight handles retained incident experience.

That means the application can keep normal records such as incident IDs and statuses in SQLite while using Hindsight to retain information that could help with future investigations.

A resolved incident can contribute information such as:

  • Confirmed root cause
  • Resolution steps
  • Runbook
  • Lessons learned
  • Prevention steps
  • Service and severity context

The important part is that memory is created from a resolved incident rather than treating every piece of incoming information as something worth remembering.

IncidentMind separates current incident state from retained incident experience.

A Production-Style Failure

To test the workflow, we used an authentication incident.

The service was an Authentication API.

A deployment introduced JWT validation middleware and changed the authentication service's signing-key configuration to use a centralized secrets provider.

After the deployment, previously authenticated users were intermittently logged out, while new authentication attempts returned HTTP 401 errors.

The logs contained a particularly useful clue:

ERROR auth-api
JWT validation failed:
InvalidSignatureError: signature verification failed

WARN auth-api
POST /auth/refresh returned status=401

WARN auth-api
Instance=auth-api-7c8d9 signing_key_version=v2

WARN auth-api
Instance=auth-api-5f2a1 signing_key_version=v1
Enter fullscreen mode Exit fullscreen mode

The interesting detail is that the database itself was healthy.

The problem was distributed across the authentication instances.

One instance was using one signing-key version while another was using a different version.

The incident logs contain the clues needed to connect the authentication failures with inconsistent signing-key configuration.

Giving the Agent the Full Incident Context

The next step is submitting the incident to IncidentMind.

Instead of sending the model only an error message, the workflow provides broader context:

  • affected service
  • incident ID
  • severity
  • recent deployment or configuration changes
  • symptoms
  • error logs

That gives the investigation agent more than a single error string to reason about.

IncidentMind collects the surrounding context before triggering the investigation.

The agent then produces a structured investigation.

In this example, the investigation connected the authentication failures with the inconsistent signing-key versions.

The important point is that the model isn't replacing the engineer.

The engineer still needs to verify the diagnosis.

The agent is helping organize the available evidence and surface a likely explanation.

The investigation view combines the incident context with the AI-generated investigation.

The Part That Changes Everything: The Resolution Becomes Memory

Finding a root cause is not the end of the workflow.

Once the cause was confirmed, the affected authentication instances were brought onto the same signing-key version, the affected instances were restarted, and authentication flows were verified again.

We also recorded the operational knowledge around the fix.

The runbook was:

RB-AUTH-KEY-08 — JWT Signing-Key Synchronization Failure

The lesson was that authentication deployments should verify signing-key consistency across instances and include automated configuration checks and controlled key rotation.

This is the information that should survive beyond the individual incident.

The resolved incident is converted into reusable operational experience.

How Hindsight Fits Into the Workflow

The Hindsight integration is deliberately small.

When the incident has been resolved, IncidentMind retains the completed incident knowledge:

response = client.retain(
    bank_id=bank,
    content=memory_payload,
    metadata=metadata,
    tags=[service, severity, "incident_resolution"]
)
Enter fullscreen mode Exit fullscreen mode

The important detail here isn't the number of lines of code.

It is what those lines represent.

The application is taking an outcome that was previously trapped inside one incident and making it available to future investigations.

Later, when another incident occurs, IncidentMind can query Hindsight for relevant historical experience:

recall_resp = client.recall(
    bank_id=bank,
    query=query,
    tags=[service] if service else None,
    budget="mid"
)
Enter fullscreen mode Exit fullscreen mode

That creates a very different investigation flow.

Without memory:

Current incident → AI investigation

With memory:

Current incident + relevant past experience → AI investigation

The retain and recall operations form the memory layer of the investigation workflow.

Why Application State and Memory Are Different

One design decision I found particularly useful was keeping application state and operational memory separate.

The application database can answer:

Which incidents exist?

Hindsight can help answer:

Have we seen something like this before, and what did we learn?

Those questions sound similar, but they serve different purposes.

SQLite gives the application predictable persistence.

Hindsight gives the agent a mechanism for retaining and recalling experience.

This separation also makes the architecture easier to reason about.

IncidentMind separates application state in SQLite from operational memory in Hindsight.

The important architectural boundary is between application state and operational memory.

Making Memory Visible

There was another lesson that wasn't obvious at first.

A memory operation can succeed in the backend without being obvious to the person using the application.

That makes debugging difficult.

If an engineer clicks a button to retain a resolved incident, they should be able to see that the operation actually contributed to the application's memory state.

We therefore made the retained-memory state visible on the dashboard.

The dashboard after successful retention shows the updated retained-memory count.

This sounds like a small UI detail, but it matters for systems involving memory.

When persistence is part of the product's behavior, it should be observable.

Otherwise, it becomes difficult to tell whether the system actually remembered something or merely appeared to.

Looking at the Code Behind the Memory Layer

The Hindsight client is initialized separately from the rest of the application logic.

The project also has a local fallback path when Hindsight credentials aren't configured, while the configured environment can use Hindsight Cloud.

That made development easier because the rest of the application didn't have to completely stop working when the external memory service wasn't available.

The memory module handles the Hindsight integration and the application's fallback behavior.

The important architectural boundary is that the rest of IncidentMind doesn't need to know every detail about how memory is stored.

It can work with two simple operations:

retain this experience

and

recall relevant experience

That keeps the memory mechanism relatively isolated from the incident workflow.

Testing the Workflow

Another thing I didn't want was a system that only looked convincing during a manual demo.

The project includes automated tests around important application behavior, including incident handling and the retained-memory dashboard counter.

The final test suite passed with:

7 passed
Enter fullscreen mode Exit fullscreen mode

Automated tests cover important application behavior, including the retained-memory flow.

This became particularly useful while changing the retention workflow.

A memory-related change can affect both backend behavior and what the dashboard reports, so having tests around that behavior provides a safety net.

What I Learned About Useful Agent Memory

The biggest lesson for me was that adding memory isn't simply adding a retain() call.

There are at least three separate questions.

What should the system remember?

For incident response, storing every raw log isn't necessarily the most useful form of memory.

A confirmed root cause, the actual resolution, the runbook, and the lesson learned are much more useful operational knowledge.

When should something become memory?

A submitted incident isn't necessarily a lesson.

The incident becomes much more valuable after the root cause has been confirmed and the resolution is known.

That is why the retention step belongs after resolution.

How should memory affect a new investigation?

This is probably the most important question.

A previous incident should provide context, not automatically become the answer.

If an old incident looks similar to a new one, the agent still needs to consider the evidence from the current incident.

Historical memory should help the investigation, not replace it.

That distinction matters if an AI system is going to be useful in an engineering environment.

What I Would Improve Next

There are several things I would explore from here.

First, I would make recalled incidents more visible during the investigation itself.

If a previous incident influenced the agent's reasoning, an engineer should be able to see which historical experience was relevant.

Second, I would build richer relationships between incidents.

For example, incidents could be connected through:

  • service
  • deployment
  • dependency
  • configuration
  • failure pattern

Third, I would measure memory quality instead of only measuring whether something was successfully stored.

A memory system can retain information perfectly and still retrieve irrelevant information.

So an important future metric is:

Did the recalled experience actually help with this incident?

Finally, I would explore how retained runbooks and prevention steps could move the system beyond diagnosis and toward improving incident-response practices over time.

The Bigger Idea

The interesting part of IncidentMind isn't that an LLM can read an incident.

The more interesting question is what happens when the system can carry useful experience from one incident into another.

A normal incident workflow often looks like:

Detect → Investigate → Fix → Close

The workflow we explored adds another step:

Detect → Investigate → Fix → Learn → Recall

That extra step changes the role of the AI assistant.

Instead of being useful only during the current incident, it can become part of a longer learning loop.

The goal isn't to make the agent automatically correct.

The goal is to make sure that when the team learns something valuable from an incident, that knowledge doesn't disappear when the incident is closed.

That's the idea behind using Hindsight with IncidentMind:

Resolve the incident. Retain the lesson. Use it when the next one arrives.

Resources

Hindsight GitHub: https://github.com/vectorize-io/hindsight

Hindsight Documentation: https://hindsight.vectorize.io/

What Is Agent Memory?: https://vectorize.io/what-is-agent-memory

IncidentMind: https://github.com/MRajeshwariReddy/incidentmind

Top comments (0)