DEV Community

Veladi Swathi
Veladi Swathi

Posted on

I Made Our Incident Agent Check Its Memory First

I Made Our Incident Agent Check Its Memory First

When an incident happens, an AI agent can produce a troubleshooting checklist in seconds.

The harder question is whether that checklist knows anything about what happened last time.

I designed our Incident Response Agent around that question.

From answering to investigating

A basic agent can receive:

Payment API is returning HTTP 500 errors.

and respond with:

Check logs.
Check database connectivity.
Check recent deployments.
Check service health.

That is useful, but generic.

Our agent adds another step:

What happened the last time we saw something like this?

That is where Hindsight enters the reasoning process.

The agent workflow

Our reasoning pipeline is:

Incident
↓
Understand symptoms
↓
Recall relevant memories
↓
Compare historical incidents
↓
Generate investigation plan
↓
Recommend runbook

The important design decision is that recall happens before the final recommendation.

[INSERT ACTUAL AGENT/REASONING CODE HERE]

// REAL CODE FROM YOUR REPOSITORY

The recalled information becomes context for the next stage.

A practical example

Suppose the current incident is:

Service:
Payment API

Symptoms:
HTTP 500 errors

Impact:
Users cannot complete payments

Hindsight recalls:

Previous incident:
Payment API HTTP 500 errors

Root cause:
Database connection pool exhaustion

Resolution:
Increased connection pool capacity

Runbook:
DB-CONNECTION-POOL

The agent can now produce a more targeted investigation:

  1. Check database connection utilization.

  2. Compare current pool usage with configured capacity.

  3. Inspect recent database-related errors.

  4. If pool exhaustion is confirmed, follow DB-CONNECTION-POOL.

The agent is still reasoning about the current incident.

The difference is that the reasoning has historical context.

Why I didn't make memory a hard rule

An early temptation with a memory-enabled agent is to say:

«"If you find a similar incident, use its solution."»

I don't think that is safe for incident response.

Two incidents can look similar while having completely different causes.

For example:

HTTP 500

could result from:

Database failure
Application exception
Dependency outage
Configuration issue
Deployment problem

So the memory should influence the investigation, not replace it.

Our agent therefore treats historical incidents as useful evidence.

Before and after

Before

Incident
↓
LLM
↓
Generic troubleshooting

After

Incident
↓
Hindsight recall
↓
Historical context
↓
LLM
↓
Targeted investigation

That small architectural change makes the agent's output much more connected to the organization's previous experience.

Turning recommendations into runbooks

One useful part of the design is connecting recalled incidents with runbooks.

A previous incident can tell us:

What happened

while the runbook tells us:

What procedure to follow

Combining both gives the agent a stronger basis for its recommendation.

For example:

Historical incident:
Database connection exhaustion

Runbook:
DB-CONNECTION-POOL

Recommendation:
Inspect pool utilization and follow the runbook
if the same failure mode is confirmed.

[INSERT YOUR ACTUAL RUNBOOK CODE/SCREENSHOT HERE]

What I learned

  1. Retrieval changes reasoning

Memory is most useful when retrieved information is available before the agent forms its recommendation.

  1. The agent should explain why it recommends something

A useful incident recommendation should expose the connection to previous incidents instead of simply producing an unexplained answer.

  1. Previous solutions should not become automatic actions

Memory can narrow the investigation without eliminating verification.

  1. Incident response needs uncertainty

A historical match is useful, but it does not prove that the current incident has the same root cause.

Closing

The interesting part of our agent is not that it can generate troubleshooting instructions.

It is that those instructions can be informed by what the system has encountered before.

Hindsight gives us the memory layer. The agent uses that memory as context, compares it with the current incident, and turns the combination into an investigation plan.

For incident response, that is a much more useful model of memory: not remembering everything, but remembering the experiences that can help with the next problem.
Add those photos on it to get good article.

Top comments (0)