DEV Community

Bhavya sri
Bhavya sri

Posted on

AI Incident Response Agent powered by Hindsight

AI Incident Response Agent Powered by Hindsight

Overview

The most dangerous thing an incident tool can do at 3 a.m. is hand an engineer a past fix and let them apply it without asking whether this incident is actually the same problem.

I built RecallOps around the opposite habit: before the agent recommends anything, it has to explain why the current incident resembles a past one and where that resemblance breaks down.

RecallOps is an AI-powered incident response agent that uses Hindsight as its memory system. Instead of simply retrieving an old solution, the agent compares the current incident with previous incidents, identifies similarities and differences, recommends checks, and only then discusses possible actions.

What RecallOps Does

RecallOps is a small Python service with a Streamlit front end.

An on-call engineer:

  1. Describes a production incident in plain text.
  2. Selects the affected service.
  3. Selects the severity.
  4. Starts the analysis.

From there:

  1. The incident description is sent to Hindsight as a recall query against a bank of past incidents.
  2. The retrieved memories and the new incident are passed into a comparison-and-reasoning prompt.
  3. The model produces a structured analysis covering similarities, differences, recommended checks, possible actions, and warnings.
  4. Once the engineer resolves the incident, the confirmed resolution and outcome are written back to Hindsight using "retain".

The important part is that the agent does not treat a retrieved incident as automatically being the same problem.

Project Architecture

The main files are:

agent.py

recall → compare → reason → retain

app.py

Streamlit UI

hindsight_memory.py

Minimal Hindsight store/recall round trip

seed_memory.py

Loads the initial incident bank into Hindsight

Hindsight supplies the memory layer behind the recall and retain operations.

The main focus of this project, however, is what happens after recall returns something: how the agent decides whether a retrieved incident is actually relevant and how it avoids treating a resemblance as a confirmed match.

Tech Stack

  • Python
  • Streamlit
  • Hindsight
  • Large Language Model
  • Git/GitHub

Why Comparison, Not Lookup?

It would be easy to build RecallOps as a simple lookup tool:

«Take the incident → find the nearest match → display the old resolution.»

But that would essentially be a FAQ bot wearing an incident-response costume.

Two incidents can have similar symptoms but completely different causes.

For example, payment timeouts with connection errors could mean database connection pool exhaustion. But they could also be caused by:

  • A slow query holding connections open
  • A stalled downstream dependency
  • A bad deployment
  • Another infrastructure problem

If an agent simply says "raise the pool limit" because that worked previously, an engineer might apply the change without verifying the current cause.

So the main design decision in RecallOps is the reasoning contract:

«The model should first explain how similar the incidents are and what that similarity actually supports before recommending an action.»

The Recall Step

The recall step is intentionally simple:

def investigate_incident(incident):
result = hindsight.recall(
bank_id=BANK_ID,
query=incident
)

memories = "\n\n".join(
    memory.text for memory in result.results
)
Enter fullscreen mode Exit fullscreen mode

The incident is passed to Hindsight in the same language the engineer used.

For example:

«"Payment API is experiencing severe timeouts. Error rate has increased to 30%. Database connection errors are also appearing."»

There is no additional classification step before recall.

The idea is to let semantic recall find potentially relevant experiences and then allow the comparison stage to decide what those experiences actually mean.

Forcing the Model to Compare Before It Concludes

The comparison prompt is the center of the project.

The agent is instructed:

IMPORTANT RULES:

  1. Do NOT blindly copy a previous solution.
  2. First identify what is similar between the new incident and the previous experience.
  3. Clearly identify what is different.
  4. A previous solution should only be recommended directly if the current evidence supports it.
  5. If the previous incident only provides a similar pattern, describe it as a reference rather than a confirmed solution.
  6. Do not claim certainty when there is insufficient evidence.
  7. Prioritize checks that an engineer should perform before making a potentially risky production change.

The output follows a fixed structure:

1. Likely Root Cause

2. Hindsight Experience

3. Why It Is Relevant

4. What Is Different

5. Recommended Immediate Checks

6. Recommended Action

7. Previous Solution

8. Warnings

Two sections do most of the comparison work:

Why It Is Relevant

and

What Is Different

They appear before Recommended Action, forcing the model to present the reasoning behind the analogy before discussing what to do.

Recommended Immediate Checks also comes before Recommended Action, so the engineer sees what should be verified before a potentially risky change.

The Previous Solution section is intentionally placed near the end and treated as a reference rather than a guaranteed answer.

A fixed structure also makes weak reasoning easier to notice. If the "What Is Different" section is empty or generic, the comparison is clearly weak.

This does not guarantee perfect reasoning, but it makes the reasoning easier for a human to inspect.

What Makes a Memory Useful?

The comparison only works if the Hindsight memory bank contains meaningful incidents.

Each stored incident follows a structure like this:

Production Incident:
Notification service experienced a 35% failure rate
with email request timeouts and intermittent 504 errors
from the external email provider.

Actual Resolution:
The external email provider was experiencing an outage.
Enabled retry with exponential backoff and queued failed
notifications for later delivery.

Outcome:
Notification failures dropped below 2% after the provider
recovered and queued notifications were delivered successfully.

Lesson Learned:
When an external notification provider returns intermittent
504 errors, verify provider health before increasing local resources.

The Lesson Learned is written as a condition rather than a command.

For example:

«"When X and Y occur, check Z first."»

This gives the comparison stage something to evaluate.

An imperative lesson such as:

«"Increase the connection pool."»

would be easier for an agent to copy without checking whether the current incident actually matches the previous one.

Using Different Types of Incidents

The memory bank should not contain only incidents that point toward the same solution.

For example, one incident may involve:

  • Database connection pool exhaustion

Another may involve:

  • Stalled workers causing a queue backlog

Another may involve:

  • An external notification provider outage

This creates useful contrast.

A comparison is only as good as the evidence and alternatives available in the memory bank.

Closing the Loop With Confirmed Outcomes

The write path is intentionally narrower than the read path.

def save_experience(incident, resolution, outcome):

experience = f"""
Production Incident: {incident}

Actual Resolution: {resolution}

Outcome: {outcome}

Lesson Learned:
This production experience should be considered
when investigating similar future incidents.
"""

hindsight.retain(
    bank_id=BANK_ID,
    content=experience
)
Enter fullscreen mode Exit fullscreen mode

I do not write the model's own analysis back into Hindsight.

Only information confirmed by the engineer is retained:

  • The incident
  • The actual resolution
  • The observed outcome

This is important because an incorrect model analysis should not become a future "past experience."

Keeping the retain path limited to confirmed outcomes helps keep future comparisons grounded in what actually happened.

What the Comparison Looks Like in Practice

An engineer enters:

Payment API is experiencing severe timeouts.
Error rate has increased to 30%.
Database connection errors are also appearing.

Hindsight may retrieve an earlier payment incident where the error rate was 35% and the cause was connection pool exhaustion.

The comparison might look like:

Why It Is Relevant

The incidents involve the same service and have a similar combination of timeouts and connection errors.

What Is Different

The current incident does not confirm that the connection pool is saturated. Connection errors alone do not prove that pool exhaustion is the root cause.

Recommended Immediate Checks

Check:

  • Current connection pool utilization
  • Whether connections are being held by long-running queries
  • Current database health
  • Recent deployment or configuration changes

Previous Solution

Increasing the connection pool helped in the previous incident, but it should be treated as a reference rather than an automatic solution.

Now consider a different incident involving a payment-provider outage with no database connection errors.

The relevant memory and comparison can change because the evidence is different.

The agent does not hard-code a specific answer for every type of incident. It compares the retrieved experience with the current evidence.

Making the Reasoning Checkable

The Streamlit interface displays the retrieved Hindsight memories alongside the generated analysis.

This is intentional.

If the model says that two incidents are similar, the engineer can inspect the actual retrieved memory and judge whether that comparison makes sense.

The goal is not simply to produce an answer. The goal is to make the reasoning easier for the human engineer to inspect.

Demo

The RecallOps demo shows the following flow:

Engineer enters incident
↓
Selects service and severity
↓
Hindsight recalls previous incidents
↓
Agent compares current and previous incidents
↓
Agent identifies similarities and differences
↓
Agent recommends immediate checks
↓
Agent discusses possible action
↓
Engineer resolves the incident
↓
Confirmed outcome is retained in Hindsight

[Add screenshot of the RecallOps Streamlit interface here]

[Add screenshot of the comparison/analysis result here]

Project Links

GitHub Repository:
[Add your GitHub repository link here]

DEV.to Article:
This article documents the design decisions and reasoning approach behind RecallOps.

Lessons Learned

  1. Retrieval isn't the hardest part

Building the recall call was relatively straightforward.

The harder part was getting the agent to reason honestly about whether a retrieved incident actually applies.

  1. Force the "What's Different?" step

A free-text prompt can allow a model to skip comparison and jump directly to a recommendation.

A required "What Is Different" section makes that behavior easier to notice.

  1. Put checks before actions

Putting Recommended Immediate Checks before Recommended Action keeps verification ahead of potentially risky changes.

Placing Previous Solution near the end also prevents the old fix from becoming the headline.

  1. Write conditional lessons

A lesson such as:

«"When X and Y occur, check Z."»

gives the agent something to compare against.

A lesson such as:

«"Do Z."»

is easier to treat as an instruction.

  1. Keep the write path narrower than the read path

Only human-confirmed outcomes are written back into Hindsight.

The model's own conclusions are not automatically turned into future memories.

Future Improvements

There are several areas I would explore next:

  • Scope comparisons to the same service by default.
  • Give more weight to very recent incidents.
  • Improve filtering of retrieved memories.
  • Evaluate comparison quality using a larger incident dataset.
  • Add more structured incident metadata.
  • Add human feedback on whether the comparison was useful.

Conclusion

RecallOps is built around a simple idea:

«An incident agent should not assume that a similar incident has the same solution.»

The useful part of incident memory is not simply retrieving what happened before. It is understanding why the previous experience is relevant, where it differs, and what evidence should be checked before taking action.

The project therefore puts comparison before recommendation and keeps confirmed outcomes separate from model-generated reasoning.

The goal is not to make the agent blindly confident. It is to make the agent show its reasoning when it says two incidents are alike — and to be willing to say "this doesn't match" when the evidence does not support the comparison.

Top comments (0)