DEV Community

Boini chandana
Boini chandana

Posted on

AI INCIDENT RESPONSE AGENT POWERED BY HINDSIGHT

AI Incident Response Agent Powered by Hindsight

The most dangerous thing an incident tool can do at 3 a.m. is confidently repeat last month's fix.

I built RecallOps around that problem: an incident response agent that remembers what happened before, but is explicitly designed not to treat that memory as the answer.

What RecallOps Does

RecallOps is a small Python service with a Streamlit front end. An on-call engineer describes a production incident in plain text, selects the affected service and severity, and starts an investigation.

Behind that button, three things happen:

  1. The incident description is used as a query against a long-term memory store.
  2. The recalled past incidents, together with the new incident, are given to a reasoning model.
  3. The model returns a structured analysis with a fixed set of sections.

After the incident is resolved, the engineer records what actually fixed it and what the outcome was. That experience is written back to the same memory store, so future investigations can use it as additional context.

The core implementation is split across the agent logic, Streamlit interface, and memory seeding code:

agent.py        # recall → reason → retain
app.py          # Streamlit UI
seed_memory.py  # loads the initial incident bank
Enter fullscreen mode Exit fullscreen mode

The memory layer is Hindsight, an open-source agent memory system. The reasoning model is openai/gpt-oss-120b served through Groq, with a temperature of 0.2. I kept the temperature low because this application needs consistent reasoning rather than creative output.

I chose not to build my own memory layer. Building a basic vector-search memory system is possible, but I wanted to spend my time on the agent's behavior rather than maintaining the retrieval infrastructure.

The more interesting problems come later: deciding what to keep, how to retrieve it from a messy incident description, and how to keep stored experiences useful as they grow.

For this project, I used Hindsight's documentation as the integration contract and kept the memory interaction focused on two operations:

retain and recall.


The Through-Line: Memory Is Evidence, Not an Answer

Here's the failure mode I designed against.

You give an LLM a new incident and a retrieved past incident that looks similar. The model sees a familiar symptom and a previous fix and says:

"Increase the connection pool from 100 to 200."

It sounds authoritative.

But similar symptoms do not necessarily mean the same root cause.

Payment timeouts with database connection errors might indicate connection pool exhaustion. They could also come from a slow query holding connections open, a downstream dependency becoming slow, or a deployment that caused connection leaks.

Increasing the connection pool in the wrong situation could make the underlying problem worse by putting additional load on an already struggling database.

So the retrieval side of RecallOps is deliberately straightforward. The more important design work is in how the retrieved memory is presented to the reasoning model.


Retrieval: One Call, No Extra Classification Step

The recall function is intentionally simple:

def investigate_incident(incident):

    result = hindsight.recall(
        bank_id=BANK_ID,
        query=incident
    )

    memories = "\n\n".join(
        memory.text for memory in result.results
    )
Enter fullscreen mode Exit fullscreen mode

The query is the raw incident description, exactly as the engineer entered it.

I don't rewrite it, extract keywords, or classify it before sending it to Hindsight. In my testing, using the incident description directly as the recall query was sufficient to retrieve relevant previous experiences without adding a separate classification step.

The retrieved memories come back as text, and that text is passed into the reasoning prompt along with the new incident.

The same memory text is also shown in the UI.

That matters because the engineer can see what the agent is actually basing its reasoning on rather than receiving a recommendation with no visible evidence.


Framing: Rules That Force a Comparison

The prompt is where the main behavior is defined.

The model isn't simply asked:

"What should we do?"

Instead, it is instructed to compare the current incident with previous experience and identify where the comparison breaks down.

The important rules are:

IMPORTANT RULES:

1. Do NOT blindly copy a previous solution.

2. First identify what is similar between the new incident
   and the previous experience.

3. Clearly identify what is different.

4. A previous solution should only be recommended directly
   if the current evidence supports it.

5. If the previous incident only provides a similar pattern,
   describe it as a reference rather than a confirmed solution.

6. Do not claim certainty when there is insufficient evidence.

7. Prioritize checks that an engineer should perform before
   making a potentially risky production change.
Enter fullscreen mode Exit fullscreen mode

The output is then locked into a fixed structure:

  1. Likely Root Cause
  2. Hindsight Experience
  3. Why It Is Relevant
  4. What Is Different
  5. Recommended Immediate Checks
  6. Recommended Action
  7. Previous Solution
  8. Warnings

The order of these sections is intentional.

"Why It Is Relevant" and "What Is Different" come before the recommendation, so the model has to examine the analogy before suggesting an action.

"Recommended Immediate Checks" comes before "Recommended Action", so the first thing an engineer sees is what to verify rather than what to change.

And "Previous Solution" appears near the bottom, with an explicit instruction not to present it as guaranteed.

The previous fix is available, but it is treated as a hypothesis rather than an automatic answer.

I can't prove that a prompt alone makes a model behave correctly, and I don't want to claim a benchmark I didn't run. What I can say is that a fixed schema with explicit "differences" and "checks" sections is easier to review than unrestricted prose. It also makes it obvious when the model skips an important part of the reasoning.


The Bank Is a Set of Lessons, Not Just a Log

The memory store starts with a handful of incidents. Each experience follows the same basic structure:

  • The incident
  • The actual resolution
  • The outcome
  • A lesson

For example:

Production Incident:
Notification service experienced a 35% failure rate with email
request timeouts and intermittent 504 errors from the external
email provider.

Actual Resolution:
The external email provider was experiencing an outage.
Enabled retry with exponential backoff and queued failed
notifications for later delivery.

Outcome:
Notification failures dropped below 2% after the provider
recovered and queued notifications were delivered successfully.

Lesson Learned:
When external notification providers return intermittent
504 errors, verify provider health before increasing local resources.
Enter fullscreen mode Exit fullscreen mode

I put particular attention into the lesson because it affects how the experience can be reused later.

The payment incident in the initial bank involves database connection problems and an increased connection pool.

Other experiences point in different directions.

For example, the notification incident suggests checking an external provider before increasing local resources. Another incident involving a growing order queue points toward checking worker health and consumer capacity.

That variety gives the agent different experiences to compare instead of repeatedly seeing the same type of solution.

The stored lessons are also phrased as conditional rules:

When X and Y occur together, check Z first.

rather than direct commands:

Do Z.

That makes the memory more useful when a future incident is similar but not identical.


Closing the Loop: Retain What Actually Happened

The part I care about most is the write path.

An AI analysis is only useful once someone has found out whether it was actually correct. In a production environment, the engineer is the person who knows what ultimately happened.

RecallOps therefore lets the engineer record the actual resolution and outcome after the incident has been handled.

The write operation is:

def save_experience(incident, resolution, outcome):

    experience = f"""
Production Incident:
{incident}

Actual Resolution:
{resolution}

Outcome:
{outcome}

Lesson Learned:
This production experience should be considered
when investigating similar future incidents.
"""

    hindsight.retain(
        bank_id=BANK_ID,
        content=experience
    )
Enter fullscreen mode Exit fullscreen mode

There are two design choices here that are important.

1. The model's investigation is not treated as ground truth

I don't store the model's investigation as if it were a confirmed fact.

The stored experience contains:

  • the incident as described,
  • the resolution entered by the engineer,
  • the outcome,
  • and a lesson-oriented note.

The model's suggestion might have been wrong, partially correct, or ignored entirely. None of that should automatically become production knowledge.

2. New experiences use the same structure

The same four-part structure is used for the initial seed incidents and experiences recorded later.

That consistency means the reasoning model sees similar information whether an experience was added during setup or recorded after a real investigation.

Because the write goes to the same bank that future investigations read from, the learning loop becomes:

Incident → Recall → Reason → Resolve → Retain

A resolved incident becomes a potential recall result for a future incident.

That is the main difference between maintaining a static collection of examples and using a memory system that can accumulate newly confirmed experiences.


What It Looks Like in Use

Consider a payment incident:

"Payment API is experiencing severe timeouts. Error rate has increased to 30%. Database connection errors are also appearing."

RecallOps can retrieve an earlier payment incident where the error rate was 35% and the cause was an exhausted database connection pool.

The analysis follows the fixed structure.

In simplified form:

Likely Root Cause:
Connection pool exhaustion is a leading hypothesis, but it is not presented as certain.

Hindsight Experience:
A previous payment incident involved timeouts, database connection errors, and an exhausted connection pool.

Why It Is Relevant:
The service and symptom pattern are similar.

What Is Different:
The current incident does not confirm pool saturation. It only reports database connection errors.

Recommended Immediate Checks:
Check current pool utilization, connection counts, and whether connections are being held by long-running queries before changing the pool limit.

Previous Solution:
Increasing the pool worked in the previous incident and could potentially be adapted, but it is presented as a reference rather than a guaranteed fix.

Now change the incident.

Suppose the report says:

"The payment provider is returning intermittent 504 errors and there are no database errors."

The previous payment incident is no longer an exact match.

A notification-provider experience becomes more relevant because it contains a similar pattern: an external dependency returning errors.

The agent can therefore recommend checking the external provider's health before assuming that the local database or connection pool is the problem.

The important part is that this behavior isn't implemented as a hard-coded rule saying:

"If you see 504, do X."

Instead, the relevant experience is retrieved from memory and compared with the current incident.


Making the Evidence Visible

The UI shows the raw retrieved memories in an expandable panel alongside the AI analysis.

That was a deliberate design decision.

If the agent cites a previous incident, the engineer should be able to read that incident and decide whether the comparison makes sense.

This creates a simple chain:

Past experience → Retrieved memory → AI reasoning → Engineer judgment

The engineer remains in the loop instead of treating the AI recommendation as an automatic production action.

For an incident-response system, being able to inspect the evidence is almost as important as generating the recommendation itself.


Lessons Learned

1. Frame retrieved memory as a hypothesis

In this project, the retrieval itself was relatively straightforward.

The harder part was deciding how the agent should use what it retrieved.

The behavior became more useful when the prompt required the model to identify:

  • what is similar,
  • what is different,
  • and what should be verified first.

2. Store lessons as conditions, not commands

"When A and B occur together, check C first" is more useful than simply storing "Do C."

A future incident can have the same general pattern without having the exact same root cause.

3. Only write down what a human confirmed

Keeping the model's investigation out of the memory store prevents the system from treating its own guesses as confirmed production knowledge.

The engineer provides the actual resolution and outcome.

4. Let the memory system handle the underlying retrieval infrastructure

Using a purpose-built memory system meant I could spend more time on the prompt, reasoning structure, and UI instead of implementing the underlying retrieval pipeline myself.

The integration remained focused on the two operations I needed: retain and recall.

5. Make the evidence visible

Showing the retrieved memories beside the analysis costs little and makes the recommendation easier to inspect.

An engineer can see not only what the agent recommends, but also which previous experience influenced that recommendation.


What I Would Improve Next

There are still things I would change with more time.

For example, recall could be filtered more explicitly by service, and recent incidents could potentially be weighted differently from older ones.

I would also like to test the system with a much larger and more varied incident dataset instead of the small set used during development.

But the core design has worked well in my testing.

An incident agent doesn't need to be clever.

It needs to remember accurately, say what it doesn't know, and get a little better every time someone tells it how things actually ended.

Top comments (0)