DEV Community

Anjali devi Muppidi
Anjali devi Muppidi

Posted on

I Built an AI Incident Agent with Hindsight That Remembers What Failed

An incident-response agent can generate a convincing answer to an error log. The harder problem is making sure it can learn from what happened the last time a similar incident occurred.

I built an incident-response agent around that idea: instead of treating every production error as a completely new problem, the agent can retrieve relevant incidents from its memory, use them as context, and remember whether its suggested fix worked or failed.

The system combines Hindsight for agent memory, Groq for the language model, and Streamlit for the interface.

The result is a simple workflow:

New incident → recall similar incidents → generate a suggestion → record the outcome → use that experience later

The problem with an agent that forgets

Imagine an engineer receives this:

Database connection timeout after 500 requests
Enter fullscreen mode Exit fullscreen mode

A language model can suggest several possible causes. Maybe the connection pool is too small. Maybe the database is overloaded. Maybe there is a network problem.

But suppose the engineering team already dealt with an almost identical incident last month.

They discovered that the connection pool was exhausted and increased the pool size from 10 to 50. The fix worked.

Without access to that history, the agent has no reason to know what happened previously.

That is the problem I wanted to address.

Instead of asking the model to solve every incident from scratch, I wanted to give it access to previous incident experience.

Making memory part of the workflow

The system stores four important pieces of information for each incident:

  • The original error log
  • The identified root cause
  • The fix that was applied
  • Whether the fix worked or failed

Hindsight provides the memory layer for this information.

The basic idea is to retain an incident when it happens and recall relevant incidents when a new one arrives.

A simplified version of the save operation looks like this:

def save_incident(log, root_cause, fix, outcome):
    text = (
        f"Error: {log}\n"
        f"Root cause: {root_cause}\n"
        f"Fix: {fix}\n"
        f"Outcome: {outcome}"
    )

    response = requests.post(
        f"{BASE_URL}/spaces/{SPACE_ID}/retain",
        headers={"Authorization": f"Bearer {HINDSIGHT_KEY}"},
        json={"content": text}
    )

    return response.json()
Enter fullscreen mode Exit fullscreen mode

When another incident arrives, the system searches the memory for something similar:

def find_similar(new_log):
    response = requests.post(
        f"{BASE_URL}/spaces/{SPACE_ID}/recall",
        headers={"Authorization": f"Bearer {HINDSIGHT_KEY}"},
        json={"query": new_log}
    )

    return response.json()
Enter fullscreen mode Exit fullscreen mode

The exact API endpoint and request format should be kept consistent with the current Hindsight documentation when the final implementation is deployed.

Recall first, generate second

One of the most important design decisions was putting memory before the language model's response.

The application first receives the new error log and calls the memory layer.

The retrieved incidents are then included in the prompt sent to Groq:

prompt = f"""
You are an on-call engineering assistant.

New incident: {log}

Similar past incidents from memory:

{similar}

Based on past incidents, suggest the likely root cause and fix.

If a past fix failed, do not suggest it again.
"""
Enter fullscreen mode Exit fullscreen mode

This changes the role of the language model.

Instead of only asking:

"What could fix this error?"

we are effectively asking:

"What could fix this error, considering what happened in similar incidents before?"

That difference is the main reason memory matters in this project.

Remembering failures is just as important

A particularly useful part of the workflow is the feedback step.

After the agent suggests a fix, the engineer can mark it as:

Worked

or

Failed

That result is stored along with the incident.

For example, imagine the agent suggests restarting a service to address repeated API timeouts.

If the engineer marks the solution as failed, that information becomes part of the incident history.

When a similar problem appears later, the previous failure can be retrieved along with the other relevant context.

This creates a feedback loop:

Incident → suggestion → outcome → memory → future suggestion

The system isn't simply storing conversations. It is storing useful information about incidents and their outcomes.

Before memory vs. after memory

The difference can be illustrated with a simple scenario.

Without memory

An engineer pastes:

Database connection timeout under heavy load
Enter fullscreen mode Exit fullscreen mode

The model generates a possible solution based on the current request.

There is no direct context about what the engineering team tried previously.

With memory

The same type of incident arrives.

Hindsight retrieves a previous incident:

Error: Database connection timeout after 500 requests
Root cause: Connection pool exhausted
Fix: Increased pool size from 10 to 50
Outcome: worked
Enter fullscreen mode Exit fullscreen mode

The language model now has concrete historical context that can influence its suggestion.

If another retrieved incident contains a failed fix, that information is also available.

The important change is not that the model suddenly becomes infallible. It is that the model is no longer operating with only the current error log.

Building an initial memory

A memory system is not very useful if it starts completely empty.

To give the application useful history, the project includes realistic example incidents covering different types of software problems, including database issues, memory leaks, API timeouts, deployment failures, and third-party service outages.

These incidents can be loaded into Hindsight before testing the application.

That means the first user interaction already has some history to work with.

The Streamlit interface then provides a simple workflow:

  1. Paste an error log.
  2. Click Get Suggestion.
  3. Retrieve similar incidents.
  4. Generate a suggested root cause and fix.
  5. Mark the result as Worked or Failed.

The interface is intentionally simple because the interesting part is what happens behind it: the interaction between memory and the language model.

What I learned

1. Memory needs useful structure

Simply saving an error message is not enough.

The combination of error, root cause, fix, and outcome gives the retrieved memory much more context.

The model can see not only what happened, but also what engineers believed caused it and whether the solution worked.

2. Failed solutions are valuable information

It is easy to focus only on successful fixes.

But knowing that a particular approach failed can be just as useful when dealing with a similar incident.

Recording the outcome turns user feedback into something the agent can use later.

3. Memory works best when it is part of the main loop

We didn't treat memory as a separate database that users occasionally inspect.

It is directly involved in the response process:

recall → reason → respond → feedback → retain

That makes memory part of how the agent operates rather than an extra feature attached to it.

4. Similar does not mean identical

There is an important limitation.

A previous incident can look similar to a new one without having the same root cause. Retrieved history should therefore be treated as context rather than absolute truth.

An engineer still needs to verify the suggested fix against the actual system.

What makes this approach interesting

The most interesting part of this project isn't simply connecting an LLM to an error log.

It is creating a feedback loop where previous engineering experience can influence future responses.

Traditional incident documentation often becomes something people search manually when a problem occurs. With an agent memory layer, that historical information can become part of the interaction itself.

The system can remember:

what went wrong, what was tried, and what happened next.

That gives the language model something much more useful than a blank conversation.

The project started with a simple question: What if an incident-response agent didn't have to forget everything after answering?

Using Hindsight as the memory layer provides one way to build that behavior.

The larger lesson for me is that an agent's usefulness is not only about how well its language model can generate an answer. It is also about what information the agent can carry forward into the next problem.


 # I Built an AI Incident Agent with Hindsight That Remembers What Failed

Top comments (0)