DEV Community

Mandadi Vennela Naga Sai
Mandadi Vennela Naga Sai

Posted on

I Gave My SRE Agent a Memory With Hindsight

Most incident-response agents can analyze an outage.

The harder question is: what happened the last time we tried that fix?

While building IncidentIQ, I wanted my incident agent to do more than analyze the current logs. I wanted it to remember previous incidents, understand which fixes worked, which failed, and which were only temporary, and use that experience when recommending what to do next.

That became the reason I added Hindsight as the persistent memory layer.

Instead of starting every incident from zero, IncidentIQ can look at what happened before and turn those experiences into evidence for the next decision.

The Problem With Starting Every Incident From Zero

Imagine a payments-api suddenly starts returning HTTP 503 errors.

The incident might look like this:

Service: payments-api
Severity: CRITICAL

Alert:
HTTP 503 surge

Logs:
Database connection pool exhausted.
Enter fullscreen mode Exit fullscreen mode

An LLM can look at those logs and suggest several possible fixes.

Restart the service.

Increase the connection pool.

Check database latency.

Inspect long-running queries.

But the LLM doesn't automatically know what happened during previous incidents in the same environment.

What if restarting the service worked once but only provided temporary relief?

What if increasing the database connection pool had already solved the same problem twice?

What if connection-timeout monitoring had previously prevented the same issue from recurring?

That history can change the recommendation.

This is the problem I wanted IncidentIQ to solve.

What I Built

IncidentIQ is an SRE incident-response system that combines a React and TypeScript frontend, a FastAPI backend, Hindsight for persistent memory, and Groq for AI reasoning.

I wanted memory to participate in the investigation itself, rather than being something separate from the incident workflow.

The basic architecture is:

React Frontend
      |
      v
FastAPI Backend
      |
      +----------> Hindsight
      |            Persistent Memory
      |
      +----------> Groq
                   AI Reasoning
Enter fullscreen mode Exit fullscreen mode

The frontend never talks directly to Hindsight or Groq.

The backend controls the memory and reasoning pipeline.

The main incident workflow is:

Current Incident
       |
       v
Hindsight Recall
       |
       v
Previous Incidents
       |
       v
Previous Actions + Outcomes
       |
       v
LLM Reasoning
       |
       v
Recommendation
       |
       v
Engineer Action
       |
       v
New Outcome
       |
       v
Hindsight Retain
Enter fullscreen mode Exit fullscreen mode

This creates a feedback loop instead of a one-time AI interaction.

Starting an Incident Investigation

The first step is describing the incident.

For example:

Service: payments-api
Severity: CRITICAL

Alert:
HTTP 503 surge

Logs:
Database connection pool exhausted.
Enter fullscreen mode Exit fullscreen mode

IncidentIQ's investigation workspace allows the engineer to provide the service, severity, alert, and logs.

It also provides quick presets for common operational problems such as:

DB connection pool
Redis failure
Auth failure
Certificate expiry
Enter fullscreen mode Exit fullscreen mode

Once the incident is submitted, the system doesn't immediately ask the LLM for an answer.

It first asks a more useful question:

Have we seen something like this before?

Giving the Agent a Memory

This is where Hindsight comes in.

IncidentIQ has a dedicated Hindsight service responsible for connecting the application to the Hindsight memory bank.

The actual implementation is small:

import os

from dotenv import load_dotenv
from hindsight_client import Hindsight


load_dotenv()

HINDSIGHT_API_KEY = os.getenv("HINDSIGHT_API_KEY")
HINDSIGHT_BASE_URL = os.getenv(
    "HINDSIGHT_BASE_URL",
    "https://api.hindsight.vectorize.io"
)
HINDSIGHT_BANK_ID = os.getenv(
    "HINDSIGHT_BANK_ID",
    "incidentiq-sre"
)

client = Hindsight(
    base_url=HINDSIGHT_BASE_URL,
    api_key=HINDSIGHT_API_KEY
)


def store_memory(content: str):
    return client.retain(
        bank_id=HINDSIGHT_BANK_ID,
        content=content
    )


def search_memory(query: str):
    return client.recall(
        bank_id=HINDSIGHT_BANK_ID,
        query=query
    )
Enter fullscreen mode Exit fullscreen mode

There are two operations that matter here:

recall()
    |
    v
Retrieve previous operational experience

retain()
    |
    v
Store new experience for future incidents
Enter fullscreen mode Exit fullscreen mode

You can explore Hindsight on GitHub and read the Hindsight documentation.

Recalling Relevant Incidents

When memory is enabled, IncidentIQ builds a query from the current incident:

query = f"""
Service: {request.service}
Severity: {request.severity}
Alert: {request.alert}
Logs: {request.logs}
"""

hindsight_response = search_memory(
    query
)
Enter fullscreen mode Exit fullscreen mode

The recalled memories are then filtered according to the affected service before being used for analysis.

The idea is straightforward:

Current Incident
       |
       v
Hindsight Recall
       |
       v
Relevant Memories
       |
       v
Historical Facts
       |
       v
Recommendation
Enter fullscreen mode Exit fullscreen mode

This is where the agent starts behaving differently from a stateless incident assistant.

The Interesting Part: Remembering What Worked

Retrieving an old incident is useful.

Retrieving the outcome of the previous solution is much more useful.

Suppose the historical memory contains:

Incident: INC-001

Action:
restart service

Outcome:
temporary
Enter fullscreen mode Exit fullscreen mode

Another memory contains:

Incident: INC-001

Action:
increase database connection pool size

Outcome:
successful
Enter fullscreen mode Exit fullscreen mode

And another contains:

Incident: INC-001

Action:
add connection timeout monitoring

Outcome:
successful
Enter fullscreen mode Exit fullscreen mode

Now the agent has more than historical context.

It has evidence about what actually happened.

The interface makes this pattern visible by connecting previous actions with their outcomes.

For the payments-api example, the system can identify that restarting the service was not the same as permanently resolving the underlying connection-pool problem.

Turning Memory Into Evidence

This was one of the most important design decisions in IncidentIQ.

I didn't want the LLM to decide historical success rates on its own.

The backend extracts historical facts from the recalled memories and calculates action statistics.

The project explicitly separates successful, failed, and temporary outcomes:

if fact["outcome"] == "successful":

    if incident_id:
        data["successful_incidents"].add(
            incident_id
        )

elif fact["outcome"] == "failed":

    if incident_id:
        data["failed_incidents"].add(
            incident_id
        )

elif fact["outcome"] == "temporary":

    if incident_id:
        data["temporary_incidents"].add(
            incident_id
        )
Enter fullscreen mode Exit fullscreen mode

The system then calculates the historical success rate from those explicit outcomes:

attempts = len(data["attempts"])
successful = len(
    data["successful_incidents"]
)

if attempts > 0:
    success_rate = successful / attempts
else:
    success_rate = None
Enter fullscreen mode Exit fullscreen mode

This means the application owns the historical evidence.

The LLM reasons over that evidence.

That separation matters because a generated explanation should not become the source of truth for historical statistics.

From Historical Evidence to a Recommendation

After Hindsight memories have been recalled and processed, IncidentIQ passes the current incident and historical memories to Groq.

The reasoning prompt explicitly tells the model to stay within the evidence:

prompt = f"""
You are an SRE incident response assistant.

Analyze the current incident using the historical
Hindsight memories provided below.

IMPORTANT:
- Base your reasoning on the evidence.
- Do not invent historical incidents.
- Do not invent evidence IDs.
- Recommend practical actions an on-call engineer can take.
- Confidence must reflect how strongly the evidence
  supports the diagnosis.
- Historical success rates should only be estimated
  from explicit outcome information in the memories.
- If there is insufficient historical evidence, say so.

CURRENT INCIDENT:
{incident}

HISTORICAL HINDSIGHT MEMORIES:
{memories}
"""
Enter fullscreen mode Exit fullscreen mode

The recommendation is also represented using a structured model:

class Recommendation(BaseModel):
    action: str
    rationale: str
    historical_success_rate: float | None = Field(
        default=None,
        ge=0.0,
        le=1.0,
    )
    drift_detected: bool = False
    evidence_ids: list[str] = []
Enter fullscreen mode Exit fullscreen mode

This gives the frontend structured information rather than one large block of generated text.

The Recommendation

For the connection-pool incident, the resulting recommendation can be:

Increase the database connection pool size
(e.g., from 50 to 100) and redeploy the
payments-api service.
Enter fullscreen mode Exit fullscreen mode

The recommendation is displayed together with its historical effectiveness, confidence, and evidence.

For example, the interface shows:

Historical Effectiveness:
100%

2 Successful / 2 Recorded

Confidence:
95%

Evidence:
INC-001
Enter fullscreen mode Exit fullscreen mode

This changes the interaction from:

AI:
"Try increasing the connection pool."
Enter fullscreen mode Exit fullscreen mode

to:

Historical Evidence:
2 successful / 2 recorded

Recommendation:
Increase the database connection pool.

Evidence:
INC-001
Enter fullscreen mode Exit fullscreen mode

The engineer can inspect the evidence and then decide what action makes sense.

What Changes With Memory?

The simplest way to understand the role of Hindsight is to compare the two workflows.

Without Memory

Incident
   |
   v
LLM
   |
   v
Possible Causes
   |
   v
Generic Recommendation
Enter fullscreen mode Exit fullscreen mode

With Hindsight

Incident
   |
   v
Hindsight Recall
   |
   v
Previous Incidents
   |
   v
Previous Actions
   |
   v
Previous Outcomes
   |
   v
LLM Reasoning
   |
   v
Evidence-backed Recommendation
Enter fullscreen mode Exit fullscreen mode

The LLM itself hasn't magically become better at debugging.

The context available to it has changed.

Memory doesn't replace reasoning. It gives reasoning something more useful to work with.

Using the Same Memory Before Deployment

I also wanted the memory layer to be useful before an incident happens.

IncidentIQ therefore includes a pre-deployment risk review.

An engineer can enter a proposed change:

Service:
payments-api

Proposed Change:
Increase connection pool 50 → 100
Enter fullscreen mode Exit fullscreen mode

[INSERT SCREENSHOT — Pre-Deploy Risk Review]

The system evaluates the proposed change using historical operational context.

The result can include a risk classification, risk score, rationale, safeguards, and related historical incidents.

This means the same memory layer can support two different moments in the software lifecycle:

Before Deployment
       |
       v
Historical Experience
       |
       v
Risk Assessment


During Incident
       |
       v
Historical Experience
       |
       v
Incident Recommendation
Enter fullscreen mode Exit fullscreen mode

The memory isn't tied to only one screen or one workflow.

Exploring What the Agent Remembers

I didn't want the memory layer to be invisible.

IncidentIQ includes a Memory Explorer where engineers can inspect the operational memories available to the system.

For example, the memory view can contain information such as:

Incident INC-001 was mitigated by restarting
the payments-api service and permanently resolved
by increasing the database connection pool size
and adding connection timeout monitoring.
Enter fullscreen mode Exit fullscreen mode

The memory explorer is useful for another reason: it makes the agent itself easier to inspect.

If a recommendation seems unexpected, an engineer can look at the memories that were available to the system.

Memory becomes something that can be investigated instead of a hidden part of the prompt.

Learning From What Actually Happened

A recommendation shouldn't automatically become a permanent truth just because an AI generated it.

The engineer still needs to record what happened.

IncidentIQ accepts an outcome and turns it into a new memory.

The backend builds an outcome memory like this:

outcome_memory = f"""
Incident {request.incident_id} affected the
{request.service} service.

Resolution attempt:
{request.action}

Outcome:
{outcome_description}

Engineer notes:
{request.notes}

Interpretation:
The engineer marked this resolution attempt as
{outcome_description}.
"""
Enter fullscreen mode Exit fullscreen mode

It then stores that experience:

result = store_memory(
    outcome_memory
)
Enter fullscreen mode Exit fullscreen mode

So the loop becomes:

Incident
   |
   v
Recall
   |
   v
Historical Evidence
   |
   v
Recommendation
   |
   v
Engineer Action
   |
   v
Real Outcome
   |
   v
Retain
   |
   v
Future Incident
Enter fullscreen mode Exit fullscreen mode

This is the part of IncidentIQ I find most interesting.

The agent isn't simply generating answers.

It has a mechanism for accumulating operational experience.

When There Isn't Enough Evidence

IncidentIQ also includes a Fix Drift view.

The idea is that a fix that worked previously should not automatically be considered reliable forever.

As more outcomes are recorded, the system can compare newer results with historical results.

But there is an important limitation.

The current interface explicitly reports when there aren't enough recorded outcomes to detect fix drift.

I prefer this behavior to manufacturing a conclusion.

If the system doesn't have enough evidence, the answer should be:

Not enough recorded outcomes to detect fix drift.
Enter fullscreen mode Exit fullscreen mode

That is much more useful than pretending the system knows something it doesn't.

What I Learned

  1. Memory Is More Useful When It Includes Outcomes

Remembering that an incident happened is useful.

Remembering what was tried and whether it worked is much more useful.

  1. Retrieval Should Be Visible

Historical context shouldn't disappear inside an LLM prompt.

Making memories and evidence visible gives engineers a way to inspect why a recommendation was produced.

  1. The LLM Shouldn't Own the Historical Facts

The application extracts historical outcomes and calculates explicit statistics.

The LLM reasons over those facts.

This separation makes the system easier to inspect.

  1. Insufficient Evidence Is Still Information

If there aren't enough recorded outcomes to detect a pattern, the system should say so.

  1. The Feedback Loop Is the Real Memory

The interesting architecture isn't:

Incident → AI → Answer
Enter fullscreen mode Exit fullscreen mode

It's:

Incident
   |
   v
Memory
   |
   v
Recommendation
   |
   v
Action
   |
   v
Outcome
   |
   v
Memory
Enter fullscreen mode Exit fullscreen mode

That is what turns previous incidents into experience that can influence future incidents.

Final Thoughts

IncidentIQ started from a simple observation:

SRE teams repeatedly solve problems that have similar histories.

An LLM can reason about the incident in front of it.

But persistent memory gives that reasoning access to what happened before.

That changes the question from:

"What should I do?"
Enter fullscreen mode Exit fullscreen mode

to:

"What happened the last time this occurred,
what actually worked,
and what evidence do we have?"
Enter fullscreen mode Exit fullscreen mode

That's what I wanted to explore with IncidentIQ: an SRE agent that doesn't just respond to incidents, but can build a persistent memory of operational experience.

The complete source code for IncidentIQ is available on GitHub:

GitHub logo Pravallika2789 / IncidentIQ-SRE-Agent

AI-powered SRE incident response agent that uses Hindsight memory to learn from past incidents, successful and failed fixes, and deployment history.

IncidentIQ - Intelligent SRE Incident Response

IncidentIQ is an intelligent SRE incident-response agent that uses Hindsight persistent memory to help engineers investigate production incidents, recall previous operational experience, learn from successful and failed resolution attempts, and make safer deployment decisions.

It combines a React frontend, FastAPI backend, Hindsight for persistent operational memory, and Groq-powered AI reasoning to turn past incidents into reusable engineering knowledge.

Team Code & Chaos

1. Reddy Gayatri Satya Sai Pravallika

2. Mandha Varshitha

3. Chenna Keerthana

4. Mandadi Vennela

Video Presentation

Watch the IncidentIQ Demo

Problem

When a production incident happens, engineers need more than an analysis of the current logs.

They need to know:

  • Has this happened before?
  • What caused the previous incident?
  • Which fix actually worked?
  • Which approaches failed?
  • Which runbook was useful?
  • What did the team learn from previous incidents?
  • Has a previously successful fix become less effective?

Solution

IncidentIQ turns previous incidents…




Top comments (0)