DEV Community

Harshini Sai Makkapati
Harshini Sai Makkapati

Posted on

How I Built an Incident Response Agent That Remembers What Killed Your Service Last Time

Every on-call engineer has lived this: it's 2am, payment-api is down, and you're staring at a dashboard trying to remember — did we see this before? What did we try? What actually worked?

The answer is almost always yes, you've seen it before. But that knowledge lives in a Slack thread from three months ago, a runbook nobody updated, or the head of the engineer who's now on vacation. So you start from scratch. You try restarting Redis. It doesn't work. You waste 20 minutes. You finally increase the database connection pool. Latency drops immediately. You write a note to yourself and forget about it.

I built IncidentRecall to break that loop. It's an incident response agent that uses Hindsight agent memory to recall what happened in similar past incidents, tell you what worked, and — critically — warn you what failed so you don't waste time on it again.


What the System Does

The core flow is simple:

Incident → Recall similar memories → LLM recommendation → Engineer feedback → New memory → Future incidents learn
Enter fullscreen mode Exit fullscreen mode

An engineer submits an incident — service name, description, symptoms, severity. The system queries Hindsight for semantically similar past incidents. Those memories become the evidence base for a structured LLM recommendation. The recommendation includes:

  • Matched historical incidents — which past incidents are similar and why
  • What to do — grounded in what actually worked before, with a confidence count
  • What NOT to try first — actions that failed in similar situations, with evidence IDs

When the engineer resolves the incident, they submit the outcome — worked or failed. That outcome gets written back to Hindsight. The next similar incident retrieves it.

The loop closes. The agent gets smarter with every incident.

Incident form
The incident form. Symptom tags are clickable — select all that apply. The agent uses these to build the Hindsight recall query.


The Memory Layer

The most important architectural decision was keeping the memory layer thin and explicit. The rest of the application — the API, the recommendation engine — only ever calls two functions:

# memory.py

def recall(query: str) -> list:
    """Retrieve relevant memories for a query."""
    results = _hindsight.recall(
        bank_id=BANK_ID,
        query=query,
        budget="low",
        max_tokens=2000,
    )
    return results.results


def retain(content: str, context: str = "incident resolution") -> None:
    """Store a memory."""
    _hindsight.retain(
        bank_id=BANK_ID,
        content=content,
        context=context,
    )
Enter fullscreen mode Exit fullscreen mode

That's the entire interface. recall() and retain(). Nothing else leaks through.

This matters because agent memory is not a database. You don't query it with SQL. You query it with natural language, and what comes back is semantically relevant text. The quality of what you get out depends entirely on the quality of what you put in.

The recall query is built to be natural, not keyword-stuffed:

query = (
    f"Service: {req.service}\n"
    f"Problem: {req.description}\n"
    f"Symptoms: {', '.join(req.symptoms)}\n"
    f"Find similar incidents, what fixes worked, what failed, root causes."
)
Enter fullscreen mode Exit fullscreen mode

And the memory written back after feedback is structured like a real incident report — not a JSON blob, not a summary, but a narrative with explicit [WORKED] and [FAILED] labels that the LLM can parse unambiguously:

memory_content = (
    f"Incident ID: {req.incident_id}\n"
    f"Service: {inc['service']}\n"
    f"Symptoms: {', '.join(inc['symptoms'])}\n"
    f"Root Cause: {req.root_cause or 'Not specified'}\n"
    f"\n"
    f"Actions Attempted:\n"
    f"  - [{outcome_label}] {req.action}\n"
    f"    Notes: {req.notes or 'None'}\n"
    f"\n"
    f"{'Successful Fix: ' + req.action if req.result == 'worked' else 'Failed Actions (do not try first): ' + req.action}\n"
)
Enter fullscreen mode Exit fullscreen mode

The [WORKED] / [FAILED] labels are load-bearing. The recommendation engine parses them to build the do_not_try_first list. If you write vague memories, you get vague recommendations.

Memory Inspector
The Memory Inspector shows exactly which memories Hindsight retrieved and their full content. INC-031D36 is a memory written during this session — the learning loop working in real time.

Memory Inspector full
The raw memory text written to Hindsight after engineer feedback. The structured narrative format — with explicit [WORKED] labels — is what makes future recall useful.


The Recommendation Engine

The LLM's job is not to be creative. It's to be a structured parser of evidence.

The prompt is explicit about this:

HISTORICAL MEMORIES FROM HINDSIGHT (your ONLY evidence source):
These are real past incidents retrieved from the memory bank. Use them as your evidence.
Do NOT invent incidents. Do NOT use general knowledge to claim something worked or failed.
If the memories do not contain relevant evidence, say so.
Enter fullscreen mode Exit fullscreen mode

The LLM returns a structured JSON object with matches, a recommendation, a do_not_try_first list, and a confidence count. The confidence count is not a probability — it's a raw count: "this fix worked in 3 out of 3 similar incidents." That's more honest and more useful than a percentage.

The recommendation engine also runs a baseline — the same incident, but with no memories passed in. This gives you a side-by-side comparison: what would a generic SRE recommendation look like versus what the memory-grounded agent recommends? The difference is the value of memory made visible.

Recommendation panel
The recommendation panel. Green = what to do, grounded in historical evidence. Red = what NOT to try first, with the exact incident IDs where it failed.


What the Learning Loop Actually Looks Like

Here's a concrete example. payment-api goes down with high latency and database connection exhaustion. Without memory, the agent says: check recent deployments, review logs, follow standard runbook. Useful, but generic.

With memory, it says:

Based on 4 similar historical incidents (INC-1002, INC-1005, INC-1008): Increase database connection pool from 50 to 150. This fix worked in 3 of 4 similar incidents retrieved.

Avoid First: Restart Redis cache — failed in INC-1002, INC-1005, INC-1008. Restart failed in 3/3 similar incidents.

That's the difference. The agent knows that restarting Redis is the instinctive wrong move for this service, because it's been tried and failed three times. It saves you 20 minutes at 2am.

Historical matches
4 historical matches retrieved from Hindsight. Each card shows root cause, what worked (green), what failed (red), and the deployment change that triggered it.

After the engineer resolves the incident, they submit the outcome via the feedback form:

Feedback form
The feedback form. The action field pre-fills from the recommendation. Both "Fix Worked" and "Fix Failed" outcomes are stored — failed outcomes are what power the Avoid First warnings.

# POST /feedback
{
  "incident_id": "INC-A1B2C3",
  "action": "Increased database connection pool from 50 to 200",
  "result": "worked",
  "notes": "Latency returned to normal within 2 minutes.",
  "root_cause": "Connection pool undersized for current traffic"
}
Enter fullscreen mode Exit fullscreen mode

The next similar incident retrieves this memory. The confidence count goes up. The pattern strengthens.

Memory updated
The Memory Updated confirmation. This outcome is now in Hindsight. The next similar incident will retrieve it.


Why Hindsight

I looked at a few options for the memory layer. The reason I chose Hindsight was the retain / recall API design. It's opinionated in the right way — you write natural language in, you get semantically relevant natural language back. There's no schema to maintain, no embedding pipeline to manage, no vector database to operate.

For an incident response use case, that matters. Incidents are described in natural language. The fixes are described in natural language. The recall query is natural language. Forcing all of that through a structured schema would lose information. Hindsight keeps the full fidelity of the text and handles the semantic retrieval.

The other thing that mattered: both retain() and recall() are single function calls. The integration is genuinely thin. The memory layer in this project is about 50 lines of real code. That's the right size for infrastructure that should be invisible.


Lessons Learned

1. Write memories like incident reports, not like database records.
The quality of recall depends on the quality of what you retain. Structured narrative with explicit outcome labels ([WORKED], [FAILED]) gives the LLM unambiguous signal. Vague summaries produce vague recommendations.

2. Store failed outcomes, not just successful ones.
The do_not_try_first list is only possible because failed actions are stored. Most systems only record what worked. That's half the information. Knowing what failed — and in which specific incidents — is often more valuable than knowing what worked.

3. The baseline comparison is the best demo you can give.
Showing the memory-grounded recommendation alongside the no-memory baseline makes the value of agent memory immediately obvious. Without it, you're asking people to imagine the difference. With it, they can see it.

4. Keep the memory interface minimal.
recall() and retain(). That's it. The rest of the application doesn't need to know how memory works. This made it trivial to swap between a live Hindsight backend and a deterministic mock for development — same interface, different implementation behind a single environment variable.

5. Confidence counts beat confidence percentages.
"Worked in 3/3 similar incidents" is more credible than "87% confidence." Engineers are skeptical. Show them the evidence count, not a number that sounds like it came from nowhere.


What's Next

The current system stores memories per-session in the demo layer and per-bank in live mode. The natural next step is per-service memory banks — so cart-service incidents don't pollute payment-api recall results, and each service builds its own institutional knowledge over time.

The other obvious extension is automatic pattern detection: if the same action fails three times for the same service, surface a proactive alert before the next incident even happens.

The foundation for both is already there. The memory layer is the hard part. Once you have reliable recall and retain, the rest is just building on top of it.


If you're building agents that need to learn from past interactions — not just retrieve static documents, but actually accumulate operational knowledge over time — Hindsight is worth looking at. The retain / recall model is the right abstraction for this class of problem.

The full project is on GitHub. The backend is FastAPI + Groq + Hindsight. The frontend is React + Vite. Demo mode requires no API keys — you can run the full learning loop locally in under five minutes.

Top comments (0)