DEV Community

Mohit Kumar
Mohit Kumar

Posted on

Hindsight Taught My SRE Agent to Stop Repeating Failed Fixes

At 3 a.m., the worst thing an on-call tool can do is suggest the fix that already failed last Tuesday. Most incident agents do exactly that, because every alert starts from a blank context window.

I built RunbookMind to fix that one specific problem: an SRE agent that remembers what actually worked, what didn't, and reranks its advice accordingly. The memory layer is Hindsight, and it ended up shaping the whole design.

What the system does

RunbookMind takes a high-severity alert, pulls in logs and deploy history, generates root-cause hypotheses, and returns a ranked list of remediation steps. Then a human tells it what happened: worked, partial, or failed. That feedback goes back into memory, and the next similar incident is ranked differently because of it.

The stack is deliberately boring:

Backend: FastAPI, Pydantic v2, SQLite for metrics
LLM: Groq, openai/gpt-oss-120b as primary with a fallback model
Memory: a Hindsight bank (devnote) accessed through a thin wrapper
UI: React + Vite + Tailwind, with a "memory on/off" toggle

The API surface is four endpoints: POST /alert, POST /feedback, GET /metrics, and GET /health, plus /reflect for cross-incident patterns.

The request path is short. An alert comes in, I recall similar past incidents, I hand alert plus recalled memory to the model, let it call up to three tools (search_logs, get_deploy_history, get_runbook) over at most four steps, parse the JSON, and rerank. The feedback path is the mirror image: format the resolution, retain it.

The through-line: outcomes are the memory, not incidents

My first instinct was to store incident descriptions and let similarity search do the rest. That gives you an agent that says "this looks like INC-2041" and then confidently suggests the same fix whether it worked or not.

The useful thing to remember isn't the incident. It's the outcome of each attempted fix. So the retained record is structured text with the rejected fixes written down explicitly:

text
INCIDENT INC-2041 | service: checkout-api | severity: SEV1
Symptoms: p99 latency 4.2s, Redis evicted_keys spiking, 500 errors
Root cause: Redis maxmemory too low after traffic growth
Fix applied: raised maxmemory 2GB->4GB, set allkeys-lru
Outcome: worked | time to resolve: 14 min
Rejected fixes: restart cache (failed), scale API pods (no effect)

That last line is the one that matters. "Restart cache (failed)" is exactly what a tired human or a naive agent would try first. Storing it means recall can surface it as a negative signal.

The wrapper around Hindsight is small enough to read in one sitting:

python
class IncidentMemory:
def recall_similar(self, alert: Alert, k: int = 5) -> list[RecalledMemory]:
query = f"{alert.service} {alert.summary} {alert.raw_log_snippet or ''}"
return hindsight.recall(bank_id=self.bank_id, query=query)

def retain_resolution(self, alert: Alert, fb: Feedback) -> None:
    hindsight.retain(bank_id=self.bank_id, content=format_resolution(alert, fb))

def reflect(self) -> str:
    return hindsight.reflect(bank_id=self.bank_id,
        query="What recurring patterns and best fixes have emerged?")
Enter fullscreen mode Exit fullscreen mode

Three calls: retain, recall, reflect. This is the main reason I used Hindsight's agent memory instead of building my own vector table. I didn't want to design chunking, embedding refresh, and a retrieval strategy just to get to the interesting part, which is what the agent does with what it remembers. The Hindsight docs cover the client details, and I'd check them before copying my wrapper, since method signatures matter.

Reranking: don't let the LLM grade its own homework

Recall alone isn't enough. If you dump five past incidents into a prompt, the model will still lean on its priors. So after the model proposes fixes, I rerank them deterministically using the outcomes from memory:

python
def score(fix, recalled):
history = [m for m in recalled if fix.matches(m)]
successes = sum(1 for m in history if m.outcome == "worked")
partials = sum(1 for m in history if m.outcome == "partial")
total = len(history) or 1
success_rate = (successes + 0.5 * partials) / total
recency_boost = 1.1 if any(m.age_days < 30 for m in history) else 1.0
return (0.6 * fix.confidence + 0.4 * success_rate) * recency_boost

The model's confidence gets 60% of the weight, historical success rate gets 40%, and anything that worked in the last 30 days gets a small boost. Partial successes count half. It's crude, and I'll say more about that below, but it's inspectable. When an SRE asks "why is this fix ranked first?", the answer is a formula and a list of incident IDs, not a vibe.

The prompt backs this up. The system message tells the model to use only recalled incidents and tool output as evidence, cite incident IDs, say so when nothing similar exists, and never invent past incidents. The response schema carries past_success_rate and source_incident_ids on every fix, so the UI can show why something is ranked where it is.

Making the effect visible: memory on/off

The most useful debugging feature I added is a memory_enabled flag on /alert. With it off, the agent behaves like every other stateless incident tool. With it on, it recalls and reranks.

Every alert is logged to SQLite with that flag, plus top1_correct, top3_correct, and minutes_to_resolve. The learning curve in the UI is just rolling accuracy over incident order. Nothing fancy, but it turns "I think it's getting better" into something I can look at, and it lets me run the same held-out alert both ways.

What it looks like in practice

Take the Redis-eviction family in the seed data. A new alert arrives for checkout-api: p99 latency climbing, 500s rising, and eviction counters spiking in the log snippet.

Memory off: the model produces reasonable but generic advice. Restarting the cache tends to come up first, because that's what the internet says to do.

Memory on: recall surfaces INC-2041. The record says the restart failed and that raising maxmemory and switching to allkeys-lru worked in 14 minutes. The reranker pushes the maxmemory fix to the top, the restart drops down with its failure history attached, and the response cites the source incident.

Then the on-call engineer marks the outcome. If the fix worked, that success rate goes up. If it didn't, that's retained too, and the next alert sees it. /reflect can then be asked what recurring patterns have emerged across the bank.

Failure handling I'm glad I did up front

LLM plumbing breaks in boring ways, so llm.py is a small state machine:

Call the primary model.
On a malformed tool call or invalid JSON, send one repair prompt with the parse error.
Still failing, retry on the fallback model.
Still failing, return a degraded response containing just the recalled incidents, with a UI warning.
20-second timeout per call, exponential backoff, max 3 retries.

Step 4 is the one I care about most. If the model is down but memory isn't, the on-call engineer still gets "here are the closest past incidents and what happened." That's useful on its own.

Lessons learned

  1. Store outcomes and rejected options, not just descriptions. The negative examples are what stop repeated mistakes. A memory of only what worked is half a memory.

  2. Structure your retained text. Fixed fields (symptoms, root cause, fix, outcome) make recall more consistent and make the records readable by humans when you audit them.

  3. Keep ranking outside the LLM. A small deterministic scorer over recalled outcomes is easier to debug, test, and explain than asking the model to weigh history in-prompt.

  4. Ship the off switch. A memory on/off toggle plus per-alert accuracy logging is how you find out whether memory is helping or just adding latency and confident-sounding noise.

  5. Know what's still crude. Two honest limitations: fix.matches(m) is fuzzy, since two fixes can be worded differently but be the same action, and the 0.6/0.4 weights and the 30-day boost are hand-picked rather than tuned. With a sparse memory bank, one lucky success can outweigh reality. The fix is more history, better fix normalization, and tuning the weights against held-out incidents.

Where this goes

The direction I care about is the feedback loop getting tighter: pulling resolution outcomes straight from ticketing and postmortems instead of waiting for a button press, and using reflection to turn a pile of incidents into runbook updates. The stateless version of this tool is fine on day one. The remembering version is what you want on day ninety.

If you want to poke at it, the code is on GitHub. And if you're building anything agentic where "it keeps making the same mistake" is a complaint, look at Hindsight on GitHub. Retain, recall, reflect is a small enough interface to try in an afternoon.

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.