DEV Community

Cover image for My Agent Remembers Every Incident, Thanks to Hindsight
Rehan
Rehan

Posted on

My Agent Remembers Every Incident, Thanks to Hindsight

TL;DR: I built TraceMind, an incident-response agent that retains every resolved incident as a post-mortem in Hindsight agent memory and recalls similar ones when a new alert fires. The interesting decision wasn't the memory layer — it was refusing to let the agent sound certain. Every recommendation is framed as a similarity hypothesis with a confidence note — the framing I'd want to see if I were the one on call.

TraceMind memory loop

The 2 AM problem

Picture it: it's 2 AM, the payments API is throwing 500s after a deploy, and you're sure you've seen this before. Somewhere there's a fix — maybe a one-line connection-pool change — buried in old chat threads and a wiki page last updated eight months ago. The knowledge exists. It just isn't retrievable at 2 AM by a tired human.

That gap — between "we solved this before" and "I can find how we solved this" — is what I built TraceMind to close.

What TraceMind does

TraceMind watches GitHub issues, treats them as incidents, and runs a four-step loop: fire, recall, recommend, resolve and retain.

When a new incident fires, the agent searches its memory of past post-mortems for similar error signatures and symptoms, then recommends investigation steps drawn from what actually worked last time — the commands engineers ran, the root causes they found, the fixes that stuck, and how long each took. When the incident is resolved, the resolution is retained as a new memory, so the next similar incident starts from a better place.

The memory layer is Hindsight, Vectorize's agent memory system. Incident knowledge is exactly the kind of thing that should compound: every outage should make the next one cheaper. The Hindsight docs got me from zero to a working retain/recall loop in an afternoon.

TraceMind memory loop architecture
The loop: every incident is recalled against, recommended from, and retained into Hindsight memory.

TraceMind live app
TraceMind live at tracemind.run.place — Trace's chat interface, where incidents are fired, recalled against past post-mortems, and resolved.

The core decision: the agent never diagnoses

Here's the through-line of the whole project, and the thing I'd defend in a design review: TraceMind never tells you the root cause. It tells you what similar incidents turned out to be, how similar they are, and what to verify.

Most incident bots I've seen do the opposite. They take an alert, run it through an LLM, and announce "Root cause: database connection exhaustion" with total confidence. Sometimes they're right. When they're wrong — and with correlated-but-distinct failure modes, they're wrong often — engineers learn to ignore the bot entirely. False certainty is worse than no information, because it spends the scarcest resource in an incident: attention.

So I built the recommendation engine around calibrated uncertainty. The investigate flow is three calls:

def investigate(self, alert: dict[str, Any], top_k: int = 3) -> dict[str, Any]:
    """Full assist flow for a new incident: analyze, recall, recommend."""
    analyzed = self.analyze(alert)
    draft = analyzed["draft"]
    matches = self.memory.search_similar(analyzed["query_text"], top_k=top_k)
    recommendation = self._recommend(draft, matches)
    ...
Enter fullscreen mode Exit fullscreen mode

Recall returns scored matches against retained post-mortems. The recommendation headline reads something like: "87% similar to INC-0007 — auth-service cascading timeouts from Stripe API." Not "the root cause is." Every recommendation ships with a confidence note, generated in code rather than by the LLM:

"confidence_note": (
    f"Based on {len(matches)} similar past incident(s). "
    f"Top match similarity: {top['score']:.0%}. "
    "Verify before acting — similar symptoms can have different root causes."
),
Enter fullscreen mode Exit fullscreen mode

That last sentence is the whole product philosophy in eleven words. The agent is an extremely well-read junior on-call engineer: it has strong recall of everything the team has ever fixed, and it knows the limits of its own knowledge.

The cold-start path gets the same treatment. When memory has nothing similar, the agent says so plainly — "I don't have enough historical context for this signature" — and falls back to generic first-response steps. It never bluffs. And it closes the loop: "Resolve this incident and I'll remember it — the next similar one will get a memory-backed recommendation."

What actually gets remembered

A memory system is only as good as what you put in it. Early on I retained free-text summaries and the recall quality was mediocre — vague summaries match vaguely. The fix was structuring the post-mortem at retain time:

def _postmortem_text(incident: dict[str, Any]) -> str:
    dep = incident.get("deployment", {}) or {}
    lines = [
        f"Incident {incident.get('id')}: {incident.get('title')}",
        f"Service: {incident.get('service')} | Severity: {incident.get('severity')}",
        f"Error signature: {incident.get('error_signature')}",
        "Symptoms: " + "; ".join(incident.get("symptoms", [])),
        f"Deployment at incident time: {dep.get('version')} (deployed {dep.get('deployed_at')})",
        f"Root cause: {incident.get('root_cause')}",
        "Investigation steps: " + "; ".join(incident.get("investigation_steps", [])),
        "Commands used: " + "; ".join(incident.get("commands_used", [])),
        f"Fix: {incident.get('fix')}",
        f"Resolved by {incident.get('engineer')} in {incident.get('mttr_minutes')} minutes.",
    ]
    return "\n".join(lines)
Enter fullscreen mode Exit fullscreen mode

The boring fields are the valuable ones. "Commands used" and "resolved in 40 minutes" are what turn a recommendation from trivia into an action plan. When the agent suggests a specific kubectl command, it's because an engineer ran exactly that command during a past incident and it worked — not because an LLM thinks it sounds plausible.

Each retained memory gets a stable document ID (incident-INC-0007), so re-resolving or re-seeding the same incident updates the record instead of duplicating it. And the retain is synchronous — no ingestion lag between "incident resolved" and "memory available":

self._bridge.call(
    "aretain",
    bank_id=self.bank_id,
    content=_postmortem_text(incident),
    context="production incident",
    timestamp=incident.get("resolved_at"),
    document_id=f"incident-{inc_id}",
    metadata={...},
    retain_async=False,  # synchronous: no ingestion lag
)
Enter fullscreen mode Exit fullscreen mode

That was a deliberate choice. In the early days, when every new incident is a potential first match, a few minutes of ingestion delay is the difference between the system working and looking broken.

A walkthrough

Here's a walkthrough using one of the seeded post-mortems, INC-0007: "auth-service cascading timeouts from Stripe API." The retained record holds the error signature (downstream Stripe p99 latency saturating the thread pool), the root cause (no bulkhead isolation; a 10s timeout with 3 retries amplifying the load), the exact commands the engineer ran (jstack thread counts, a Stripe status check, kubectl set env to cut timeouts and retries), the fix (2s timeout, single retry with jittered backoff, circuit breaker, graceful degradation), and the resolution time: 40 minutes.

Walk through what happens when a new issue fires with similar symptoms — checkout errors climbing, threads saturating. The agent recalls INC-0007 as the top match, and the recommendation says: here's the match, here's what worked last time, here are the exact commands, here's how long it took — and here's the caveat that similar symptoms can have different root causes, so verify the downstream latency before changing anything. When the engineer resolves it, the new incident — with its own root cause, commands, and timing — goes into memory for the next one.

That's the compounding effect. The seeded post-mortems aren't just a dataset; they're the beginning of an institutional memory that doesn't quit, forget, or go on vacation.

What I'd do differently

Seed with real post-mortems, not reconstructed ones. Structured fields beat prose, but the highest-signal fields — the actual commands, the actual timeline — only exist if engineers record them during the incident. I'd integrate the retain step into the incident channel itself, capturing commands as they're run rather than reconstructing them afterward.

Similarity isn't causality, and the UI has to keep saying so. I put the caveat in the confidence note, but I keep being tempted to make it more prominent. The failure mode I worry about most is over-trusting a high similarity score — the framing is load-bearing, and it deserves as much design attention as the recall quality.

Build the boring storage first. I built the Hindsight integration before the local SQLite fallback was solid. Debugging the memory layer without a local store to compare against is miserable — the hybrid store (local SQLite as the database of record, Hindsight for semantic recall) is the right shape. Build the boring half first.

The takeaway

The lesson I keep coming back to: for an agent that advises humans under pressure, honesty is a feature, not a limitation. "87% similar to INC-0007, verify before acting" gets used. "Root cause: connection exhaustion" gets ignored after the first miss. Memory made the agent knowledgeable; refusing to overclaim made it trustworthy.

If you're building agents that touch production systems, give them a memory — Hindsight is a solid place to start — and then spend as much effort on how the agent expresses uncertainty as you do on making it knowledgeable. The engineers on call at 2 AM will thank you.

Top comments (0)