The Incident Copilot Lesson I Didn't Expect: Remember the Failures Too
At 3 a.m., the costliest mistake an on-call engineer can make is a sensible-looking action that deepens the outage. Restarting pods is sensible. Adding replicas is sensible. For one class of incident, both are exactly wrong, and the only safeguard is a teammate who remembers how badly it went last time.
I built an incident copilot to be that teammate. It takes a new alert, recalls similar past incidents, and produces a plan that leads with what worked and spells out what to avoid. The memory layer is Hindsight, and the decision that paid off most was an unglamorous one: I store failures with the same care as fixes.
What the system does
Incident Copilot is a small FastAPI service with a single-page UI. The layout is intentionally plain:
Browser UI (app/static/index.html)
|
FastAPI (app/main.py)
/plan /action /close /patterns
| |
agent.py memory.py ----> Hindsight (retain / recall / reflect)
recall -> 1 LLM call
|
Groq: openai/gpt-oss-120b
Four endpoints carry the whole workflow:
• POST /plan takes an alert and its symptoms, recalls memories, and returns a structured plan.
• POST /action records the outcome of a step (worked, failed, harmful) while the incident is still open.
• POST /close retains the complete postmortem after resolution.
• GET /patterns asks Hindsight to reflect on the recurring root causes and failed fixes it has accumulated.
Every Hindsight call sits in app/memory.py. The agent is a single recall followed by a single LLM call: no planner loop, no tool-calling maze, and no vector-store plumbing for me to run. I wanted memory to be the interesting part and everything around it to be readable in one sitting.
If agent memory is a new concept, Vectorize has a solid explainer on what agent memory is. In short, the model stays stateless and something else holds the history. I chose Hindsight for that role because its three operations, retain, recall, and reflect, line up neatly with how on-call work actually flows. The Hindsight docs describe the full API; I use only those three.
The through-line: outcomes are data
Most postmortem tooling is built around the resolution. What fixed it? Record that, since it's what you'll want next time.
That's only half true. In practice the resolution is usually the last of several attempts, and the earlier attempts consume the time. A typical incident on checkout-service runs like this: restart the pods (harmful, the reconnect storm exhausted the pool again), scale to 12 replicas (harmful, more replicas opened more DB connections and hit max_connections), roll back config deploy #4471 (worked, the pool size had been cut from 50 to 10). Reaching the rollback took 82 minutes, and most of that was the two wrong moves.
A memory holding only "rolled back deploy #4471" saves the next responder nothing. A memory that also holds the two harmful attempts and why each failed lets them skip both.
So each action in a retained incident carries an explicit outcome label, written into the text at retain time:
def retain_incident(inc: dict):
actions = "\n".join(
f"- Attempted: {a['action']}. Outcome: {a['outcome'].upper()}. {a['note']}"
for a in inc["actions"]
)
content = (
f"Incident {inc['id']} on {inc['service']} ({inc['severity']}).\n"
f"Alert: {inc['alert']}\nSymptoms: {inc['symptoms']}\n"
f"Actions:\n{actions}\n"
f"Root cause: {inc['root_cause']}\nResolution: {inc['resolution']}\n"
f"Time to resolve: {inc['ttr_minutes']} minutes."
)
_retain(
bank_id=BANK,
content=content,
context="incident postmortem",
timestamp=inc["resolved_at"],
document_id=f"incident-{inc['id']}",
tags=[f"service:{inc['service']}", f"severity:{inc['severity']}"],
metadata={"incident_id": inc["id"]},
)
Three choices here deserve a defense.
WORKED, FAILED, and HARMFUL are uppercase words in the prose, not just JSON fields. Hindsight builds memories by extracting from natural language, so the outcome has to survive that extraction. "Restarting pods was harmful" should be a fact the memory layer holds, not metadata that gets discarded.
timestamp is when the incident was resolved, not when I ingested it. This matters more than it seems. Backfill old postmortems with a default timestamp of now and every memory looks fresh, so the agent can't distinguish a six-month-old runbook from last week's. Real dates are what make staleness detection possible.
document_id is stable. It's incident-, so retaining the same postmortem again replaces it rather than duplicating it. Postmortems get edited: typos fixed, follow-ups added. Without an idempotent key, each edit would create a second, conflicting copy, and recall would return both.
The same principle applies during an incident. The UI puts Worked / Failed / Harmful buttons on every recommended step, and a click calls /action, which retains a short note right away:
def log_action(incident_id: str, n: int, action: str, outcome: str, note: str = ""):
"""Mid-incident feedback: the Worked / Failed / Harmful buttons."""
_retain(
bank_id=BANK,
content=f"During incident {incident_id}, attempted: {action}. Outcome: {outcome.upper()}. {note}",
context="live incident action",
document_id=f"incident-{incident_id}-action-{n}",
)
If one engineer marks "restart pods" as harmful at 03:10, a second engineer paged for a related alert at 03:25 already sees it. Nobody waits for the postmortem.
Recall across services on purpose
The design choice I'd defend hardest is what I leave out at recall time: there is no service filter.
def recall_for_alert(alert: str, symptoms: str = ""):
query = f"{alert} {symptoms}".strip()
try:
resp = client.recall(bank_id=BANK, query=query, budget="high", max_tokens=3000)
except TypeError:
resp = client.recall(bank_id=BANK, query=query)
return resp.results
The query is only the alert text plus symptoms. The service name is in there, but nothing requires a match on it. That's deliberate, because an outage's cause is rarely a property of the service. A config deploy that shrinks a connection pool looks identical on checkout-service, orders-api, and payments-api, and the stack trace even reads the same:
HikariPool-1 - Connection is not available, request timed out after 30000ms
My test history contains four incident families: connection pool exhaustion, cache stampede after a Redis failover, expired internal mTLS certificates, and disks filled by unrotated logs. Each family spans multiple services. The held-out incident I use most is on payments-api, while earlier incidents with the same root cause are on checkout-service, orders-api, and inventory-service. A per-service index returns nothing. Recall keyed on symptoms returns all three.
I set the recall budget to "high" with 3,000 tokens. That's a conscious cost trade: recall runs once per alert, mid-incident, when latency matters far less than surfacing the right memory. If it ran per keystroke, I'd choose differently.
One bit of honest friction: the try/except TypeError is there because the Hindsight client's keyword arguments changed between versions I worked with, and I was tired of upgrades breaking the app. _retain has the same shim, dropping tags and metadata if a client rejects them. I prefer one slightly ugly compatibility layer to version checks sprinkled across the codebase. Pin your client version and keep the wrapper in one findable place.
Making the model use it honestly
Recall hands the model memories, but doesn't ensure it uses them well. The system prompt is where I constrain that:
SYSTEM = """You are an on-call incident copilot. You get a NEW ALERT and MEMORIES from past incidents.
Rules:
- Put fixes that worked in past similar incidents first.
- Explicitly list actions that were HARMFUL or FAILED in similar incidents under "avoid", with the reason.
- Cite the incident id for every claim. Never invent incident ids.
- If a memory is older than 6 months (compare with TODAY), add a staleness note.
- If memories are NONE, weak or unrelated, say so and give general advice, clearly labelled as general advice with source null. Return ONLY JSON: ...""" Three of those rules target failure modes I care about:
- Citations on every claim. Each step in the UI has a source link, and clicking it highlights the matching card in the "Memories used" panel. An engineer under pressure shouldn't have to trust the model. One click should be enough to check.
- A dedicated avoid list. Without it, harmful actions get buried in prose or dropped. A required structured field forces the model to surface them.
- Explicit "I don't know" behavior. When recall finds nothing relevant, the plan says so and labels its advice as general, with a null source. A confident answer built on no evidence is worse than none. The staleness rule works thanks to the resolution timestamp I insisted on earlier. The agent receives TODAY in the prompt and each memory's date in context, so "this was true in February" is something it can reason about. Behavior: what a plan looks like Here's the alert I use most: payments-api authorization requests stalling, queue depth 3.2k and rising Card authorizations queue up while pod CPU stays low. Traces show requests idle waiting on the Postgres client, and logs show 'HikariPool-1 - Connection is not available, request timed out after 30000ms'. The UI has a Memory ON/OFF toggle that sends use_memory to /plan. It's the best debugging tool in the project because it gives a like-for-like comparison: same model, same prompt, same alert, with and without recall. With memory off, the plan is competent but generic: check the pool, review recent deploys, consider restarting or scaling. That advice is reasonable in isolation, and it includes the two actions that have historically made this exact failure worse. With memory on, the plan changes shape. Step one is to roll back the recent config deploy, citing the earlier pool-exhaustion incidents. The avoid section lists restarting pods and scaling replicas, each with the reason from the incident where it backfired and a link to its source. If a cited incident is over six months old, a staleness note appears above the plan. The "Memories used" panel shows the recalled records, typed and dated. What convinced me was watching a citation land on an incident from a different service. The agent wasn't matching names. It was matching what actually went wrong. One more endpoint is worth mentioning. /patterns calls Hindsight's reflect operation with the query "What recurring root causes and failed fixes have you learned? Cite incident IDs." It's the nearest thing to an automatic review of the postmortem archive. I use it less during incidents and more between them, to see which failure classes keep coming back. Evaluating it without cheating Memory systems make self-deception easy. Test on incidents already in the store and you're measuring exact-match retrieval, not learning. The evaluation script replays incident history in date order against a fresh bank. For each incident it plans with memory off and on, grades both, and only afterward retains that incident: for n, inc in enumerate(incidents): row = {"id": inc["id"], "n_memories_before": n} for mode in ("off", "on"): p = agent.plan(inc["alert"], inc["symptoms"], use_memory=(mode == "on")) row[mode] = grade(p, inc) results.append(row) retain_incident(inc) # only AFTER planning, so it never sees itself time.sleep(5) # give consolidation a moment That comment carries the whole design. Retaining after planning means the agent never sees the answer to the question it's asked, and each incident is judged against only the history that would have existed then. The output is a learning curve: the cumulative wrong-first-suggestion rate against the number of past incidents in memory. A caveat, stated plainly: an LLM judge does the grading, checking whether the first step matches the known-correct fix and whether any known-bad action is recommended as something to do. That's good for trends and unreliable for exact numbers. I spot-check the judged rows by hand, and you should too if you reuse this approach. Lessons learned Store what failed, not only what worked. Failed attempts make up most of an incident's elapsed time, and a responder can skip them if they know about them. Put the outcome in the text, not just a field. Pass real timestamps when you retain. Backfilled memories with ingest-time dates all look current. With real dates, staleness handling becomes a prompt rule instead of a project. Give records stable IDs. Postmortems get edited. An idempotent document_id stops a corrected record from turning into two conflicting ones. Don't scope recall more narrowly than the cause. Root causes cross service boundaries, and a service filter would have hidden exactly the incidents that matter. Tags remain available for slicing later; I just don't gate recall on them. Build the off switch first. The memory ON/OFF toggle gave me a clean A/B on every alert, and the eval script uses the same flag. If you can't turn memory off, you can't tell what it contributes. Expect a delay after writing. Retained content is consolidated in the background, so a recall right after a write can come back empty or thin. The seed script prints a reminder to wait a minute or two, and the eval loop sleeps between steps. I learned this by staring at an empty result and blaming my code. To dig into the memory layer itself, start with the Hindsight GitHub repository, and see the Hindsight documentation for retain, recall, and reflect in more depth than I have room for here.



Top comments (0)