
The plan looked great, and I couldn't tell whether it was using any of our incident history.
That was the real problem when I built Incident Copilot. Getting an LLM to write a fluent incident response plan is easy. Knowing whether that plan came from our past outages or from the model's general training is hard. A confident answer proves nothing, so I added a switch.
What the system does
Incident Copilot is a small FastAPI service with a single-page UI. You paste an alert and symptoms, and it returns a ranked plan. Beside the plan, a "Memories used" panel shows the exact records the plan was built from.
The flow is one recall followed by one LLM call. There is no planner loop and no vector store I maintain. The model is openai/gpt-oss-120b on Groq, and the memory layer is Hindsight, the open-source agent memory system. Four endpoints do the work:
• POST /plan recalls memories and returns a structured plan.
• POST /action records a step's outcome (worked, failed, harmful) mid-incident.
• POST /close retains the full postmortem.
• GET /patterns asks Hindsight to reflect on recurring root causes.
If you're new to the idea, Vectorize has a good overview of what agent memory is, and the Hindsight docs cover retain, recall, and reflect, the three operations I use.
Why I needed an off switch
Without memory, an LLM gives you the textbook plan: check the pool, look at recent deploys, restart the pods, scale up. Those are sensible in isolation, which is the trap. If recall silently breaks (empty results, wrong bank, a bad query), the plan still reads well and nothing tells you.
So /plan takes a use_memory flag. Same model, same prompt, same alert, and the only variable is whether recall runs. The same flag drives my evaluation script:
python
for mode in ("off", "on"):
p = agent.plan(inc["alert"], inc["symptoms"], use_memory=(mode == "on"))
row[mode] = grade(p, inc)
Once that flag existed, every question became an A/B test instead of a hunch. The toggle turned "does memory help?" into something I could answer per alert in about ten seconds.
What I put in memory: failures with outcomes
A toggle only shows a difference if there is something worth recalling. My seed data is in data/incidents.json, and each incident records every action taken and how it turned out:
json
{
"id": "INC-101",
"service": "checkout-service",
"alert": "checkout-service p99 latency > 4s for 5m",
"actions": [
{ "action": "Restarted checkout pods", "outcome": "harmful",
"note": "Latency got worse; the reconnect storm exhausted the pool..." },
{ "action": "Scaled checkout to 12 replicas", "outcome": "harmful",
"note": "More replicas opened more DB connections and hit max_connections..." }
]
}
Restarting and scaling are the two things everyone reaches for at 3 a.m., and in this failure family both make it worse. At retain time I render each outcome into the text as an uppercase word (Outcome: HARMFUL), because Hindsight extracts memories from natural language and I wanted "restarting pods was harmful" to survive as a fact, not as a dropped metadata field.
Reading the toggle: the same alert, two ways
I ran an alert that isn't in the seed data: "Kafka consumer lag growing on the orders topic," with the symptom "Consumer group keeps rebalancing every few minutes."
Memory ON: the plan's first step was to roll back config deploy #5102, which lowered the orders-api connection pool size and idle timeout. It cited INC-104, where the same kind of settings change caused instability. A yellow banner warned that several referenced incidents were older than six months. In the Memories used panel, one recalled record reads: "Restarting orders-api pods worsened the incident by triggering a reconnect storm that exhausted the connection pool."
The panel is what made this a debugging tool and not just a demo. Each recalled memory is tagged world, experience, or observation, and dated. When a plan looked wrong, I could see in one glance whether retrieval had fetched the wrong record or the model had misused the right one.
The loop that makes it learn
Every step in a plan has Worked, Failed, and Harmful buttons. Clicking one calls /action, which retains a short note right away, keyed to a stable document_id so edits replace records instead of duplicating them:
python
_retain(
bank_id=BANK,
content=f"During incident {incident_id}, attempted: {action}. Outcome: {outcome.upper()}. {note}",
context="live incident action",
document_id=f"incident-{incident_id}-action-{n}",
)
I marked a pool-sizing step as worked for incident LIVE-001 and the UI showed "Worked (saved)." A later alert on payments-api (HikariPool-1 - Connection is not available, request timed out after 30000ms) got a plan that cited LIVE-001 directly, saying a recent deploy had altered pool settings and caused the same timeout. An incident I had just closed was already shaping the next recommendation. The plan also cited INC-108 for the rollback step.
The "Show patterns" button calls Hindsight's reflect operation and returns a "Recurring Root Causes and Failed Fixes" summary across the whole bank, starting with families like expired mTLS certificates. I use it between incidents more than during them.
A choice that mattered: recall across services
I don't filter recall by service:
python
resp = client.recall(bank_id=BANK, query=f"{alert} {symptoms}".strip(),
budget="high", max_tokens=3000)
A config change that shrinks a connection pool looks the same on checkout-service, orders-api, and payments-api. In my data the same root cause appears across several services, and a per-service index would have found nothing for a new one. The Kafka run above worked because it recalled an orders-api incident for a Kafka consumer symptom, which shows the flip side too: recall keyed on symptoms can pull a plan toward a root cause that doesn't fit. Nothing in the Kafka alert mentioned a config change, yet the plan led with one.
Lessons learned
- Build the off switch first. It is cheaper than any evaluation framework, and without it you can't tell what memory contributes.
- Put the outcome in the text. "Harmful" as a searchable word in the memory, not just a JSON field, is what let recall surface warnings.
- Show the recalled memories next to the answer. It separates retrieval bugs from generation bugs faster than anything else I tried.
- Real timestamps matter. I retain each incident with its resolution date, so the agent can reason about staleness. The six-month banner in the Kafka run comes from that.
- Expect over-anchoring. Cross-service recall is powerful and occasionally too confident.
- Give memory time to settle. Retained content is consolidated in the background, so recall right after seeding can look thin. I lost time to that once before I read the seed script's reminder. Where this goes next The code is at https://github.com/Hasinireddy2407/incident-copilot. If you want to try the memory layer yourself, start with the Hindsight GitHub repository.
Top comments (0)