I used Code.in as my AI coding agent for this project.
When a production alert fires, the first useful question is usually "have we seen this before?" Answering it means reconstructing history from someone's memory, old tickets, chat threads and runbooks. That takes time, and the result depends on who is on call.
I built an incident-response agent that answers that question from persistent memory, using Hindsight as the memory layer. The most useful thing I learned came from a bug: at one point the agent cited the right past incident and still got its root cause wrong.
A scope note before the details. This is a demonstration built on six synthetic incidents that I wrote myself. None of it comes from a real production system, and the agent doesn't learn from conversations. Incidents enter its memory only when I run an ingest script.
How the pieces fit
There are two jobs:
-
Retain:
ingest.pytakes each synthetic incident and stores it in a Hindsight memory bank. -
Recall and diagnose:
agent.pytakes a new incident description, asks Hindsight for related memories, and passes both to an LLM. The model isopenai/gpt-oss-120bon Groq, called through the OpenAI-compatible client againsthttps://api.groq.com/openai/v1.
The repo has no vector database, embedding code or retrieval logic. Hindsight stores and retrieves, and my code calls its client.
Each incident is formatted as natural-language text with its ID, severity, service, title, symptoms, root cause, resolution, runbook and timestamp. The retain step is one call per incident:
content = format_incident_for_memory(incident)
client.retain(bank_id=BANK_ID, content=content)
Recall and the baseline
diagnose() takes a use_memory flag. With use_memory=False, Hindsight is skipped and the model receives "(no memories retrieved)". That gives me a fair baseline: same prompt, same model, and only the memory differs. With use_memory=True, the agent calls recall:
recall_result = client.recall(
bank_id=BANK_ID,
query=new_incident_description,
max_tokens=2000,
)
seen_texts = set()
for m in getattr(recall_result, "results", []):
text = m.text.strip()
if not text or text in seen_texts:
continue
seen_texts.add(text)
memories.append(text)
if len(memories) >= MAX_MEMORIES:
break
MAX_MEMORIES is 8. The system prompt does the rest. It tells the model that memory entries are fragments to combine, and that a root cause or resolution may be attributed to a past incident only if it appears in the recalled text:
- Only state a root cause or resolution "from" a past incident if it
appears in the MEMORY text. If the memory shows the symptoms but not the
root cause or fix, say that plainly instead of guessing what it was.
Before and after
The demo incident is:
Checkout API is returning intermittent 502s again, started about 10 minutes ago. We're seeing it spike right after that promo email went out this morning. p99 latency on /checkout/create is climbing fast.
Without memory, the agent says it has no similar past incident and produces a generic diagnostic plan based only on the symptoms. It lists possible causes (traffic surge, downstream overload, connection-pool exhaustion, circuit-breaker or timeout settings, recent configuration changes, network issues) and then diagnostic steps. Connection-pool exhaustion is one candidate among several, and nothing in the plan is specific to my system.
With memory, Hindsight returned eight entries. The important ones described INC-1042:
- intermittent 502 Bad Gateway errors on
checkout-apiafter a marketing email blast, with p99 on/checkout/createrising from about 300 ms to several seconds - a documented root cause: a PgBouncer connection-pool limit of 20 connections for the
checkout-dbPostgreSQL database, so requests queued and exceeded the 5-second upstream gateway timeout - a documented resolution: raise
max_client_connfrom 20 to 100, enable transaction pooling, and add a CloudWatch alarm on connection-wait time
The diagnosis called the new alert a likely repeat of INC-1042 and tied the two together: same service, same endpoint, intermittent 502s, a marketing email traffic spike, and rapidly rising p99 latency. It then used the documented PgBouncer cause and fix instead of a generic guess. Its first checklist step was to verify the current PgBouncer configuration rather than assume it.
Both runs use the same model, so the difference is the context Hindsight supplied.
The bug: the right incident with the wrong facts
My first recall logic deduplicated by incident ID. Hindsight returns memory as fragments, and I assumed several fragments from one incident would just be noise. So I kept the first fragment per incident and capped the list at three.
In practice that was a mistake. The fragment that survived for INC-1042 was the symptoms. The fragments holding the root cause and the resolution were discarded by my own code. The model received symptoms without a cause, and it filled the gap. It wrote a "known root cause" for INC-1042 that described a generic traffic surge, and a "resolution" made of ordinary scaling advice. The documented PgBouncer limit appeared nowhere.
This is worse than having no memory. The answer cited a real incident ID, so it looked verified, but the facts attached to that ID were the model's guesses.
I made two changes. Deduplication now removes only exact duplicate text, so sibling fragments survive, and the limit went up to eight. I also added the system-prompt rule quoted above, so the model can only attribute what's in the recalled text. On the next run, the root-cause and resolution fragments were in the memory panel and the diagnosis used them.
Limits and what production would need
- The data is synthetic and small. Six incidents I authored, with a new alert I wrote to resemble one of them. This is a best case. I haven't tested alerts with no good match, and I haven't measured recall quality, accuracy or latency.
-
The agent doesn't learn on its own.
diagnose()never writes to memory. Incidents are added only throughingest.py. - Recall still returns near-duplicates. Some entries described the same fact in different words, and exact-text dedupe doesn't catch those. That spends context on repetition.
- Production needs more. Real incident data raises questions about access control, retention and sensitive content in logs and tickets. It would need validation of recalled facts against the source records, and integration with real incident systems such as ticketing, paging and postmortem tools. None of that exists here.
What I'd take from this
- Recall returns fragments, so don't deduplicate by entity. Collapsing to one fragment per incident kept the symptoms and lost the cause.
- A cited ID is not a verified fact. Tell the model to attribute only what appears in the retrieved text, and make it admit when the cause is missing.
- Print the recalled memory next to the answer. I found the bug because the demo shows both panels.
-
Keep the no-memory baseline in the same code path. A
use_memoryflag makes the comparison fair. - Treat memory as history, not current state. Recalled facts describe what was true when the incident happened, so the agent should check the current configuration.



Top comments (0)