Every team I've worked on has a folder of post-mortems that nobody opens during the next incident. The write-ups are fine. The problem is that at 2 a.m. you don't go looking for "that thing from last spring"; you start from scratch. I wanted to see how small I could make the loop that puts an old post-mortem in front of you when a new incident looks like it.
The result is a small Python project, Incident-Response-Agent. It stores past incidents in Hindsight, an open-source agent memory system, and looks up the closest one when a new incident description comes in.
One note on scope: the incident records in this repo are three sample records I wrote to exercise the pipeline. They are not real production incidents, and nothing here is a benchmark. This is a write-up of how the memory layer is wired and what I learned from wiring it.
What the project does
“The whole thing is a few small Python files plus a data file:”
-
data/incidents.jsonholds the incident records. -
src/memory.pywraps Hindsight with two functions: one to store an incident, one to search for similar ones. -
src/agent.pyloads the records, stores them, takes a new incident description, and prints the best match along with its runbook and lesson.
There is no web UI, no ticketing integration, and no alerting hook. The agent is a script that runs build_memory() and then investigate() on a query string. I kept it that small on purpose, because I wanted to see what the memory layer does on its own before building anything around it.
How the incidents are stored
Each record in data/incidents.json has an id, a title, a root cause, a resolution, a runbook name, and a lesson:
{
"id": "INC-003",
"title": "Notification service failure",
"root_cause": "Message queue stopped processing events",
"resolution": "Restarted the queue consumer and replayed failed messages",
"runbook": "notification-queue-recovery",
"lesson": "Check queue consumer health before investigating application code"
}
There are three of these: a payment API outage caused by connection pool exhaustion, a checkout timeout caused by slow queries, and the notification queue failure above. They deliberately overlap in vocabulary (two of them involve databases, all three involve restarts), so retrieval has to do more than match a single keyword.
Before storing anything, build_memory() flattens each record into one sentence-shaped string:
text = (
f"Incident: {incident['title']}. "
f"Root cause: {incident['root_cause']}. "
f"Resolution: {incident['resolution']}. "
f"Runbook: {incident['runbook']}. "
f"Lesson: {incident['lesson']}."
)
remember_incident(text)
I chose this over storing raw JSON because the memory system is searching over text. A flat, labelled paragraph keeps the root cause, the fix, and the runbook in the same chunk, so whatever comes back from a search carries all of them. The trade-off is that the structure is now encoded in labels like Runbook: and I have to parse it back out later. More on that below.
Retain and Recall

The memory layer is src/memory.py, and it's short enough to show whole:
from hindsight import HindsightEmbedded
memory = HindsightEmbedded(
profile="incident-response",
llm_provider="none",
llm_api_key=""
)
BANK_ID = "incidents"
def remember_incident(text):
memory.retain(
bank_id=BANK_ID,
content=text,
context="production incident post-mortem",
document_id=text.split(".")[0]
)
def find_similar_incident(query):
return memory.recall(
bank_id=BANK_ID,
query=query,
min_scores={"reranker": 0.5}
)
A few things here are worth explaining, because I made each choice deliberately.
Embedded, with no LLM configured. I'm using HindsightEmbedded, which runs the memory system inside my own process rather than as a separate service, and I set llm_provider="none". The project only ever calls retain and recall. It never asks a model to generate anything, so I didn't wire one in. If you want to see what else Hindsight can do, the Hindsight documentation covers the wider API. I haven't used those parts here and I'm not going to pretend I have.
One bank for all incidents. Everything goes into a memory bank named incidents, and context="production incident post-mortem" tags what kind of content it is. If I later add other kinds of memory (on-call notes, customer-facing comms), separate banks keep them from polluting each other's searches.
A stable document_id. I derive it from the first sentence of the stored text, which is Incident: <title>. The point is that each incident maps to a predictable identifier rather than a random one. It is a shortcut, though. Two incidents with the same title would collide, and the id field from the JSON (INC-001 and so on) isn't included in the stored text at all. Using it would be the obvious fix.
A score floor on recall. min_scores={"reranker": 0.5} tells Hindsight to drop results whose reranker score is below 0.5. I picked 0.5 as a starting point, not from tuning. It matters because of what I wanted the agent to do when nothing matches: say so, rather than surface the least-bad record and present it as a suggestion.
From match to recommendation
investigate() takes the new incident description and calls find_similar_incident(). If results come back, it takes the first one and pulls the runbook and lesson out of the text by splitting on the labels:
best = results[0].text
if "Runbook:" in best:
runbook = best.split("Runbook:")[1].split(".")[0]
print("Recommended Runbook:", runbook)
if "Lesson:" in best:
lesson = best.split("Lesson:")[1].split(".")[0]
print("Lesson:", lesson)
If the result list is empty, it prints that no similar incident was found.
The query the script ships with is "The notification service is failing because messages are stuck in the queue." It is written to line up with the third sample record, so the path being exercised is: recall returns the notification-service incident, and the agent prints the notification-queue-recovery runbook and the lesson about checking queue consumer health before digging into application code. I'm describing what the code is built to do with that input; I haven't included captured output or any hit-rate numbers, because three records can't support them.
This parsing is the most fragile part of the project. It works because I control the format of the text going in, and it breaks if a runbook name or lesson ever contains a period. Splitting on "." is a stand-in for storing runbook and lesson as real structured fields.
What I learned
Decide what a memory contains before you decide how to search it. Most of the design work was in that flattened string. Putting cause, fix, runbook, and lesson into one labelled chunk meant a single recall call returned everything the agent needed.
You don't need a model in the loop to get value from memory. Retrieval alone is enough for "have we seen this before?" Leaving the LLM out made the project easier to reason about, because every behavior comes from what I stored and what I asked for.
Set a relevance floor, and decide what "no match" means up front. An agent that always answers is worse than one that admits it has nothing. The reranker threshold is a blunt instrument, but it makes the empty-result branch real.
Don't round-trip structure through prose if you can avoid it. Encoding fields as labels and parsing them back with split works for a demo-sized dataset and nothing more. I'd store runbook and lesson as structured fields next to the text.
Be honest about your test data. Three sample records can show that the plumbing works. They can't show that retrieval quality holds up on a real archive, and I'd want a much larger, messier set before claiming anything about that.
If you want to try the same approach, start with the Hindsight repository. For the broader idea of why agents benefit from persistent memory, Vectorize's explainer on what agent memory is is a good place to begin.


Top comments (0)