DEV Community

Mohammed Omer
Mohammed Omer

Posted on

Only Resolved Incidents Survive a Restart in My Hindsight Bank

I built an incident-diagnosis agent that remembers every outage we've resolved.

The interesting part isn't the diagnosis. It is the persistence boundary I ended up drawing, and how much of the system's value lives on one side of it.

The rule ended up being one line:

Incidents live in a process-local list and vanish on restart. Only resolved incidents get written into a Hindsight bank and outlive everything.

That constraint shaped every other decision in the codebase.

Demo

Here is a short demonstration of BlameLess in action:

What the system does

BlameLess has one workflow, and it runs in a loop:

  1. Recall: a new incident arrives as raw log paste. Semantic search over the Hindsight bank finds past incidents with similar symptoms.

  2. Reflect: the same Hindsight bank receives a natural-language question. Given these symptoms, what is the most likely root cause and fix based on past incidents? This becomes the diagnosis.

  3. Retain: when a human marks the incident as resolved, the postmortem is written back into Hindsight so the next similar incident can use it.

The division of responsibility matters:

The LLM does not diagnose.

Hindsight's reflect() performs the reasoning over the accumulated incident memory. The LLM call in llm.py exists only to reshape that result into UI sections.

If I take the memory away, the quality visibly changes. That is something I can demonstrate directly with a button.

Every Hindsight call lives in a single file, memory.py.

Nothing else in the project imports hindsight_client.

That gives me one place to log,
time, and error-handle the memory layer. It also means the Hindsight Python SDK can be changed without touching the rest of the application.

The persistence boundary

Here is the whole asymmetry in code:


python
# main.py - session state. Resets on restart, deliberately.
SESSION_INCIDENTS: list[dict[str, Any]] = []

# main.py - POST /incidents/{incident_id}/resolve
incident["root_cause"] = payload.root_cause
incident["resolution"] = payload.resolution
incident["status"] = "resolved"

retained = False
warning: str | None = None

try:
    await mem.retain_incident(incident)
    retained = True
except HindsightError as exc:
    # Resolving an incident must never lose the incident,
    # so we return 200 with a warning rather than erroring out.
    warning = str(exc)

That try/except returning HTTP 200 with a retained: false flag is one of the most important design decisions in the system.
An incident that has been diagnosed, root-caused, and fixed by a human is valuable information. If the Hindsight retain operation fails, the resolution itself should not be treated as a failed operation.
So the application records the resolution, reports the Hindsight failure honestly, and lets the operator decide what to do next.
The unresolved incident is different.
You can always re-paste logs.
You cannot easily re-create the reasoning that led to a root cause and a successful resolution.
That asymmetry is why the Hindsight bank is append-only from the application's perspective and why there is no delete endpoint in main.py.
What actually goes into Hindsight
Before memory can be useful, it has to be written consistently.
incident_to_memory_text() renders every postmortem into the same labelled layout:

parts = [
    f"INCIDENT {incident['id']}: {incident['title']}",
    f"Severity: {incident.get('severity', 'unknown')}",
    f"Affected system: {incident.get('system', 'unknown')}",
    "",
    "SYMPTOMS:",
    str(incident.get("symptoms", "")).strip(),
]

if incident.get("root_cause"):
    parts += [
        "",
        "ROOT CAUSE:",
        str(incident["root_cause"]).strip(),
    ]

if incident.get("resolution"):
    parts += [
        "",
        "RESOLUTION:",
        str(incident["resolution"]).strip(),
    ]

The structure is not cosmetic.
Hindsight chunks and embeds this text, so a predictable layout gives the memory system a consistent representation of every incident.
The retain call also passes:
document_id=incident["id"]
timestamp
metadata={"severity", "system"}
document_id makes a retained memory traceable back to the incident that produced it.
timestamp makes temporal questions possible, such as identifying what happened during a particular period.
This is where I started thinking about Hindsight as more than a vector store.
The memory needs to behave like an operational record because the questions being asked are operational questions.
Giving the Hindsight bank a mission
The identity of the memory bank is also explicitly defined.
create_bank() takes a mission and a disposition:

BANK_MISSION = (
    "I am an on-call incident analyst. I remember past incidents, their root "
    "causes and fixes, and use them to diagnose new ones. I focus on systemic "
    "causes, never on blaming people."
)

BANK_DISPOSITION = {
    "skepticism": 4,
    "literalism": 3,
    "empathy": 2,
}
The mission tells Hindsight what the accumulated memory is supposed to represent.
The disposition affects how the bank approaches retained information and future reasoning.
The important part is that the constraint is attached to the Hindsight bank itself rather than being treated as a temporary instruction in one application prompt.
That makes the memory layer responsible for maintaining the intended character of the accumulated incident knowledge.
Two things that bit me
1. The sync Hindsight client does not work in FastAPI
Hindsight's synchronous helpers use loop.run_until_complete().
Calling them from inside an async def route raises:

RuntimeError: event loop is already running
The fix is to use the asynchronous variants.
It also lets recall and reflect run concurrently because they are independent operations against the same Hindsight bank:
memories, reflect_text = await asyncio.gather(
    mem.recall(payload.symptoms),
    mem.reflect(payload.symptoms),
)
That concurrency is important because otherwise the two operations would unnecessarily wait for each other.
2. create_bank fails on the second run
A Hindsight bank that already exists returns 409.
Since the whole purpose of the bank is to outlive the application process, recreating it every time the application starts would be the wrong model.
I therefore treat startup as:
Ensure the bank exists.
rather than:
Create the bank.
The application can restart while the Hindsight memory remains intact.
That is one of the central properties of the system.
Recall returns facts, not documents
Another important detail is how Hindsight recall behaves.
arecall() on a 20-incident bank returns roughly 45 results.
Hindsight returns facts, so a single incident can appear as several fragments containing its symptoms, root cause, and resolution.
Rendering all of those fragments directly would produce an unreadable wall of text.
Instead, I collapse them back into one card per incident:
def group_memories(
    memories: list[dict[str, Any]],
    limit: int = 8
) -> list[dict[str, Any]]:
    grouped: dict[str, dict[str, Any]] = {}

    for m in memories:
        key = (
            m.get("document_id")
            or f"obs:{m.get('text', '')[:60]}"
        )

        ...

    ranked = sorted(
        grouped.values(),
        key=lambda c: (
            c["document_id"] is None,
            -c["score"],
        ),
    )
Enter fullscreen mode Exit fullscreen mode

Top comments (0)