Why Hindsight Needs the Reason a Fix Failed
The most useful thing an incident agent can remember is not that someone restarted a service. It is that the restart bought 25 minutes, the errors returned, and the connection leak was still there.
That distinction is the center of the incident memory system I built. It collects what engineers learned during an outage, retrieves relevant experience when another incident arrives, and gives an investigator that context without letting history masquerade as current evidence. I use Hindsight for the long-term memory layer because incident knowledge is more than a pile of similar-looking error messages.
From incident form to durable memory
An incident starts with the information an on-call engineer actually has: a title, service, severity, symptoms, and perhaps an error string. That is enough to begin an investigation, but not enough to make a useful memory. The record becomes valuable after someone adds a root cause, the steps that exposed it, the resolution, and the attempts that did not work.
The project keeps that record behind a small interface. The investigation flow does not need to know how memory is stored; it asks the bank to retain an incident or recall a few records for a new one:
export interface MemoryBank {
readonly name: string;
readonly backend: "hindsight" | "local-fallback";
retain(record: MemoryRecord): Promise<void>;
recall(incident: NewIncident, limit?: number): Promise<RecalledMemory[]>;
all(): Promise<MemoryRecord[]>;
clear(): Promise<void>;
}
In the production system, the Hindsight-backed adapter maps those operations to a dedicated incident memory bank. Keeping that boundary explicit matters: the incident workflow owns the meaning of a record, while Hindsight owns long-term retention and retrieval. The Hindsight GitHub repository and Hindsight documentation describe the retain, recall, and reflect operations that make this more than a transcript archive.
The record is deliberately structured. A failed attempt is an action paired with its explanation, not a note buried in a postmortem paragraph:
interface FailedAttempt {
action: string;
why_it_failed: string;
}
interface MemoryRecord extends Incident {
source: "seed" | "engineer_saved" | "engineer_correction";
verified_at: string | null;
corrections: string[];
}
This structure captures provenance as well as content. A seed record, an engineer-confirmed resolution, and a later correction should not carry identical authority. When an engineer saves an incident, the workflow records the root cause, resolution, failed action, and lesson, and marks the record verified. When someone corrects a recalled incident, the correction remains attached to that incident rather than becoming an unrelated new document.
Why the failure belongs beside the fix
The payment incident in the repository makes the design choice concrete. Under peak load, the Payments API exhausted its database connection pool. Restarting the pods cleared the pools and reduced errors for about 25 minutes. Then the leak filled them again. Increasing the database connection limit bought more time, but also increased memory pressure. Neither action fixed the code path that kept connections open inside a retry loop.
The eventual repair was to release connections in a finally block, cap pool size per pod, and alert on pool utilization. If memory retained only “restarted pods” and “raised max_connections,” a later investigation could repeat both interventions while missing the actual lesson. If it retained only the final resolution, it would omit the strongest diagnostic clue: a restart that briefly helps can point toward a resource leak.
So I send failed attempts and their explanations along with the successful resolution. This is where Hindsight’s model of agent memory is useful. Retain stores an experience; recall finds relevant experiences later; reflect can reason across those experiences to form a broader view. The Vectorize overview of agent memory draws a useful line between remembering text and learning from accumulated experience. In incident response, that line is the difference between finding an old ticket and understanding what the old ticket should change about today’s investigation.
Recall is context, not a verdict
When a new incident comes in, the application queries memory using the title, service, symptoms, and error message. The production Hindsight integration can use more than literal string overlap: its retrieval combines semantic, keyword, entity, and temporal signals. That matters when the new incident says “pool checkout timed out” and the old postmortem says “connection slots exhausted.” The wording differs; the operational mechanism may not.
The application passes the recalled records into the investigation as a separate section. It labels the current event and historical memories explicitly, then gives the model a strict rule:
- Keep CURRENT INCIDENT and PAST INCIDENT MEMORIES strictly separate.
- Never state that a past incident IS the current root cause.
- Clearly distinguish SUCCESSFUL PREVIOUS ACTIONS from FAILED PREVIOUS ACTIONS.
- If conflicting memories are provided, say additional current evidence is required.
That instruction is not decorative prompt hygiene. Similar symptoms do not prove the same cause. A service can time out because of a connection leak, a slow dependency, a network change, or a bad rollout. The old incident should help the engineer choose what to inspect, not allow the agent to skip inspection.
For each recalled memory, the interface also shows why it was returned: matching terms, service, age, engineer verification, and any corrections. Confidence is a property of the evidence we have, not a number the language model gets to invent. Older or unverified memories can still be useful, but the engineer should be able to see their limits.
What a repeat incident looks like
Suppose checkout errors return with a database timeout during peak traffic. Hindsight recalls the earlier Payments API incident because it shares the service and connection symptoms, even if the error message has changed. The investigation can then say: a previous incident found connections leaking in a retry helper; restarting pods only hid the issue temporarily; increasing the connection ceiling delayed failure and raised database memory pressure.
That is a better starting point than “try restarting the pods.” The agent can recommend checking active and idle database connections, comparing pool checkout latency with the deploy timeline, and tracing the retry path. It should still ask for current evidence before calling the root cause. If the current connection counts are normal and a dependency is timing out, the old incident is a useful contrast, not the answer.
The workflow also handles disagreement deliberately. If two recalled records for the same service point to different causes, the system surfaces both and asks for more evidence instead of blending them into a confident-sounding compromise. Hindsight can help retrieve and synthesize the history; the application still has to define what uncertainty means during an incident.
What I learned
Store the negative result, not just the action. “Restarted pods” is an event. “Errors returned after 25 minutes because connections continued leaking” is experience an investigator can use.
Preserve provenance. Engineer verification and corrections change how much weight a memory deserves. Keep that history visible rather than flattening every record into the same text field.
Use memory to choose tests, not to declare causes. Historical evidence is valuable because it narrows what to inspect. Production evidence still decides what happened this time.
Treat disagreement as data. Conflicting postmortems can reflect different failure modes, changed architecture, or weak records. Hiding the conflict makes the answer cleaner and the investigation worse.
Make retention part of incident closure. If saving the resolution is optional, the memory bank will mostly contain incidents that were easy to document. Capture the failed actions and the final explanation while the people who know the details are still there.
Incident response already produces the raw material for a useful memory system: hypotheses, failed mitigations, confirmed causes, and corrections. Hindsight gives that history a long-lived place to accumulate and a way to return when it matters. The engineering work is deciding what the agent is allowed to conclude from it. My rule is simple: remember exactly why the last fix failed, then make the next investigator prove whether the same failure is happening again.






Top comments (0)