Why Hindsight Needs the Reason a Fix Failed
The most useful thing an incident agent can remember is not simply that someone restarted a service. It is that the restart reduced errors for 25 minutes, the failures returned, and the real problem was still unresolved.
That idea became the foundation of an incident memory system I built using Hindsight as the long-term memory layer. Instead of storing only successful resolutions, the system also remembers failed troubleshooting attempts, engineer corrections, and the reasoning behind each decision. The goal is simple: help future investigators learn from previous incidents without confusing historical evidence with current facts.
From Incident Report to Long-Term Memory
An investigation usually begins with limited information:
- Incident title
- Service name
- Severity
- Error message
- Symptoms
This information is enough to start investigating, but it is not enough to create valuable long-term knowledge.
Only after an engineer identifies the root cause and completes the investigation does the incident become useful for future incidents. At that point the memory stores:
- Root cause
- Final resolution
- Failed attempts
- Why those attempts failed
- Lessons learned
- Engineer verification
- Corrections made later
Rather than letting the investigation flow interact directly with the storage backend, I created a small abstraction layer.
export interface MemoryBank {
readonly name: string;
readonly backend: "hindsight" | "local-fallback";
retain(record: MemoryRecord): Promise<void>;
recall(
incident: NewIncident,
limit?: number
): Promise<RecalledMemory[]>;
all(): Promise<MemoryRecord[]>;
clear(): Promise<void>;
}
This keeps the investigation workflow independent from the storage implementation. The application decides what should be remembered, while Hindsight decides how memories are retained and retrieved.
Why Failed Attempts Matter
One design decision turned out to be surprisingly important.
Instead of storing only the successful solution, every failed troubleshooting attempt is stored together with the reason it failed.
interface FailedAttempt {
action: string;
why_it_failed: string;
}
This small structure changes how future incidents are investigated.
Imagine a Payments API incident where the database connection pool becomes exhausted during peak traffic.
The engineering team tries several actions.
Attempt 1
Restart the application pods.
Result:
- Errors disappear.
- Twenty-five minutes later they return.
Attempt 2
Increase the database connection limit.
Result:
- Failures occur later.
- Database memory usage increases.
Eventually the engineers discover that a retry helper never releases database connections.
The permanent fix is to:
- release connections in a
finallyblock, - limit the pool size,
- monitor pool utilization.
If memory stores only the final resolution, future engineers lose valuable diagnostic information.
If memory stores only "Restarted pods", future engineers may waste valuable time repeating the same temporary workaround.
The useful knowledge is actually:
Restarting the service temporarily reduced errors because leaked connections were cleared, but the leak continued to exist.
That explanation becomes reusable engineering knowledge.
Preserving Trust in Memory
Not every memory should carry the same level of confidence.
Some incidents are imported as sample data.
Some are verified by engineers.
Others are corrected later when new information becomes available.
To preserve that history, each memory keeps additional metadata.
interface MemoryRecord extends Incident {
source: "seed" | "engineer_saved" | "engineer_correction";
verified_at: string | null;
corrections: string[];
}
This makes it possible to distinguish between:
- demonstration data,
- engineer-confirmed incidents,
- corrected historical records.
Instead of overwriting previous knowledge, corrections become part of the memory's history.
Using Historical Memory Carefully
When a new incident arrives, the application searches memory using information such as:
- service
- title
- symptoms
- error message
Hindsight retrieves the most relevant previous incidents.
Those memories are then presented as historical context—not as conclusions.
The investigation prompt follows a few simple rules.
- Keep current incident evidence separate from historical memories.
- Never assume a previous root cause is the current root cause.
- Clearly distinguish successful and failed previous actions.
- If multiple historical incidents disagree, request more current evidence.
This prevents historical knowledge from becoming false certainty.
For example, suppose a checkout service begins timing out again during heavy traffic.
The system recalls a previous Payments API incident involving connection leaks.
Instead of saying:
This is definitely another connection leak.
the investigation says something closer to:
A previous incident involving this service experienced a connection leak. Restarting the service only reduced errors temporarily. Check current connection usage, retry logic, and pool utilization before concluding the same root cause exists.
Historical experience guides the investigation without replacing current evidence.
Handling Conflicting Memories
Real systems evolve.
The same service might fail for completely different reasons six months apart.
One historical incident might point to a database leak.
Another might identify an upstream dependency failure.
Rather than combining these into a single confident answer, the application surfaces both memories and explicitly asks for additional evidence.
Conflicting memories are treated as useful information rather than mistakes.
This encourages investigation instead of assumption.
What I Learned
Building this project changed how I think about long-term memory for AI agents.
The most valuable memory is often not the successful fix.
It is understanding why previous fixes failed.
Negative results prevent engineers from repeating ineffective troubleshooting steps.
Engineer verification increases confidence.
Corrections preserve knowledge instead of hiding mistakes.
Most importantly, historical memory should narrow the investigation—not replace it.
Hindsight provides a powerful foundation for retaining and retrieving engineering experience, but the surrounding application still needs clear rules about what memories mean and how they should influence decision making.
For incident response, memory should answer one question:
"What should I investigate next?"
—not—





Top comments (0)