
# I Built Hindsight's Safety Gate Around Uncertainty
The dangerous moment in incident response is not when an agent has no answer. It is when it has an answer that looks just plausible enough to execute.
I built RECALL around that problem: turning historical incidents into usable operational memory without treating every retrieved memory as an instruction. The memory layer combines incident history, semantic lessons, runbooks, service topology, incident fingerprints, freshness, and contradiction handling. I use Hindsight GitHub as the long-lived memory layer in the production architecture, while RECALL owns the operational policy that decides when remembered information is trustworthy enough to influence an incident response.
That distinction ended up being more important than the retrieval itself.
The system is really a memory pipeline
RECALL starts with a live incident: a service, symptoms, error signatures, deployment context, and severity. The first job is retrieval. The second is interpretation. The third is deciding how much authority the retrieved history deserves.
The repository makes those stages explicit.
The MemoryEngine keeps three useful forms of operational memory:
- Episodic memory: individual incidents and their resolutions.
- Semantic memory: lessons distilled from several incidents.
- Procedural memory: runbooks with execution history and win rates.
That maps naturally onto how I use Hindsight's agent memory: past events are useful because they preserve context, not because they produce a single magic similarity number.
The incident itself is represented by an eight-dimensional Incident DNA: symptoms, error severity, service depth, deploy proximity, blast radius, configuration drift, network pressure, and resolution complexity.
Retrieval then combines several signals instead of trusting one search method. The code calculates BM25 relevance, exact or adjacent service relationships, symptom overlap, deployment and error-signature alignment, and the historical runbook's success rate.
The result is deliberately explainable.
const overallScore = (
bm25Normalized * 0.35 +
serviceMatch * 0.25 +
tagOverlap * 0.20 +
signalAlignment * 0.10 +
runbookConfidence * 0.05
);
I like this approach because an engineer can challenge the score. If a match is high, I can point to the service, symptoms, deployment proximity, and the runbook history that produced it. If it is wrong, there are identifiable inputs to inspect.
But there is a problem hidden inside that score.
A historically similar incident can still be operationally wrong.
The memory can be right while the recommendation is wrong
One of the most useful cases in the repository is Redis.
An older incident describes a Redis eviction storm and contains a successful mitigation. At retrieval time, it can look extremely relevant to a new Redis incident. Same service. Similar symptoms. Similar error signatures. Possibly the same deployment pattern.
The problem is that the infrastructure changed.
The old environment used self-hosted Redis on AKS virtual machines. The newer architecture uses Azure Cache for Redis Enterprise. A historical instruction such as restarting redis-server can therefore be perfectly faithful to the old incident and completely inappropriate for the current system.
This is where I put the safety boundary around Hindsight.
Hindsight is useful for remembering what happened. RECALL decides whether that memory still applies to the infrastructure in front of me.
The repository models known architectural migrations explicitly:
{
targetService: 'redis-cluster',
migrationDate: '2025-01-15',
oldArch: 'Self-hosted Redis 6.2 on AKS VMs',
newArch: 'Azure Cache for Redis Enterprise (Active-Active Geo)',
impactedFixKeyword: 'restart redis-server',
confidenceDeduction: 29,
}
The important design choice is that staleness does not simply delete the memory.
I still want the incident because the failure pattern may be valuable. What I do not want is to silently turn an old remediation into a current command.
So the system separates similarity from actionability.
if (doc.stalenessContext?.isStale) {
const deduction = doc.stalenessContext.confidenceDeduction || 29;
stalenessWarning = `Fix confidence adjusted: ${Math.round(overallScore)}% → ` +
`${Math.max(15, Math.round(overallScore) - deduction)}%.`;
}
That separation became the core of my Hindsight integration.
I wanted uncertainty to be a state, not an error
The easiest mistake to make with an incident agent is to assume that every incident has a historical precedent.
Real systems eventually produce something you have never seen before.
RECALL therefore has a hard retrieval boundary. When the best historical match falls below 45%, the system enters Novel Incident Mode instead of manufacturing confidence from weak evidence.
The UI makes that state explicit: no authoritative historical match, exploratory diagnostics instead of autonomous mitigation.
That sounds simple, but it changes the architecture.
A conventional retrieval system asks:
“What is the closest thing I remember?”
My safety gate asks two questions:
“How similar is it?”
“How safe is it to act on that similarity?”
Those are different questions.
Hindsight is particularly useful here because its value is not limited to returning the nearest text fragment. A durable memory system can preserve the relationships around an experience: what happened, what was learned, and what should be remembered later. RECALL then adds operational constraints around that memory before it becomes an execution recommendation. The Hindsight documentation describes this broader memory model, which fits the way I wanted incident history to behave.
The second safety boundary is execution
Even a fresh, high-confidence memory is not permission to change production.
That is why retrieval and execution are separate stages in RECALL.
Each procedural memory contains not just a command, but a dry-run command, rollback command, risk level, blast radius, estimated execution time, and historical success rate.
export interface RunbookStep {
id: string;
stepNumber: number;
title: string;
command: string;
dryRunCommand: string;
rollbackCommand: string;
risk: RiskLevel;
blastRadius: string;
estimatedSeconds: number;
}
The operator gets a dry-run path before production execution. Destructive operations require explicit human approval. A rollback path is part of the procedure rather than something the agent invents after a failure.
This is important because uncertainty exists at more than one layer.
The retrieval layer can be uncertain about whether an old incident applies. The execution layer can be uncertain about whether the proposed command is safe in the current environment. The system should not collapse both uncertainties into one confidence percentage.
Contradictions are memory, too
Another design problem appeared once I treated operational knowledge as something that changes over time: sometimes two historical memories disagree.
The repository includes a contradiction registry for exactly this case. One incident recommends increasing Kafka's max.poll.interval.ms. A later incident explains why that change can mask deadlocked consumer threads and recommends reducing max.poll.records instead.
I don't want retrieval to pick whichever document happens to rank higher and call the problem solved.
Instead, RECALL exposes the conflict and records a reconciliation decision.
public reconcileContradiction(id: string, resolutionNote: string) {
const c = this.contradictions.find(item => item.id === id);
if (c) {
c.status = 'reconciled';
c.suggestedReconciliation = resolutionNote;
}
}
This is another place where Hindsight changes the way I think about memory. A useful memory system should preserve not only successful answers, but also the fact that the organization once disagreed about an answer and eventually learned why.
That is operational knowledge that would otherwise disappear into old post-mortems.
What happens after the incident is resolved
The memory loop is incomplete if incidents only get retrieved. The system has to learn from its own outcomes.
When an incident resolves, RECALL's consolidation stage creates a new episodic memory, updates the selected runbook's execution statistics, creates a semantic lesson, and records a post-mortem summary.
The runbook statistics matter because procedural memory should have evidence behind it. If a runbook repeatedly succeeds, that history becomes part of future retrieval. If it starts failing, its influence should change.
The repository models this directly:
rb.timesExecuted += 1;
rb.timesSucceeded += 1;
rb.winRate = Math.round(
(rb.timesSucceeded / rb.timesExecuted) * 100
);
The important part is the loop:
incident → retrieval → guarded recommendation → execution → outcome → memory update
That is much more useful than simply attaching a chatbot to a pile of runbooks.
A concrete incident flow
Consider the Redis eviction scenario.
A production alert arrives for cart-service. The symptoms include cache pressure and eviction errors shortly after a deployment. RECALL searches historical incidents using lexical relevance, service relationships, symptoms, error signatures, and deployment context.
The old Redis incident rises to the top.
The system can explain why: exact service relationship, overlapping symptoms, deployment proximity, and a previously successful runbook.
Then the staleness guard notices that the historical remediation predates the Redis migration.
The old incident remains valuable as evidence about the failure pattern, but its operational confidence is reduced. The current recommendation uses the cloud-native recovery procedure rather than blindly replaying the old VM restart.
The operator can dry-run the proposed step, inspect the blast radius, and approve production execution.
After resolution, the new outcome becomes another memory. The runbook's history changes, the semantic lesson is reinforced, and the next incident has one more piece of evidence available.
That is the behavior I wanted from Hindsight: memory that gets more useful with experience without becoming more authoritative merely because it is older and well indexed.
What I learned building the safety boundary
1. Retrieval confidence is not execution confidence
A 94% historical match does not mean a 94% safe command. Similarity answers whether two situations resemble each other. Safety requires additional evidence about infrastructure, provenance, permissions, and blast radius.
2. Stale memory is still useful memory
Deleting old incidents because they are outdated throws away the failure pattern. I would rather preserve the incident and downgrade its authority when the architecture has changed.
3. Contradictions deserve first-class storage
If two experienced engineers reached different conclusions, that disagreement is part of the organization's knowledge. Hiding it behind ranking makes the system look cleaner while making the reasoning worse.
4. Human approval belongs after reasoning, not before it
A human should not have to manually reconstruct the agent's reasoning from logs. The system should surface the evidence, confidence, staleness warnings, proposed command, dry-run result, and rollback path before asking for approval.
5. Memory needs an outcome loop
The best incident memory is not the largest database. It is the memory whose usefulness changes when reality proves or disproves it. Every executed runbook and resolved incident should feed that loop.
The part I would protect most carefully
If I had to preserve only one architectural decision from RECALL, it would be the separation between remembering and acting.
Hindsight gives me a strong foundation for persistent agent memory. RECALL adds the operational context that memory alone cannot safely infer: architectural drift, contradictions, service topology, runbook history, execution risk, and explicit human approval.
That boundary is what keeps a useful memory system from becoming an automation system that confidently repeats yesterday's assumptions.
The goal was never to make the agent remember everything.
It was to make it remember enough to help, and know when remembering is not enough.
Top comments (0)