Restarting the checkout pods bought us eleven minutes. Scaling the database bought us fifteen. Neither one fixed anything, and both are the kind of thing a post-mortem mentions in one line, if at all, before moving on to "what finally worked."
That gap is what I built CausalOps around. It's an incident-response system whose memory holds the fixes that failed, the conditions they failed under, and the side effects they caused, right next to the ones that worked. The memory layer is Hindsight, an open-source agent memory system, and most of this article is about how I wired it in and where it surprised me.
What it does
When an incident opens, CausalOps builds a structured state: service, deployment version, error rate, p95 latency, DB connection saturation, plus the conditions and symptoms attached to it. Then it:
Recalls relevant history from Hindsight.
Scores every past intervention against the current incident.
Hands both to an LLM (gpt-4o-mini, temperature 0.2) that returns a structured analysis: root cause hypothesis, causal chain, four counterfactual branches (rollback, restart, scale, do nothing), and an explicit list of uncertainties.
Stops. requiresHumanApproval is always true, and the system prompt says "Never execute an operational action."
A human approves one branch. The outcome, the recovery time, and a learned rule get written back, both to Prisma/SQLite for the audit ledger and to Hindsight for future recall.
The stack is Next.js 14, Prisma, SQLite, React Flow for the causal graph, and the @vectorize-io/hindsight-client package. Hindsight runs as its own service on port 8888.
The through-line: failures are first-class memories
Most "similar incident" tooling indexes what fixed the problem. I think that's backwards. When a database is at 96% connection saturation, the dangerous move isn't the obscure one, it's the plausible one. Scaling the database is a reasonable-sounding response to a saturated database. It's wrong when the actual cause is a connection leak in the application, because the leak just eats the new capacity. That's what happened in INC-0762, and it's what CausalOps needs to remember.
So the retain path treats a failed intervention as a document with the same weight as a successful one. Here's the payload shape, trimmed:
const outcomeContent = [
`[Intervention Outcome: ${intv.actionName} on ${intv.incidentId}]`,
`Outcome Status: ${outcomeStatus}`, // SUCCESS | TEMPORARY_RECOVERY | FAILED
`Observed Result: ${intv.observedResult}`,
whyFailed ? `Why It Failed / Failure Reason: ${whyFailed}` : null,
intv.sideEffects ? `Side Effects Observed: ${intv.sideEffects}` : null,
`Operating Conditions Context: ${conditionsList.join(', ')}`,
`Runbook Execution Steps:\n${stepsFormatted}`,
].filter(Boolean).join('\n');
await retain(DEFAULT_MEMORY_BANK, outcomeContent, {
documentId: `intervention_${intv.id}`,
metadata: { type: 'intervention', incidentId: intv.incidentId, status: outcomeStatus },
tags: ['outcome', intv.incidentId, outcomeStatus.toLowerCase()],
});
Two decisions in there matter more than they look.
The failure reason is a labeled line. "Why It Failed" isn't buried in prose. When Hindsight extracts facts from the document, the causal claim ("scaling didn't help because the leak consumed the added capacity") survives as its own fact instead of dissolving into a summary.
The operating conditions travel with the outcome. A restart that works on a transient deadlock fails on a leak. If you retain "restart: worked" without the conditions, you've stored a lie that's true only sometimes. The memory has to carry the "under what conditions" or recall becomes confident and wrong.
I also give the bank a mission when I create it, so extraction is steered toward what I care about:
await client.createBank(DEFAULT_MEMORY_BANK, {
name: 'CausalOps Incident Memory',
mission: 'Store historical incident facts, causal graphs, runbook execution steps, ' +
'and intervention outcomes (both successful and failed) to guide safe remediation.',
});
Post-mortems go in through the same door. Pasting one into the memory page retains the raw text under a postmortem_* document ID and separately extracts a single rule with positive and negative conditions, for example: rollback is effective when a checkout deployment is followed by rapid DB connection growth, not when the saturation is database-only. That "not" is the whole point. It's the part a plain text search never gives you. Read what agent memory actually is if you want the longer argument for why retrieval over raw logs isn't the same thing.
Recall, and why I didn't let it decide alone
At analysis time, the query is built from the incident's own state, not from a free-text description:
const recallQuery = `Symptoms: ${currentSymptoms.join(', ')}. ` +
`Conditions: ${currentConditions.join(', ')}. ` +
`Service: ${incident.service.name}. ` +
`Suspected root cause: ${incident.rootCause || ''}`;
const recallResult = await recall(DEFAULT_MEMORY_BANK, recallQuery, {
budget: 'high',
preferObservations: true,
});
I use budget: 'high' because incident analysis isn't latency-sensitive at the level of milliseconds; a slower, deeper recall is the right trade. preferObservations biases toward Hindsight's consolidated observations rather than raw retained facts, which is what I want when the question is "what have we learned about this pattern?" The Hindsight documentation covers the recall budgets and memory types in detail.
Here's the design choice I'd defend in a code review: Hindsight doesn't produce the final ranking. Every past intervention also gets a deterministic score:
relevance = 0.35·conditions + 0.20·service + 0.15·deployment
+ 0.20·causal + 0.10·recency
Conditions are a Jaccard-style overlap with partial credit for substring matches. Recency decays over 90 days. Anything Hindsight recalls gets a boost on top:
if (actionMatch || incidentMatch || docMatch) {
exp.recalledViaHindsight = true;
exp.relevanceScore = Math.min(0.99, Number(((exp.relevanceScore || 0.7) + 0.18).toFixed(2)));
}
The two signals cover for each other. The formula is transparent and I can explain any number in it to an on-call engineer at 3 a.m. Hindsight catches what the formula can't: an incident on a different service, with different condition strings, that has the same causal shape. And when the UI shows a "Recalled via Hindsight" badge next to a card, the engineer can see which signal surfaced it.
What it looks like on a real incident
My regression case is INC-1042: checkout v4.7.2, 37% error rate, 4.8s p95, 96% DB connection saturation. With memory on, the analysis pulls three prior incidents into the frame:
INC-0981: rolled back a checkout deployment with the same connection growth. Worked, MTTR 4m 12s.
INC-0873: restarted pods. Temporary recovery, recurred after 11 minutes.
INC-0762: scaled the database. Failed; the leak consumed the added capacity.
The output recommends the rollback, and it explains the other branches by their history instead of just ranking them lower. Restart is flagged as "temporary relief followed by probable recurrence." Scaling is flagged as not addressing the cause. Doing nothing is marked HIGH risk and HARD to reverse. The language is deliberately hedged ("historical evidence suggests," "estimated") and the uncertainty list still says telemetry can't rule out concurrent batch-job load.
The before/after is one query parameter. POST /api/incidents/INC-1042/analyze?memory=off skips recall entirely, and the response reports hindsight.enabled: false. I built that toggle so I could see, incident by incident, what the memory layer changes rather than assuming it helps. Approving the rollback then writes the outcome (28-second recovery), a decision-ledger entry, and a new learned rule back into the bank, so the next incident starts with one more data point.
I want to be straight about what I haven't done: I haven't run a systematic evaluation of recommendation quality with and without memory. The toggle makes that experiment cheap, and I trust the behavior on the incidents I've walked through, but I'm not going to put a percentage on it.
Memory is never on the critical path
An incident tool that hangs because its memory service is down is worse than no memory at all. So the client wrapper does two boring things on purpose:
const controller = new AbortController();
const timeout = setTimeout(() => controller.abort(), 3500);
// ...
} catch (err: any) {
console.warn(`[Hindsight] recall() failed or unreachable (${err?.message || err}).`);
return null;
}
Every retain and recall has a 3.5-second abort and returns null on failure. Callers treat null as "no memory available" and continue with the structured score alone. There's also an isHindsightAvailable() health check the seeding script uses before it starts pushing history in. Memory advises. It doesn't gate.
What I'd tell someone else
Retain the reason, not just the result. "Failed" is nearly useless. "Failed because the leak consumed the added capacity" is a reusable fact. Structure your retained documents so the causal sentence is easy to extract.
Store negative conditions. The most valuable line in the whole system is the one that says when not to do the obviously good thing.
Keep a scorer you can explain next to the memory layer. I don't want the only reason a recommendation surfaced to be "the memory system returned it." Hybrid ranking gave me something to debug.
Build the off switch first. ?memory=off cost me ten lines and is the only reason I can say anything honest about what recall contributes.
Know where your joins are brittle. This is the part I like least. Matching Hindsight results back to my Prisma rows currently checks whether the recalled text contains the action name or incident ID, with a fallback on document IDs. Substring matching on prose is fragile, and the +0.18 boost is a hand-tuned constant, not a learned one. The fix is to make the documentId I already set on every retain the only join key, and to calibrate the boost against outcomes once there's enough history to do it. The service-similarity table in the scorer is also hard-coded to my own topology, and that won't survive contact with anyone else's.
I started this thinking the hard part would be getting an LLM to reason causally. It turned out the harder problem was making sure the system remembered being wrong. If you're building anything where an agent recommends actions with real consequences, give the failures a home in memory. The code is all built on Hindsight, and the docs are a good place to start if you want to try the same retain/recall pattern on your own runbooks.
Top comments (0)