The first time I measured my on-call agent with memory against the same model without it, memory lost. Replayed
over six months of incidents, the agent that could remember every postmortem found the root cause 50% of the
time. The one that remembered nothing got 61%.
This is the story of why that happened, and what it took to build an agent where memory makes the plan better
without making the diagnosis worse.
What I built
Déjà Vu is an incident-response agent for the engineer holding the pager. When an alert fires, it recalls what
happened the last time this service broke and writes a plan: what to do, what not to do (with the incident
where it backfired), and who fixed it before. When the incident is over, the engineer marks each suggested step as
worked, no effect or made it worse, and that verdict becomes memory for the next page.
The memory layer is Hindsight, the open-source agent memory system from Vectorize.
I didn't want a vector store I'd have to babysit. I wanted something that behaves like a colleague who reads every
postmortem: it extracts facts, links entities, consolidates repeated evidence into beliefs, and answers questions
across all of it. If you're new to the idea, Vectorize has a good explainer on
what agent memory is and why it differs from retrieval.
Every postmortem goes in with its real timestamp, a stable document id and service tags, so later recalls can be
scoped and cited:
async def retain_incident(self, inc: dict[str, Any]) -> None:
await self.client.aretain(
bank_id=self.bank_id,
content=data.postmortem_text(inc),
context=f"incident postmortem for {inc['service']}",
timestamp=datetime.fromisoformat(inc["started_at"]),
document_id=inc["id"],
tags=["kind:postmortem", service_tag(inc["service"]), f"severity:{inc['severity']}"],
metadata={"incident_id": inc["id"], "service": inc["service"], "severity": inc["severity"]},
)
Team rules ("rollbacks must use the exact argocd app rollback command") and engineer verdicts go in the same way.
Each service also gets a Hindsight mental model: a standing question ("how does checkout-api fail, what worked, what
made it worse?") that Hindsight re-answers after new memories consolidate. Those became living runbooks nobody
maintains by hand. The Hindsight documentation covers retain, recall, reflect and
mental models in detail.
How I measured it
A demo where the agent says the right thing once proves nothing, so I built a replay. The history is six months of a
(fictional) UPI and card payments gateway: 23 incidents with real-looking HikariCP, pgbouncer, Kafka and CoreDNS log
lines, and every action someone tried along with its outcome. Six failure modes recur, usually with a different
trigger each time. Connection-pool exhaustion shows up four times, caused by a missing index, a migration lock, a
connection leak and autoscaling. Eight incidents are one-offs, and the first occurrence of each recurring failure is new too. Those 14
first-time incidents are the control group.
The replay walks the history in date order into a fresh memory bank. Before each incident is retained, two agents
triage it blind: the same model (gpt-oss-120b) with the same instructions, one with memory and one without. A
grader model compares both plans with the real postmortem. Because an incident is only retained after it has been
graded, nothing is ever in memory before it happens.
Failure one: memory as answers
My first prompt said, in effect, "trust memory over generic advice." The memory agent recognised every repeat
incident (9 of 9) and halved harmful suggestions. Its diagnosis was also worse than the agent with no memory at all.
The clearest case was a ledger consumer falling behind. The logs said:
WARN consumer poll timeout has expired ... longer than the configured max.poll.interval.ms
ledger-worker: fx_rates.get(currency='USD') took 4212ms (timeout 5000ms)
A four-second FX lookup inside the poll loop. The memory agent answered "max.poll.records is too high", because
that had been the cause the last time this service lagged. Same symptom, new trigger, wrong answer. It's exactly how
a mediocre senior engineer behaves: pattern-match first, read second.
The fix was to treat memory as hypotheses. The plan must now quote the evidence in this alert, mark for every
past incident it matches whether that incident's trigger is actually present (same_trigger), and say what is
different this time. In the UI that became a distinct verdict: "Seen this symptom before, but the trigger is new."
Failure two: the same thing, more subtly
That helped, but the replay caught it again. A checkout alert carried a stack trace from HikariCP's leak detector,
pointing at DbRetryTemplate.execute. The agent without memory read it correctly: a new retry path wasn't returning
connections. The agent with memory blamed pgbouncer's connection limit. pgbouncer appears nowhere in that alert. It
came from two past pool-exhaustion incidents that did involve pgbouncer.
So I changed the order of operations. Déjà Vu now reads the evidence first, with no memory at all, and memory has to
earn the right to change that diagnosis:
FIRST_READ = """
Your own first read of this alert, made from the evidence alone before looking at memory:
- evidence: {evidence}
- likely root cause: {root}
It is the working diagnosis. Change the root cause only if a past incident's own signature (its log line, component
or metric) appears in this alert and explains the evidence better. Either way, use memory for everything else below.
"""
In the side-by-side view this costs nothing extra: the memory lane reuses the no-memory lane's read, then checks it
against what Hindsight recalls. Run alone, it makes its own read while the recalls are in flight.
What memory is actually good for
Here is the final replay, graded blind:
| With memory | Without | |
|---|---|---|
| Root cause right, first-time incidents (control) | 75% | 75% |
| Root cause right, repeat incidents | 78% | 72% |
| Plan included the fix that actually worked, repeats | 6 / 9 | 4 / 9 |
| Plan included the fix that actually worked, all 23 | 14 / 23 | 9 / 23 |
| Recommended something that backfired last time | 2 | 2 |
The control row is the one I care about most. On incidents with no history, memory neither helps nor hurts, which
is what evidence-first is for. Diagnosis on repeats is only a little better, and that's honest: both agents read the
same logs. The real gap is operational knowledge. With memory, the plan contains the fix that worked last time and
the rollback in the form this team uses, and it knows what not to touch.
On the festive-sale alert (26 pods, no more connections allowed (max_client_conn)), the agent without memory
suggests raising the pgbouncer limit and resizing the pool. Sensible and generic. Déjà Vu says "Déjà vu: we have
seen this before", links INC-2107, applies the fix that worked there, rolls back with the exact Argo CD command the
team mandated, warns against restarting pgbouncer (it did nothing in INC-2014) and pages the engineer who fixed it
last time. The answer without memory arrives in about four seconds; Déjà Vu's in about eleven.
The grader was part of the system too
One thing surprised me. My first grader scored each plan on its own, and it gave near-identical answers different
scores: 1.0 in one lane, 0.5 in the other, for the same root cause. Now one call grades both plans side by side,
blind and in a fixed random order, at temperature 0, with the explicit rule that equivalent answers get equal grades.
The gap between the lanes shrank, and I trust it more.
What I learned
- Memory makes anchoring the default. Retrieval plus "trust your memory" produces an agent that diagnoses the last incident instead of this one. Make it read the evidence first.
- Run a control group. If memory "helps" on incidents it has never seen, something is leaking. The first-time incidents told me when a change was real.
- Measure the thing memory is for. Root-cause accuracy barely moved. "Did the plan contain the fix that worked?" and "did it avoid what backfired?" moved a lot.
- The grader is code you have to debug. Grade both answers in one blind call, or you'll optimise for noise.
- Keep the loop honest. Each fix here came from reading where the previous version missed, on the same 23 incidents. That's fine for learning; just don't call the result a benchmark.
The most useful thing an on-call agent can say isn't "I know what this is." It's "Don't do that. It made this
worse on 22 April, and here's who fixed it."
Code: https://github.com/Rushikumar-06/dejavu · Built with the Déjà Vu team · Thanks to @Code.in



Top comments (0)