The fact that would have stopped my on-call agent from making a database outage worse was already in memory. It was
recall result #30 of 35, and my code passed the top 18 to the model.
I build Déjà Vu, an incident-response agent that remembers every postmortem, every "that made it worse", and every
verdict an engineer gives on its suggestions. This post is about the least glamorous part of it, retrieval, and why
getting it right mattered more than any prompt I wrote.
The setup
When an alert fires, Déjà Vu recalls what happened the last time this service broke and writes a plan: what to do,
what not to do (and where it backfired), and who fixed it before. Memory lives in
Hindsight, an open-source agent memory system. Postmortems, team rules
and engineer verdicts are retained with their real timestamps, a document id per incident, and tags like
service:checkout-api. Hindsight extracts facts from them, links entities, and consolidates repeated evidence into
observations such as "during INC-2041, scaling checkout-api from 6 to 10 pods exacerbated the issue by exhausting
pgbouncer max_client_conn."
Recall is hybrid (semantic, BM25, entity graph and temporal, then reranked), so I assumed the hard part was done.
It wasn't.
The bug that wasn't in the model
My first agent searched memory with the alert text, took the results in order, and let the model run up to three
rounds of "is anything missing?" tool calls before writing a plan. On a festive-sale alert, checkout returning 5xx
because pgbouncer had hit max_client_conn after the autoscaler reached 26 pods, it produced a confident plan that
said scaling had made INC-2107 worse. It hadn't. Scaling made INC-2014 and INC-2041 worse; INC-2107's fix was
capping the autoscaler and resizing the pool.
When I printed exactly what reached the model, the reason was obvious. Searching with the alert text ranks memories
that look like the alert first: other alerts, other error lines. Hindsight had returned 35 relevant memories, and
the two that mattered were near the bottom:
- #29: Remediation for INC-2107: lowered Hikari maximumPoolSize, raised max_client_conn, capped HPA maxReplicas. This worked.
- #30: Scaling checkout-api from 6 to 10 pods during INC-2041 MADE IT WORSE.
My code kept the first 18. The model never saw either one. It filled the gap with a plausible guess and a
plausible-looking citation.
Team rules had a different problem: they never arrived at all. I scoped recall to the alert's service with
tags_match="any", which in Hindsight also keeps untagged memories. I'd assumed team rules would slip through
that way. But I had tagged them kind:note, so they were neither untagged nor matching. The Argo CD rollback
convention the team cared about never reached the festive-sale triage at all.
Ask memory the questions you actually have
The fix wasn't a better embedding. It was asking better questions. The agent now makes three recalls in parallel,
each with a clear job:
await asyncio.gather(
recall("similar incidents", query=query, tags=scope, max_tokens=1500),
recall("what was tried", query=outcome_query(alert, service), tags=scope, max_tokens=1000),
recall("team rules", query=query, tags=["kind:note"], strict=True, budget="low", max_tokens=500),
)
The second query is written in the language of outcomes, not symptoms: "On checkout-api, for incidents like ...,
which remediation actions WORKED, had NO EFFECT or MADE IT WORSE, and who resolved them?" The third uses strict
tag matching, so it returns only team rules and nothing else can crowd them out:
resp = await self.client.arecall(
bank_id=self.bank_id,
query=query,
types=["world", "experience", "observation"],
budget=budget,
max_tokens=max_tokens,
query_timestamp=at,
tags=tags,
tags_match="any_strict" if strict else "any",
)
The scope is the service plus platform, because DNS, NAT and certificates break everyone. The
Hindsight documentation is precise about these match modes; I just hadn't read
that part carefully enough.
Shape the context, then check the citations
Passing everything recalled isn't enough if it arrives as a flat list. Facts about four different pool-exhaustion
incidents look alike, and a model will happily credit one incident's outcome to another. So recalled memories are
now grouped per incident, oldest first, with team rules in their own section. The model reads something closer to a
stack of postmortems than a bag of sentences.
The UI links every incident ID to its postmortem, so a wrong citation is worse than none. The last step is plain
code:
def _check_citations(plan: dict[str, Any], memories: list[dict[str, Any]]) -> dict[str, Any]:
"""The UI links every incident ID, so a wrong citation is worse than none: keep only incidents that were recalled."""
known = {i for m in memories for i in m["incidents"]}
plan["matched_incidents"] = [m for m in plan["matched_incidents"] if m["id"] in known]
plan["seen_before"] = plan["seen_before"] and bool(plan["matched_incidents"])
...
Runbooks nobody writes
The other Hindsight feature I lean on is mental models. Each service gets one: a standing question that Hindsight
re-answers when new memories consolidate.
await self.client.acreate_mental_model(
bank_id=self.bank_id,
id=rb["id"],
name=rb["name"],
source_query=rb["query"],
tags=[service_tag(rb["service"])] if rb["service"] else None,
max_tokens=1200,
trigger={"refresh_after_consolidation": True},
)
After an incident, the engineer marks each suggested step as worked, no effect or made it worse. That verdict is
retained as a new incident document. In one run I recorded that capping the autoscaler fixed the festive-sale alert
in 12 minutes. It became INC-2117, the next triage of the same alert led with that step and cited INC-2117, and a
minute later the checkout-api runbook had rewritten itself to include it.
What changed
On the festive-sale alert, measured against the same memory bank on Groq's free tier:
| Before | After | |
|---|---|---|
| Time until the plan is on screen | 61 s | 11 s |
| LLM tokens for a side-by-side triage | 16.2k | 9.1k |
| Waits on the per-minute rate limit | 4 (44 s) | 0 |
The tool-call loop had used more than half of the tokens, and at 8k tokens per minute it spent most of its time
waiting. Three recalls take about a second and a half on Hindsight Cloud. The plan now cites INC-2107 for the fix
that worked there, warns against restarting pgbouncer (it did nothing in INC-2014), and rolls back with
argocd app rollback checkout-api <history-id>, the exact form the team's rule demands.
Over a replay of six months of synthetic but realistic history, graded blind against the real postmortems, plans
with memory contained the fix that actually worked in 14 of 23 incidents, against 9 of 23 for the same model without
memory.
What I learned
- Similarity to the alert is not usefulness. The alert tells you what's broken; the memory you need is what fixed it and what backfired. Ask for those directly.
-
Know your tag semantics.
anykeeping untagged memories is a feature, until your rules are tagged. Give each kind of memory a recall that can't be crowded out. - Shape the context before blaming the model. Grouping facts per incident fixed a citation error no prompt had.
- Verify what the model cites, in code. It's ten lines, and it's the difference between a tool engineers trust and one they double-check.
- The cheapest agent loop is often no loop. Three parallel, well-aimed recalls beat three rounds of "is anything missing?" on quality, latency and cost.
If you're building on agent memory, print what your model actually
sees. Mine had the right answer the whole time. It was sitting at position 30.
Code: https://github.com/Rushikumar-06/dejavu · Built with the Déjà Vu team · Thanks to @Code.in


Top comments (0)