What I built
This is an on-call assistant that diagnoses production alerts using an LLM, but the interesting part isn't the LLM call, it's what happens before and after it. Before the model sees an alert, the system recalls anything relevant from a persistent memory store. After an engineer tries a fix, the outcome, worked or failed, gets written back to that same store. Nothing about the model itself changes between those two events. The prompt is identical. The weights are identical. What changes is what the model is allowed to remember.
That memory layer is Hindsight, and the reason I picked it over rolling my own retrieval is that incident data doesn't arrive as clean keyword matches. "Gateway timeouts" and "504 errors" and "requests are timing out" all describe the same failure in different words, and a system that only matches on exact terms misses the connection an on-call engineer would make instantly. Hindsight's recall is built to catch that kind of paraphrase, along with time-based queries like "what broke last month," which matters more than it sounds like it should when you're trying to figure out if a problem is new or recurring.
The experiment: teach it nothing, then teach it once
Here's the alert I used to test the system's honesty: billing-service double-charged customers after a scheduler timezone change. Nothing about billing, timezones, or double-charging existed in the incident history. The closest thing by category was an unrelated cron job overlap in a different service, with a completely different root cause.
With memory turned on, the agent correctly said no match and gave general advice. That mattered more to me than a good answer would have, because the failure mode I was actually worried about was the opposite one: a retrieval system that always returns something, and a model that mistakes "something" for "relevant" just because it was handed to it. Getting a confident wrong guess back from a memory-backed system is worse than getting no guess at all, since it looks authoritative in exactly the situation where it's making things up.
I resolved the incident, described the real fix, and told the system it worked:
def record_outcome(alert: str, fix_tried: str, worked: bool):
if worked:
text = (f"For the alert '{alert}', the fix '{fix_tried}' WORKED and resolved the problem. "
f"Recommend this fix for similar alerts.")
else:
text = (f"For the alert '{alert}', the fix '{fix_tried}' FAILED and did NOT resolve the problem. "
f"Do not recommend '{fix_tried}' for similar alerts.")
with memory_client() as m:
m.retain(bank_id=BANK, content=text, context="fix outcome feedback")
One retain call. No retraining step, no batch job, no redeploy. About a minute later, once Hindsight finished processing the new memory, I sent a second alert, worded differently on purpose: billing-service charged customers twice after a daylight saving time change. Same underlying bug, different phrasing, and it had existed in memory for less time than it takes to make coffee.
The agent recalled the fix by name and recommended it directly, without me touching the prompt, the model, or the codebase in between. That's the entire pitch of building on a memory layer instead of hardcoding a lookup table: the thing you taught it generalizes to a differently-worded version of the same problem, because recall works on meaning, not string matching.
How the recall side actually decides what's relevant
The diagnosis function is short, and most of its value is in what it refuses to do rather than what it does:
def diagnose(alert: str, use_memory: bool = True) -> str:
past_text = ""
if use_memory:
with memory_client() as m:
past = m.recall(bank_id=BANK, query=alert)
past_text = "\n".join(f"- {r.text}" for r in past.results[:12])
prompt = f"""You are an on-call incident response assistant.
New alert: {alert}
Past incidents from memory:
{past_text or "None available."}
Rules:
- Line 2: which past incident it matches (or "no match").
- Use ONLY facts from the past incidents above. Never invent statistics.
- Only mention table or column names that appear in the past incidents."""
return ask_llm(prompt)
I capped the retrieved memories at twelve, not because Hindsight can't return more, it happily returned around ninety hits for a well-populated query, but because handing an LLM ninety loosely-related facts and expecting it to weigh them correctly is optimistic. Retrieval quality and prompt discipline are two separate problems, and building a good memory system doesn't excuse you from solving both.
The "use only facts from the memories above" rule exists because I watched the model, unprompted, invent a specific confidence percentage that appeared nowhere in my data. It sounded plausible. It was fabricated. Retrieval-augmented generation doesn't automatically make a model honest, it just gives it better material to be dishonest with if you don't explicitly constrain it.
Loading history without duplicating it
The billing incident above joined a dataset of pre-existing incidents I'd already loaded, and getting that loading step right mattered more than I expected the first time I had to re-run it:
client.retain(
bank_id=BANK,
content=inc["text"],
context="production incident post-mortem",
timestamp=f"{inc['date']}T10:00:00Z",
document_id=inc["id"],
)
The document_id field is what makes this safe to re-run. Without it, every re-run of a loader script duplicates every incident, and a memory store full of triplicate facts doesn't just waste space, it skews what the model treats as consensus. If three copies of the same connection-pool incident all show up in a retrieval result, the model has no way to know they're the same fact stated three times rather than three independent confirmations. Idempotent writes aren't a nice-to-have here, they're what keeps the memory trustworthy as the dataset grows.
The timestamp field matters for a related reason: it's how the system can eventually answer "has this happened recently" instead of just "has this ever happened." A fix that worked eight months ago on a system that's since been rearchitected is a very different signal than one that worked last week, and without an explicit timestamp on every memory, that distinction disappears.
What I'd tell someone building something similar
Test the "nothing matches" case on purpose, before you ever demo the "something matches" case. It's tempting to only show off the impressive retrievals. The alert with zero relevant history is the one that tells you whether your prompt is actually constraining the model or just decorating its guesses with citations.
A single retain call is a bigger deal than it looks. There's no fine-tuning step between "the agent doesn't know this" and "the agent knows this." That's the entire value proposition of Vectorize's agent memory approach over baking knowledge into a model: the update is immediate and it's just data, not a training run.
Idempotent writes are not optional once your dataset is anything other than a one-time script. I only added document_id after nearly duplicating my incident history on a second run. Do it from the first script, not the one after you get burned.
Cap what you hand the model, even when the memory system can return more. Ninety results is a great sign your retrieval works. It's a bad idea to send all ninety into a prompt and hope the model sorts out relevance on its own.
Generalization across phrasing is the actual test of a memory system, not exact recall. Anyone can build a lookup table keyed on the exact alert string. The point of putting an embedding-based memory layer underneath an LLM is that "gateway timeouts" and "daylight saving time change" describing the same root cause both land on the same memory. If your system only works when the wording matches exactly, you've built a cache, not a memory.



Top comments (0)