DEV Community

Tanvi reddy
Tanvi reddy

Posted on

A Memory ON/OFF Toggle Was My Best Hindsight Debugging Tool

The most expensive thing an on-call engineer can do at 3 a.m. is a reasonable-looking action that makes the outage worse. Restarting pods is reasonable. Scaling up replicas is reasonable. In one class of incident, both are exactly wrong, and the only thing that stops you is someone on the team remembering that it went badly last time.
I built an incident copilot around that problem. It takes a new alert, recalls what happened in similar past incidents, and produces a plan that leads with what worked and explicitly lists what to avoid. The memory layer is https://github.com/vectorize-io/hindsight, and the decision that ended up mattering most was a boring one: I store the failures with the same care as the fixes.
What the system does
Incident Copilot is a small FastAPI service with a single-page UI. The shape is deliberately simple:
Browser UI (app/static/index.html)
|
FastAPI (app/main.py)
/plan /action /close /patterns
| |
agent.py memory.py ----> Hindsight (retain / recall / reflect)
recall -> 1 LLM call
|
Groq: openai/gpt-oss-120b

Four endpoints do all the work:
POST /plan takes an alert and symptoms, recalls memories, and returns a structured plan.
POST /action records the outcome of a step (worked, failed, harmful) while the incident is still open.
POST /close retains the full postmortem once the incident is resolved.
GET /patterns asks Hindsight to reflect on what recurring root causes and failed fixes it has learned.
All Hindsight calls live in one file, app/memory.py. The agent itself is one recall followed by one LLM call. There is no planner loop, no tool-calling maze, and no vector-store plumbing that I maintain. I wanted the memory layer to be the interesting part, and for the rest to be easy to read in one sitting.
If you're new to the idea, Vectorize has a good overview of (https://vectorize.io/what-is-agent-memory). The short version is that the model stays stateless and something else holds the history. I picked Hindsight for that "something else" because it exposes three operations that map cleanly onto an on-call workflow: retain, recall, and reflect. The (https://hindsight.vectorize.io/) cover the full API; I only use those three.
The through-line: outcomes are data
Most postmortem tooling optimizes for the resolution. What fixed it? Write that down. That's the thing you'll want next time, right?
Partly. But if you look at how incidents actually go, the resolution is usually the last of several attempts. The earlier attempts are where the time goes. In my seed data, an incident on checkout-service looks like this: restart the pods (harmful, the reconnect storm exhausted the pool again), scale to 12 replicas (harmful, more replicas opened more DB connections and hit max_connections), roll back config deploy #4471 (worked, the pool size had been cut from 50 to 10). The rollback took 82 minutes to reach. Most of those minutes were the two wrong moves.
A memory that only stores "rolled back deploy #4471" lets the next responder skip zero of that. A memory that stores the two harmful attempts, with the reason each one failed, lets them skip both.
So every action in a retained incident carries an explicit outcome label, and I render it into the text at retain time:
def retain_incident(inc: dict):
actions = "\n".join(
f"- Attempted: {a['action']}. Outcome: {a['outcome'].upper()}. {a['note']}"
for a in inc["actions"]
)
content = (
f"Incident {inc['id']} on {inc['service']} ({inc['severity']}).\n"
f"Alert: {inc['alert']}\nSymptoms: {inc['symptoms']}\n"
f"Actions:\n{actions}\n"
f"Root cause: {inc['root_cause']}\nResolution: {inc['resolution']}\n"
f"Time to resolve: {inc['ttr_minutes']} minutes."
)
_retain(
bank_id=BANK,
content=content,
context="incident postmortem",
timestamp=inc["resolved_at"],
document_id=f"incident-{inc['id']}",
tags=[f"service:{inc['service']}", f"severity:{inc['severity']}"],
metadata={"incident_id": inc["id"]},
)

A few choices in there are worth defending.
WORKED, FAILED, and HARMFUL are uppercase words in the text, not just fields in a JSON blob. Hindsight extracts memories from natural language, so the outcome needs to be something that survives extraction. I wanted "restarting pods was harmful" to be a fact the memory layer can hold, not metadata that gets dropped on the floor.
timestamp is the resolution time, not the ingest time. This matters more than it looks. If you backfill old postmortems and let the timestamp default to now, every memory looks fresh, and the agent can't tell a six-month-old runbook from last week's. Passing the real date is what makes staleness detection possible later.
document_id is stable. It's incident-, so retaining the same postmortem twice replaces it rather than duplicating it. Postmortems get edited. People fix typos and add follow-ups. Without an idempotent key, every edit would create a second, contradictory copy of the same incident, and recall would happily return both.
The same idea applies mid-incident. The UI shows Worked / Failed / Harmful buttons on each recommended step, and clicking one hits /action, which retains a short note immediately:
def log_action(incident_id: str, n: int, action: str, outcome: str, note: str = ""):
"""Mid-incident feedback: the Worked / Failed / Harmful buttons."""
_retain(
bank_id=BANK,
content=f"During incident {incident_id}, attempted: {action}. Outcome: {outcome.upper()}. {note}",
context="live incident action",
document_id=f"incident-{incident_id}-action-{n}",
)

If the on-call engineer marks "restart pods" as harmful at 03:10, a second engineer who gets paged for a related alert at 03:25 already sees it. Nobody has to wait for the postmortem.
Recall across services on purpose
The design decision I'd argue for hardest is what I don't do at recall time: I don't filter by service.
def recall_for_alert(alert: str, symptoms: str = ""):
query = f"{alert} {symptoms}".strip()
try:
resp = client.recall(bank_id=BANK, query=query, budget="high", max_tokens=3000)
except TypeError:
resp = client.recall(bank_id=BANK, query=query)
return resp.results

The query is just the alert text plus the symptoms. The service name shows up in that text, but nothing forces a match on it. That's intentional, because the cause of an outage is rarely a property of the service. A config deploy that shrinks a connection pool will look the same on checkout-service, orders-api, and payments-api. The stack trace even says the same thing:
HikariPool-1 - Connection is not available, request timed out after 30000ms

In the data I built to exercise this, there are four incident families: connection pool exhaustion, cache stampede after a Redis failover, expired internal mTLS certificates, and disks filled by unrotated logs. Each family spans several different services. The held-out incident that I test with is on payments-api; the seeded ones with the same root cause are on checkout-service, orders-api, and inventory-service. A per-service index would find nothing. A recall keyed on symptoms finds all three.
I set the recall budget to "high" and give it 3,000 tokens. That's a cost decision I made on purpose: recall happens once per alert, during an incident, when latency matters far less than getting the right memory. If I were running this per keystroke I'd feel differently.
One piece of honest friction: the try/except TypeError in there exists because the Hindsight client's keyword arguments differed between versions I was working with, and I got tired of the app breaking on upgrades. The same shim appears in _retain, which drops tags and metadata if a client rejects them. I'd rather have a slightly ugly compatibility layer in one file than scatter version checks around the codebase. Pin your client version, and keep the wrapper where you can find it.
Making the model use it honestly
Recall gives the model memories. It doesn't guarantee the model will use them well. The system prompt is where I constrain that:
SYSTEM = """You are an on-call incident copilot. You get a NEW ALERT and MEMORIES from past incidents.
Rules:

  • Put fixes that worked in past similar incidents first.
  • Explicitly list actions that were HARMFUL or FAILED in similar incidents under "avoid", with the reason.
  • Cite the incident id for every claim. Never invent incident ids.
  • If a memory is older than 6 months (compare with TODAY), add a staleness note.
  • If memories are NONE, weak or unrelated, say so and give general advice, clearly labelled as general advice with source null. Return ONLY JSON: ..."""

Three of those rules exist because of failure modes I care about:
Citations on every claim. In the UI, each step has a source link. Clicking it highlights the matching card in the "Memories used" panel. An engineer under pressure shouldn't have to trust the model. They should be able to check in one click.
A dedicated avoid list. Without it, harmful actions get buried in prose or omitted. Making it a required structured field forces the model to surface them.
Explicit "I don't know" behavior. If recall returns nothing relevant, the plan says so and labels its advice as general, with a null source. A confident answer built on no evidence is worse than no answer.
The staleness rule works because of that resolution timestamp I insisted on earlier. The agent gets TODAY in the prompt and each memory's date in the context, so "this was true in February" is something it can actually reason about.
Behavior: what a plan looks like
Take the alert I use most:
payments-api authorization requests stalling, queue depth 3.2k and rising
Card authorizations queue up while pod CPU stays low. Traces show requests idle waiting on the Postgres client, and logs show 'HikariPool-1 - Connection is not available, request timed out after 30000ms'.

The UI has a Memory ON/OFF toggle, which sends use_memory to /plan. It's the most useful debugging tool in the project, because it gives me a like-for-like comparison: same model, same prompt, same alert, with and without recall.
With memory off, the plan is competent and generic. Check the pool, look at recent deploys, consider restarting or scaling. That advice is sensible in isolation, and it includes the two actions that have historically made this specific failure worse.
With memory on, the plan changes shape. The first step is to roll back the recent config deploy, citing the earlier pool-exhaustion incidents. The avoid section lists restarting pods and scaling replicas, each with the reason from the incident where it backfired, and each linked to its source. If the cited incident is more than six months old, a staleness note appears above the plan. In the "Memories used" panel, I can see the recalled records, typed and dated.
The moment that sold me on the approach was seeing the citation land on an incident from a different service. The agent wasn't pattern-matching on names. It was matching on what actually went wrong.
The other endpoint worth mentioning is /patterns, which calls Hindsight's reflect operation with the query "What recurring root causes and failed fixes have you learned? Cite incident IDs." It's the closest thing to an automatic review of the postmortem archive. I use it less during incidents and more between them, to see which failure classes keep recurring.
Evaluating it without cheating
It's easy to fool yourself with memory systems. If you test on incidents that are already in the store, you're measuring retrieval of an exact match, not learning.
The evaluation script replays the incident history in date order against a fresh bank. For each incident, it plans with memory off and with memory on, grades both, and only then retains that incident:
for n, inc in enumerate(incidents):
row = {"id": inc["id"], "n_memories_before": n}
for mode in ("off", "on"):
p = agent.plan(inc["alert"], inc["symptoms"], use_memory=(mode == "on"))
row[mode] = grade(p, inc)
results.append(row)
retain_incident(inc) # only AFTER planning, so it never sees itself
time.sleep(5) # give consolidation a moment

That comment is the whole point. Retaining after planning means the agent never sees the answer to the question it's being asked, and each new incident is evaluated against only the history that would have existed at that time. The output is a learning curve of the cumulative wrong-first-suggestion rate against the number of past incidents in memory.
One caveat I'll state plainly: the grading is done by an LLM judge that checks whether the first step matches the known-correct fix and whether any known-bad action is recommended as something to do. That's useful for trends and unreliable for exact numbers. I spot-check the judged rows by hand, and you should too if you reuse this approach.

Lessons learned
Store what failed, not only what worked. Failed attempts are most of the elapsed time in an incident, and they're the part a responder can skip if they know about it. Write the outcome into the text, not just a field.
Pass real timestamps when you retain. Backfilled memories with ingest-time dates make everything look current. Once the dates are real, staleness handling is a prompt rule instead of a project.
Give records stable IDs. Postmortems get edited. An idempotent document_id keeps a corrected record from becoming two conflicting ones.
Don't scope recall more tightly than the cause is scoped. Root causes cross service boundaries. Filtering by service would have hidden exactly the incidents that matter. Tags are still there for slicing later; I just don't gate recall on them.
Build the off switch first. A memory ON/OFF toggle gave me a clean A/B on every alert, and the eval script uses the same flag. If you can't turn memory off, you can't tell what it's contributing.
Expect to wait after writing. Retained content is consolidated in the background, so a recall run immediately after a seed can look empty or thin. The seed script prints a reminder to wait a minute or two, and the eval loop sleeps between steps. I learned this by staring at an empty result and assuming my code was broken.
If you want to look at the memory layer itself, the https://github.com/vectorize-io/hindsight is where I'd start, and the https://github.com/vectorize-io/hindsight explains retain, recall, and reflect in more depth than I have room for here.

Top comments (0)