DEV Community

Navya Reddy
Navya Reddy

Posted on

A Memory ON/OFF Toggle Was My Best Hindsight Debugging Tool

The most useful feature in our incident assistant isn't the recall pipeline or the prompt. It's a button labeled "Memory" that flips between ON and OFF, and it's the only reason I can tell whether the memory layer is doing anything at all.

What Incident Copilot does

Incident Copilot is an on-call assistant that remembers past incidents. You paste in a new alert and its symptoms. It recalls similar past incidents, recommends what worked, and tells you which fixes made things worse last time. Every claim in the plan has to cite an incident ID.

The stack is deliberately small:

A FastAPI backend with four endpoints: /plan, /action, /close and /patterns
A single-page UI that renders the plan and lets the on-call engineer mark each step as Worked, Failed or Harmful
One LLM call per plan (openai/gpt-oss-120b on Groq)
Hindsight as the memory layer, behind a single module, memory.py

There is no agent loop, no tool-calling and no planner. Recall happens in code, the results go into one prompt, and the model returns JSON. I made that choice on purpose. If the model decides whether to consult memory, a bad answer could be a retrieval failure, a tool-selection failure or a reasoning failure, and I can't tell which. With one fixed path, only two things can differ between runs.

That is also why the toggle exists.

The through-line: make memory a variable you can flip

Agent memory is hard to evaluate because a plausible answer looks the same whether it came from memory or from the model's general knowledge. "Check the connection pool" is a good suggestion for a latency alert with or without a memory system. If I only ever run the system with memory on, I can convince myself it's working.

So the whole request path takes a use_memory flag:

python
def plan(alert: str, symptoms: str = "", use_memory: bool = True) -> dict:
memories = []
if use_memory:
try:
memories = [_mem_dict(m) for m in recall_for_alert(alert, symptoms)]
except Exception as e:
return _error(f"Memory unavailable (is Hindsight running?): {e}")

mem_text = "\n".join(
    f"[{m['type']}] ({m['date'] or 'undated'}) {m['text']}" for m in memories
) or "NONE"
Enter fullscreen mode Exit fullscreen mode

The OFF path isn't a separate code path. It sends the same prompt, the same model and the same temperature, with the memory block replaced by the literal string NONE. The system prompt already handles that case: if memories are missing, weak or unrelated, say so and label any advice as general, with source: null.

The UI toggle is a boolean passed to /plan. When a plan looks good, I can flip the button, rerun the same alert, and see what the model says without agent memory.

The same flag drives the automated evaluation, which I get to below.

How Hindsight is used

All Hindsight calls live in one file. The Hindsight docs describe retain, recall and reflect as separate operations, and I use all three for different jobs.

Retain: postmortems as documents with outcomes. Each resolved incident is stored as a structured postmortem. The important part is that every attempted action carries an explicit outcome:

python
actions = "\n".join(
f"- Attempted: {a['action']}. Outcome: {a['outcome'].upper()}. {a['note']}"
for a in inc["actions"]
)
...
_retain(
bank_id=BANK,
content=content,
context="incident postmortem",
timestamp=inc["resolved_at"],
document_id=f"incident-{inc['id']}",
tags=[f"service:{inc['service']}", f"severity:{inc['severity']}"],
metadata={"incident_id": inc["id"]},
)

Three details matter here:

WORKED, FAILED and HARMFUL are written into the text. Restarting the pods and rolling back the config both appear in an incident, and the memory has to say which one helped.
The timestamp is the resolution time, not the ingestion time. The prompt tells the model to add a staleness note when a memory is older than six months, and that only works if dates mean something.
document_id is stable (incident-INC-113), so retaining an incident twice updates it instead of duplicating it. That makes the seed script safe to re-run.

Recall: across services, not within one. The recall query is the alert text plus the symptoms, with a high budget. I deliberately don't filter by service. The tags are there for inspection, not for narrowing. This turned out to be the most valuable behavior in the system, and the next section shows why.

Live feedback. During an incident, the Worked / Failed / Harmful buttons call /action, which retains each outcome immediately under its own document ID (incident-{id}-action-{n}). If the second responder gets paged twenty minutes later, they see what the first one already tried.

Reflect. The Patterns panel asks Hindsight one question: what recurring root causes and failed fixes it has learned, citing incident IDs. I don't post-process the answer.

What the toggle showed

The seed data has four failure families: connection pool exhaustion after a config deploy, cache stampedes after a Redis failover, expired internal TLS certificates, and disks full of unrotated logs. Each family has three incidents on three different services in the seed set, plus one held out for testing.

The held-out pool-exhaustion incident is INC-113 on payments-api. The alert reads "authorization requests stalling, queue depth 3.2k and rising". Nothing in that line mentions a database, a pool or a deploy. The symptoms field has the HikariPool-1 - Connection is not available log line, and that is the only overlap with the earlier incidents, which happened on checkout-service, orders-api and inventory-service. No prior incident involved payments-api.

With memory OFF, the model only has the alert text, and the prompt tells it to label any advice as general. Nothing in that text rules out a pod restart, which is a common first move, and it is exactly the action that made things worse in every earlier incident in that family, because the reconnect storm exhausts the pool again.

With memory ON, recall surfaces the earlier pool-exhaustion incidents from other services. The plan leads with checking for a recent config deploy that changed pool settings, and the restart shows up under avoid with the reason and the incident IDs attached.

That contrast is the argument for the whole design. The flip side is just as useful, though. When memory is ON and the answer is still bad, I can look at _memories, which the API returns alongside every plan, and decide whether recall missed or the model ignored what it was given. Being able to answer that quickly is why I trust the system.

Evaluating it without fooling myself

The toggle also drives scripts/eval_learning_curve.py, which replays every incident in date order against a fresh bank:

python
for n, inc in enumerate(incidents):
row = {"id": inc["id"], "n_memories_before": n}
for mode in ("off", "on"):
p = agent.plan(inc["alert"], inc["symptoms"], use_memory=(mode == "on"))
row[mode] = grade(p, inc)
results.append(row)
retain_incident(inc) # only AFTER planning, so it never sees itself
time.sleep(5) # give consolidation a moment

Every incident is planned twice, once with memory and once without, and only afterwards retained. The comment on that retain call is there because an incident can't be allowed to recall itself; otherwise the comparison is meaningless.

The output is a cumulative wrong-first-suggestion rate as the bank grows. The first few incidents should look identical between modes, because there's nothing to recall. If ON pulls away from OFF as the count rises, memory is contributing.

Two caveats I'd want to hear as a reader:

The grader is an LLM judge checking whether the first step is the correct fix and whether any known-bad action is recommended. I spot-check its output, and you should too.
The incident set is small, with four failure families. It shows the mechanism working. It says nothing about how much a real on-call rotation would gain, and I'm not going to quote a percentage the data can't support.
Things that hurt

Consolidation delay. Hindsight processes what you retain before it becomes fully useful for recall. Recall right after seeding can be thin. The seed script prints a reminder to wait a minute or two, and the evaluation sleeps between incidents so the memory-ON path isn't unfairly penalized.

Client API drift. memory.py has a _retain wrapper that retries without tags and metadata if the installed client rejects them. It's ugly, and it's there because client versions differ in which keyword arguments they accept. I'd rather have the ugliness in one file than scattered through the app.

Synthetic corpus, real patterns. To build test data I fixed each family's root cause and log fragments, then had a model write varied incident records around them, with a validity check that the last action is always the one that worked. That is fine for testing the mechanism. It means the corpus is tidier than a real postmortem archive, which is another reason to treat the eval as a mechanism check and not as a forecast.

Lessons learned
Ship the ablation switch in the product. A use_memory flag on the request path costs a few lines and gives you a permanent way to answer "is this thing doing anything?" from the UI, the API and the eval script.
Store outcomes, not just events. "Restarted the pods" is useless as a memory. "Restarted the pods, HARMFUL, reconnect storm exhausted the pool" changes the recommendation. The ability to say "don't do that" is worth more than the ability to say "try this".
Don't scope recall too early. The interesting retrievals crossed service boundaries. Filtering on the alerting service would have hidden the exact matches that mattered.
Return the evidence with the answer. Every plan ships with the memories it was built from. Debugging memory needs the retrieval result next to the output, so it's part of the response instead of a log line.
Keep the pipeline boring. Recall in code plus one LLM call means the variables are the memories and the prompt. I'd add tools and loops only after the boring version stops being good enough, and I'd keep the toggle when I did.

The Hindsight GitHub repo has the memory layer itself, and the Hindsight docs cover retain, recall and reflect in more detail than I have here. If you're building anything that claims to remember, build the OFF switch first.



Top comments (0)