<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Annie</title>
    <description>The latest articles on DEV Community by Annie (@annie_431579b6bb6cc6b8ea3).</description>
    <link>https://dev.to/annie_431579b6bb6cc6b8ea3</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4150621%2F75e6af38-07dc-4e0c-9e71-0cbe9b3c04f8.png</url>
      <title>DEV Community: Annie</title>
      <link>https://dev.to/annie_431579b6bb6cc6b8ea3</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/annie_431579b6bb6cc6b8ea3"/>
    <language>en</language>
    <item>
      <title>A Memory ON/OFF Toggle Was My Best Hindsight Debugging Tool</title>
      <dc:creator>Annie</dc:creator>
      <pubDate>Tue, 29 Sep 2026 17:40:12 +0000</pubDate>
      <link>https://dev.to/annie_431579b6bb6cc6b8ea3/a-memory-onoff-toggle-was-my-best-hindsight-debugging-tool-202b</link>
      <guid>https://dev.to/annie_431579b6bb6cc6b8ea3/a-memory-onoff-toggle-was-my-best-hindsight-debugging-tool-202b</guid>
      <description>&lt;p&gt;Why My Incident Copilot Remembers What Failed, Not Just What Worked&lt;br&gt;
The most expensive thing an on-call engineer can do at 3 a.m. is take a reasonable-looking action that makes the outage worse. Restarting pods is reasonable. Scaling up replicas is reasonable. In one class of incident, both are exactly wrong, and the only thing that stops you is someone on the team remembering that it went badly last time.&lt;br&gt;
I built an incident copilot around that problem. It takes a new alert, recalls what happened in similar past incidents, and produces a plan that leads with what worked and explicitly lists what to avoid. The memory layer is Hindsight, and the decision that ended up mattering most was a boring one: I store the failures with the same care as the fixes.&lt;br&gt;
What the system does&lt;br&gt;
Incident Copilot is a small FastAPI service with a single-page UI. The shape is deliberately simple:&lt;br&gt;
Browser UI (app/static/index.html)&lt;br&gt;
        |&lt;br&gt;
   FastAPI (app/main.py)&lt;br&gt;
   /plan  /action  /close  /patterns&lt;br&gt;
        |                    |&lt;br&gt;
   agent.py              memory.py  ----&amp;gt;  Hindsight (retain / recall / reflect)&lt;br&gt;
   recall -&amp;gt; 1 LLM call&lt;br&gt;
        |&lt;br&gt;
   Groq: openai/gpt-oss-120b&lt;br&gt;
Four endpoints do all the work:&lt;br&gt;
• POST /plan takes an alert and symptoms, recalls memories, and returns a structured plan.&lt;br&gt;
• POST /action records the outcome of a step (worked, failed, harmful) while the incident is still open.&lt;br&gt;
• POST /close retains the full postmortem once the incident is resolved.&lt;br&gt;
• GET /patterns asks Hindsight to reflect on the recurring root causes and failed fixes it has learned.&lt;br&gt;
All Hindsight calls live in one file, app/memory.py. The agent itself is one recall followed by one LLM call. There is no planner loop, no tool-calling maze, and no vector-store plumbing for me to maintain. I wanted the memory layer to be the interesting part, and everything else to be readable in one sitting.&lt;br&gt;
If the idea is new to you, Vectorize has a good overview of what agent memory is. The short version: the model stays stateless and something else holds the history. I picked Hindsight for that "something else" because it exposes three operations that map cleanly onto an on-call workflow: retain, recall, and reflect. The Hindsight docs cover the full API; I only use those three.&lt;br&gt;
The through-line: outcomes are data&lt;br&gt;
Most postmortem tooling optimizes for the resolution. What fixed it? Write that down, because that's what you'll want next time.&lt;br&gt;
Partly. But in real incidents the resolution is usually the last of several attempts, and the earlier attempts are where the time goes. A typical incident on checkout-service goes like this: restart the pods (harmful, the reconnect storm exhausted the pool again), scale to 12 replicas (harmful, more replicas opened more DB connections and hit max_connections), roll back config deploy #4471 (worked, the pool size had been cut from 50 to 10). The rollback took 82 minutes to reach, and most of those minutes were the two wrong moves.&lt;br&gt;
A memory that only stores "rolled back deploy #4471" lets the next responder skip nothing. A memory that stores the two harmful attempts, along with why each one failed, lets them skip both.&lt;br&gt;
So every action in a retained incident carries an explicit outcome label, rendered into the text at retain time:&lt;br&gt;
def retain_incident(inc: dict):&lt;br&gt;
    actions = "\n".join(&lt;br&gt;
        f"- Attempted: {a['action']}. Outcome: {a['outcome'].upper()}. {a['note']}"&lt;br&gt;
        for a in inc["actions"]&lt;br&gt;
    )&lt;br&gt;
    content = (&lt;br&gt;
        f"Incident {inc['id']} on {inc['service']} ({inc['severity']}).\n"&lt;br&gt;
        f"Alert: {inc['alert']}\nSymptoms: {inc['symptoms']}\n"&lt;br&gt;
        f"Actions:\n{actions}\n"&lt;br&gt;
        f"Root cause: {inc['root_cause']}\nResolution: {inc['resolution']}\n"&lt;br&gt;
        f"Time to resolve: {inc['ttr_minutes']} minutes."&lt;br&gt;
    )&lt;br&gt;
    _retain(&lt;br&gt;
        bank_id=BANK,&lt;br&gt;
        content=content,&lt;br&gt;
        context="incident postmortem",&lt;br&gt;
        timestamp=inc["resolved_at"],&lt;br&gt;
        document_id=f"incident-{inc['id']}",&lt;br&gt;
        tags=[f"service:{inc['service']}", f"severity:{inc['severity']}"],&lt;br&gt;
        metadata={"incident_id": inc["id"]},&lt;br&gt;
    )&lt;br&gt;
A few choices in there are worth defending.&lt;br&gt;
WORKED, FAILED, and HARMFUL are uppercase words in the text, not just fields in a JSON blob. Hindsight extracts memories from natural language, so the outcome has to survive extraction. I wanted "restarting pods was harmful" to be a fact the memory layer can hold, not metadata that gets dropped.&lt;br&gt;
timestamp is the resolution time, not the ingest time. This matters more than it looks. If you backfill old postmortems and let the timestamp default to now, every memory looks fresh, and the agent can't tell a six-month-old runbook from last week's. Passing the real date is what makes staleness detection possible.&lt;br&gt;
document_id is stable. It's incident-, so retaining the same postmortem twice replaces it instead of duplicating it. Postmortems get edited: people fix typos and add follow-ups. Without an idempotent key, every edit would create a second, contradictory copy of the same incident, and recall would happily return both.&lt;br&gt;
The same idea applies mid-incident. The UI shows Worked / Failed / Harmful buttons on each recommended step, and clicking one hits /action, which retains a short note immediately:&lt;br&gt;
def log_action(incident_id: str, n: int, action: str, outcome: str, note: str = ""):&lt;br&gt;
    """Mid-incident feedback: the Worked / Failed / Harmful buttons."""&lt;br&gt;
    _retain(&lt;br&gt;
        bank_id=BANK,&lt;br&gt;
        content=f"During incident {incident_id}, attempted: {action}. Outcome: {outcome.upper()}. {note}",&lt;br&gt;
        context="live incident action",&lt;br&gt;
        document_id=f"incident-{incident_id}-action-{n}",&lt;br&gt;
    )&lt;br&gt;
If the on-call engineer marks "restart pods" as harmful at 03:10, a second engineer paged for a related alert at 03:25 already sees it. Nobody has to wait for the postmortem.&lt;br&gt;
Recall across services on purpose&lt;br&gt;
The design decision I'd argue for hardest is what I don't do at recall time: I don't filter by service.&lt;br&gt;
def recall_for_alert(alert: str, symptoms: str = ""):&lt;br&gt;
    query = f"{alert} {symptoms}".strip()&lt;br&gt;
    try:&lt;br&gt;
        resp = client.recall(bank_id=BANK, query=query, budget="high", max_tokens=3000)&lt;br&gt;
    except TypeError:&lt;br&gt;
        resp = client.recall(bank_id=BANK, query=query)&lt;br&gt;
    return resp.results&lt;br&gt;
The query is just the alert text plus the symptoms. The service name appears in that text, but nothing forces a match on it. That's intentional, because the cause of an outage is rarely a property of the service. A config deploy that shrinks a connection pool looks the same on checkout-service, orders-api, and payments-api. Even the stack trace says the same thing:&lt;br&gt;
HikariPool-1 - Connection is not available, request timed out after 30000ms&lt;br&gt;
The incident history I test against has four families: connection pool exhaustion, cache stampede after a Redis failover, expired internal mTLS certificates, and disks filled by unrotated logs. Each family spans several services. The held-out incident I use most is on payments-api; the earlier incidents with the same root cause are on checkout-service, orders-api, and inventory-service. A per-service index would find nothing. Recall keyed on symptoms finds all three.&lt;br&gt;
I set the recall budget to "high" with 3,000 tokens. That's a deliberate cost decision: recall happens once per alert, during an incident, when latency matters far less than surfacing the right memory. If I were running it per keystroke I'd feel differently.&lt;br&gt;
One piece of honest friction: the try/except TypeError exists because the Hindsight client's keyword arguments differed between versions I worked with, and I got tired of upgrades breaking the app. The same shim appears in _retain, which drops tags and metadata if a client rejects them. I'd rather have one slightly ugly compatibility layer than version checks scattered around the codebase. Pin your client version and keep the wrapper where you can find it.&lt;br&gt;
Making the model use it honestly&lt;br&gt;
Recall gives the model memories. It doesn't guarantee the model uses them well. The system prompt is where I constrain that:&lt;br&gt;
SYSTEM = """You are an on-call incident copilot. You get a NEW ALERT and MEMORIES from past incidents.&lt;br&gt;
Rules:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Put fixes that worked in past similar incidents first.&lt;/li&gt;
&lt;li&gt;Explicitly list actions that were HARMFUL or FAILED in similar incidents under "avoid", with the reason.&lt;/li&gt;
&lt;li&gt;Cite the incident id for every claim. Never invent incident ids.&lt;/li&gt;
&lt;li&gt;If a memory is older than 6 months (compare with TODAY), add a staleness note.&lt;/li&gt;
&lt;li&gt;If memories are NONE, weak or unrelated, say so and give general advice, clearly labelled as general advice with source null.
Return ONLY JSON: ..."""
Three of those rules exist because of failure modes I care about:&lt;/li&gt;
&lt;li&gt; Citations on every claim. In the UI, each step has a source link. Clicking it highlights the matching card in the "Memories used" panel. An engineer under pressure shouldn't have to trust the model; they should be able to check in one click.&lt;/li&gt;
&lt;li&gt; A dedicated avoid list. Without it, harmful actions get buried in prose or omitted. Making it a required structured field forces the model to surface them.&lt;/li&gt;
&lt;li&gt; Explicit "I don't know" behavior. If recall returns nothing relevant, the plan says so and labels its advice as general, with a null source. A confident answer built on no evidence is worse than no answer.
The staleness rule works because of the resolution timestamp I insisted on earlier. The agent gets TODAY in the prompt and each memory's date in the context, so "this was true in February" is something it can actually reason about.
Behavior: what a plan looks like
Take the alert I use most:
payments-api authorization requests stalling, queue depth 3.2k and rising
Card authorizations queue up while pod CPU stays low. Traces show requests idle waiting on the Postgres client, and logs show 'HikariPool-1 - Connection is not available, request timed out after 30000ms'.
The UI has a Memory ON/OFF toggle, which sends use_memory to /plan. It's the most useful debugging tool in the project because it gives a like-for-like comparison: same model, same prompt, same alert, with and without recall.
With memory off, the plan is competent and generic: check the pool, look at recent deploys, consider restarting or scaling. That advice is sensible in isolation, and it includes the two actions that have historically made this specific failure worse.
With memory on, the plan changes shape. The first step is to roll back the recent config deploy, citing the earlier pool-exhaustion incidents. The avoid section lists restarting pods and scaling replicas, each with the reason from the incident where it backfired, and each linked to its source. If a cited incident is more than six months old, a staleness note appears above the plan. In the "Memories used" panel I can see the recalled records, typed and dated.
The moment that sold me was seeing a citation land on an incident from a different service. The agent wasn't pattern-matching on names; it was matching on what actually went wrong.
The other endpoint worth mentioning is /patterns, which calls Hindsight's reflect operation with the query "What recurring root causes and failed fixes have you learned? Cite incident IDs." It's the closest thing to an automatic review of the postmortem archive. I use it less during incidents and more between them, to see which failure classes keep recurring.
Evaluating it without cheating
It's easy to fool yourself with memory systems. If you test on incidents already in the store, you're measuring retrieval of an exact match, not learning.
The evaluation script replays the incident history in date order against a fresh bank. For each incident, it plans with memory off and on, grades both, and only then retains that incident:
for n, inc in enumerate(incidents):
row = {"id": inc["id"], "n_memories_before": n}
for mode in ("off", "on"):
    p = agent.plan(inc["alert"], inc["symptoms"], use_memory=(mode == "on"))
    row[mode] = grade(p, inc)
results.append(row)
retain_incident(inc)          # only AFTER planning, so it never sees itself
time.sleep(5)                 # give consolidation a moment
That comment is the whole point. Retaining after planning means the agent never sees the answer to the question it's being asked, and each incident is evaluated against only the history that would have existed at that time. The output is a learning curve: the cumulative wrong-first-suggestion rate against the number of past incidents in memory.
One caveat, stated plainly: grading is done by an LLM judge that checks whether the first step matches the known-correct fix and whether any known-bad action is recommended as something to do. That's useful for trends and unreliable for exact numbers. I spot-check the judged rows by hand, and you should too if you reuse this approach.
Lessons learned
Store what failed, not only what worked. Failed attempts account for most of the elapsed time in an incident, and they're exactly what a responder can skip if they know about them. Write the outcome into the text, not just a field.
Pass real timestamps when you retain. Backfilled memories with ingest-time dates make everything look current. Once the dates are real, staleness handling is a prompt rule instead of a project.
Give records stable IDs. Postmortems get edited. An idempotent document_id keeps a corrected record from becoming two conflicting ones.
Don't scope recall more tightly than the cause is scoped. Root causes cross service boundaries, and filtering by service would have hidden exactly the incidents that matter. Tags are still there for slicing later; I just don't gate recall on them.
Build the off switch first. A memory ON/OFF toggle gave me a clean A/B on every alert, and the eval script uses the same flag. If you can't turn memory off, you can't tell what it's contributing.
Expect to wait after writing. Retained content is consolidated in the background, so a recall run immediately after a write can look empty or thin. The seed script prints a reminder to wait a minute or two, and the eval loop sleeps between steps. I learned this by staring at an empty result and assuming my code was broken.
If you want to look at the memory layer itself, the Hindsight GitHub repository is where I'd start, and the Hindsight documentation explains retain, recall, and reflect in more depth than I have room for here.
&lt;a href="![%20](https://dev-to-uploads.s3.us-east-2.amazonaws.com/uploads/articles/6f8778h16mqz7obtu0cl.jpeg)"&gt;&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvz7lcs1xpowvcx6uv2j7.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvz7lcs1xpowvcx6uv2j7.jpeg" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fanh8ba92z8r2qp19hn14.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fanh8ba92z8r2qp19hn14.jpeg" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
  </channel>
</rss>
