DEV Community

Rama krishna
Rama krishna

Posted on

My Agent Remembered the Fix and Was Wrong

Last month my incident-recovery agent looked at a failing search query, remembered it had fixed the same symptom before, and proposed the same fix. It would have been the wrong one.

The bug was a search gateway returning results from the wrong corpus version. The first time, an alias was pointing at the old collection. This time the alias was already correct, and the real problem was somewhere else.

What AfterTrace does

AfterTrace is a command-line tool that helps recover from incidents. It follows one loop: detect, diagnose, propose, approve, fix, verify, retain.

The rule I built around is: memory proposes, live evidence disposes. The agent can remember, but it never acts on memory alone.

It's built from:

Qdrant Cloud for the vector data the search gateway queries
Hindsight for the agent's long-term memory
SQLite (one file) to log incidents, events, aliases, and cache state
A Python CLI that runs the loop and asks a human to approve every write

You can see the full project here:
https://github.com/Ramakrishna1967/AfterTrace

The idea

Most talk about agent memory is about the upside: the agent remembers, so it gets better. But a memory that makes an agent quicker at being right also makes it quicker at being wrong when things change.

So I gave Hindsight recall one small job: decide what to check first. A recalled fix never gets applied on its own. Before any change, the agent re-checks the live state that fix depends on.

I tested this with three scenarios, each run in a fresh process so the only thing carried over is what's stored in memory.

Scenario 1: cold start

A new corpus, B, is loaded correctly, but the alias still points at the old one, A. A query that should return B returns A.

With no memory, the agent checks that B is fine, sees the alias points at A, proposes switching it, and waits for approval. After the change it re-queries B, confirms it works, and saves the real incident to Hindsight. I only save what actually happened, never a made-up summary.

Scenario 2: transfer

A different corpus, same kind of bug. A fresh process calls recall through the Hindsight API, gets the experience from Scenario 1, and checks the alias first. That's the payoff of persistent memory: it went straight to the likeliest cause. It still checked live state before writing anything, and the query passed.

Scenario 3: the important one

Same symptom, but now the alias already points at B. The real cause is a stale cache serving old results.

Recall returns the old alias fix. The agent re-checks the live alias, sees the fix no longer applies, and prints:

REJECTED recalled alias fix

It then looks for another cause, finds the stale cache, and clears only the cache. The alias is left alone.

Without the check, Scenario 3 ends with the agent rewriting a healthy alias while the real fault stays. With it, memory decides what to look at and live evidence decides what to change.

What I learned
Use recall to prioritize, not to authorize. Ranking suspects by past incidents helps a lot. Acting on them blindly is the risk.
Re-check what a recalled fix depends on. A remembered fix assumes the world is unchanged. Test that.
Only save real incidents. Made-up memories poison future recalls.
Test the failure case first. The scenario where memory is wrong shows whether your design is safe.
Fresh processes show what memory really does. Restarting between scenarios made it clear what came from Hindsight and what was just in RAM.

What's next

Right now the checks cover alias and cache faults. Other fault types need their own checks. I also want to save rejected recalls, so the rejection itself becomes something the agent learns from.

If you're building an agent with persistent memory, Hindsight is worth trying. The biggest lesson for me was deciding exactly what memory is allowed to decide.

Top comments (1)

Collapse
 
reidmarlow profile image
Reid Marlow •

Splitting recalled memories into an assertion check versus the mutation saves a lot of wasted agent turns. When an agent recalls an action directly, it tends to rationalize why the live environment fits the old fix. Storing the broken precondition as a separate check lets the harness discard the candidate before the model even starts reasoning about applying it.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.