DEV Community

Cover image for The night my agent warned me not to restart the pod
P Ajith kumar
P Ajith kumar

Posted on AI-assisted

The night my agent warned me not to restart the pod

What I built

I built an on-call assistant for a payments platform. It takes an alert, like payments-api is timing out and returning 504 errors, and gives an engineer a diagnosis: likely root cause, which past incident it resembles, and an ordered list of fixes. Nothing unusual so far. Every AI vendor ships some version of "paste your logs, get an answer."

The part that makes it worth writing about is what happens after the fix. When an engineer tries something and it works, or doesn't, they tell the agent. That outcome gets stored, and the next time a similar alert fires, the agent's answer is different because of it. It isn't fine-tuned on this feedback. It's retrieved.

That retrieval layer is Hindsight, and it's the reason this project works at all. Underneath, the agent is a thin layer: recall relevant history, hand it to an LLM alongside the new alert, and let the model reason over both. The interesting engineering isn't the LLM call. It's deciding what "relevant history" means and making sure the model doesn't quietly ignore it.

The story: connection pools, refund handlers, and a reflex that doesn't work

Here's the incident that made the project click for me. payments-api had a history of 504 errors, and every time, the instinct was the same: restart the service. Sometimes that worked. Once, it didn't, because the root cause wasn't a stuck process, it was a connection leak in the refund handler that quietly ate every connection in the pool. A restart bought a few minutes of headroom and then the same failure mode came right back.

I wanted the agent to remember that distinction specifically, not just "payments-api had an outage," but "payments-api had an outage where the obvious fix was wrong." That's a much stronger signal than a generic postmortem summary, and it's the whole reason I don't treat "with memory" and "without memory" as marketing language in this project. They're two different prompts, run side by side, and the difference between them is the entire pitch.

Here's the function that draws the line:

def diagnose(alert: str, use_memory: bool = True) -> str:
    past_text = ""
    if use_memory:
        with memory_client() as m:
            past = m.recall(bank_id=BANK, query=alert)
            past_text = "\n".join(f"- {r.text}" for r in past.results[:12])

    prompt = f"""You are an on-call incident response assistant.

New alert: {alert}

Past incidents from memory:
{past_text or "None available."}

Rules:
- Line 1: most likely root cause.
- Line 2: which past incident it matches (or "no match").
- Then 2-3 fix steps, ordered most likely to work first.
- If a past incident says a fix FAILED or DID NOT WORK, do NOT recommend
  it as a step. Put it on a final line starting with "Avoid:".
- If a past fix WORKED, prefer it and say so.
- Use ONLY facts from the past incidents above. Never invent statistics."""

    return ask_llm(prompt)
Enter fullscreen mode Exit fullscreen mode

The use_memory flag is the whole experiment. Flip it off, and the model gets nothing but the alert text, so it falls back to textbook troubleshooting: check downstream health, look at logs, maybe scale out. Flip it on, and Hindsight's recall pulls in the closest matching incidents by meaning, not just keyword overlap, so "gateway timeouts" correctly surfaces a memory that says "504 errors."

The rule that matters most in that prompt is the Avoid: instruction. Retrieval alone doesn't prevent a model from suggesting a fix that already failed, because a similarity search will happily return a failed attempt right alongside a successful one. It's the prompt's job to tell the model what to do with a memory that says "restart did not help." Get that wrong, and you've built an agent that quotes its own postmortems back to you without learning anything from them.

What the two answers actually look like

For the alert above, here's the model with no history:

Likely cause: upstream dependency or resource exhaustion causing the API to exceed its timeout.
No previous fixes to reference.

  1. Check health and latency of downstream services.
  2. Restart the payments-api pods to clear stuck connections.

Reasonable. Also exactly the reflex that doesn't work here.

With memory:

Root cause: Postgres connection pool exhaustion caused by a connection leak in the refund handler.
Matches: payments-api 504 incident (connection leak, fixed by increasing max_connections and patching the handler).

  1. Increase max_connections and patch the refund handler to release connections properly.
  2. Enable monitoring on connection pool usage. Avoid: Simply restarting the pods, as that did not resolve the timeout in a prior incident.

Same model, same alert, same timeout budget. The only input that changed is twelve lines of retrieved history. That's the entire argument for treating memory as infrastructure rather than a feature checkbox: it changes what the model is willing to recommend, not just how it phrases the recommendation.

Closing the loop is the part nobody demos

The recall side gets all the attention because it's visually satisfying, two columns, one obviously better than the other. The retain side is less glamorous and more important, because it's the only mechanism that keeps the "with memory" column from going stale.


def record_outcome(alert: str, fix_tried: str, worked: bool):
    if worked:
        text = (f"For the alert '{alert}', the fix '{fix_tried}' WORKED and resolved the problem. "
                f"Recommend this fix for similar alerts.")
    else:
        text = (f"For the alert '{alert}', the fix '{fix_tried}' FAILED and did NOT resolve the problem. "
                f"Do not recommend '{fix_tried}' for similar alerts.")
    with memory_client() as m:
        m.retain(bank_id=BANK, content=text, context="fix outcome feedback")
Enter fullscreen mode Exit fullscreen mode

I tested this on an alert that had never appeared anywhere in the system: a scheduler timezone bug that double-charged customers. With no history, the agent guessed at a plausible cause and suggested rolling back the last deployment, a reasonable guess, but not the actual cause. I told it the real fix, marked it as successful, and waited about a minute for Vectorize's agent memory to finish processing the new fact. On the next similar alert, phrased differently, the agent cited the exact fix I'd just taught it and recommended it directly. Nobody touched the prompt or the model in between. The only thing that changed was what it could recall.

What I'd do differently

Retrieval will hand the model a bad match, and the model will sometimes use it anyway. I added an explicit "use only facts from the memories above, never invent statistics" rule after watching the model claim a confidence number that existed nowhere in my data. Constrain the output at the prompt level; don't trust the model to police itself.

Memories can lose meaning when they're split apart. Hindsight extracts individual facts out of longer incident write-ups, which is usually a strength, it's how a single postmortem becomes several retrievable, dated facts. But a poorly worded input sentence can occasionally get split in a way that drops the outcome ("worked" vs. "failed") from the fact that survives. Writing retain calls with the outcome stated explicitly and redundantly fixed this for me, and it's a habit I'd recommend to anyone storing structured feedback rather than prose.

"No match" is a feature, not a failure state. The tempting bug is a model that gets handed weak, tangentially related memories and decides they must be relevant because they're the closest thing available. I had to explicitly instruct it to say "no match" and fall back to general advice when nothing genuinely applies, otherwise it forces connections between unrelated incidents just because a similarity search always returns something.

Two-column comparisons are a legitimately good way to evaluate a memory system. Running the same prompt with and without the retrieved context, side by side, made every regression and improvement immediately visible in a way that a single output never would have. If you're building anything on top of Hindsight or a similar memory layer, keep the no-memory path around even after you're confident the memory path works. It's the only honest baseline you have.

The value isn't the LLM call. It's what you choose to keep. Anyone can wire an alert to a model and get a diagnosis. The part that actually took thought was deciding what counts as a fact worth retaining, how to phrase it so a later retrieval doesn't lose the nuance, and how to keep a fix that failed from ever becoming a fix the model suggests again. That's the whole job, and it's the part that doesn't show up in a demo unless you go looking for it.

Top comments (0)