My On-Call Agent Stopped Repeating the Same Mistake
At 2 AM, a payments service throws a wall of 5xx errors. An on-call engineer who joined three months ago gets paged. They have two choices: guess, or know. Most of the time, they guess, because the knowledge of "we already tried that, it didn't work" lives in someone else's head, or buried in a Slack thread from six weeks ago.
I built an agent that remembers instead of guessing.
What the system does
On-Call Memory is an incident-response assistant. You paste in a live alert, and it gives you a diagnosis, a ranked list of fixes, and, critically, a list of fixes to avoid because they already failed on a similar incident. It's built around Hindsight, a memory layer for AI agents, paired with an LLM served through Groq and a small Streamlit front end.
The architecture is simple:
Alert comes in
-> recall() queries Hindsight for similar past incidents
-> alert + recalled memories go to the LLM
-> LLM returns diagnosis, fix, and what NOT to try
-> after resolution, retain() stores the new outcome
Nothing exotic. The interesting part isn't the plumbing, it's what you choose to remember.
Storing outcomes, not just facts
Most memory demos I'd seen store what happened. Mine stores what happened and whether the fix worked. That distinction turned out to matter more than I expected.
Here's the retain function:
python
def retain_incident(inc: dict) -> None:
text = (
f"INCIDENT {inc['id']} | service: {inc['service']} | date: {inc['date']}\n"
f"Symptom: {inc['symptom']}\n"
f"Root cause: {inc['root_cause']}\n"
f"Fix attempted: {inc['fix']}\n"
f"Outcome: {inc['outcome']}"
)
client().retain(bank_id=BANK, content=text)
Every incident gets five fields, and the last one, outcome, is the whole point. Without it, an agent can tell you "someone restarted the service once." With it, an agent can tell you "restarting the service failed twice, don't do that."
Recall is just as simple:
python
def recall(query: str, limit: int = 10) -> list[str]:
resp = client().recall(bank_id=BANK, query=query)
results = getattr(resp, "results", None) or []
return [getattr(r, "text", None) or str(r) for r in results][:limit]
The system prompt does the rest of the work, instructing the model to rank fixes by their recorded outcome and to never recommend something that's already on record as failed for a similar root cause.
What it looks like in practice
I seeded the memory with ten synthetic but realistic incidents across five services, deliberately pairing some of them: a failed fix followed by the fix that actually worked. Then I ran the same alert with memory off and with memory on.
Memory off, given a payments-service alert about connection pool exhaustion, the agent said: check logs and metrics, then restart the service. Reasonable-sounding. Also wrong, because a restart had already been tried twice and failed both times.
Memory on, same alert, the agent recalled four related incidents and answered differently: don't restart, that failed on INC-101 and INC-103. Roll back the deploy or revert the config instead, those are what worked on INC-102 and INC-104. It cited the incident IDs directly, so the recommendation wasn't just plausible, it was traceable.
To make sure this wasn't a fluke tied to one scenario, I tried a completely unrelated alert: an email queue backing up due to SMTP throttling. The agent switched entirely to a different pair of incidents, correctly recommended failing over to a secondary provider instead of scaling up workers (which had made things worse before), and even reasoned that restarting an unrelated service like payments wouldn't help here. That last part surprised me. The model wasn't just pattern-matching text, it was reasoning about relevance.
The feedback loop
The part that makes this a learning system rather than a static lookup table is the feedback form. After resolving an incident, you record what you did and whether it worked, and it's retained immediately:
python
if st.button("Save to memory") and fix:
memory.retain_incident(dict(
id=f"INC-{st.session_state.count}", service=service,
date=dt.date.today().isoformat(), symptom=alert,
root_cause=root or "unknown", fix=fix, outcome=outcome))
Run the same alert again immediately afterward, and the new fix shows up in the recalled memories. No retraining, no batch job, just a write followed by an immediately-useful read.
Lessons learned
Outcome-aware memory beats fact memory. A system that remembers "this happened" is a search engine. A system that remembers "this happened, and here's what fixed it" is an advisor.
Vague input breaks retrieval. Early on, I tested the agent with an unrelated, vague sentence, not an actual alert, and it still confidently generated a payments-related answer, because those memories happened to be textually closest. The fix was a prompt guardrail: if the input doesn't describe a specific service or symptom, say so instead of forcing a match. Retrieval systems will always return something; deciding when to trust that something is a separate design problem.
A restart isn't always the wrong answer. One of my seeded incidents involved a stale signing key, where restarting the affected pod actually was the correct fix. I kept it in deliberately, so the agent couldn't just learn "never restart." It has to reason per-incident, based on the specific root cause, not a blanket rule.
Synchronous glue code needs care. Streamlit re-runs the whole script on every interaction, which doesn't play well with a cached async client. I ended up creating a fresh Hindsight client on every call rather than reusing one across re-runs, a small detail that cost more debugging time than it should have.
Small scope, deep behavior beats broad scope, shallow behavior. I kept this to one workflow (incident response), one persona (the on-call engineer), and one clear value proposition (don't repeat failed fixes). That constraint made it possible to actually demonstrate the memory improving, rather than building five shallow features.
Where this goes next
The pattern here, remembering outcomes rather than just facts, isn't specific to incident response. The same idea applies to support tickets (don't suggest a troubleshooting step that already failed for this customer), sales (don't repeat a pitch that already got rejected by this type of prospect), or maintenance logs (don't repeat a repair that didn't hold last time). Incident response was just the clearest place to prove it, because "worked" and "failed" are unambiguous and fast to verify.
If you want to look at the code, it's on github https://github.com/Girish6516/oncall-memory--1-. The docs for the memory layer are at Hindsight's documentation, and there's a good overview of the underlying idea at Vectorize's page on agent memory.
Top comments (0)