The first time our agent fixed a bug, it took 7 investigation steps and 11 LLM calls, and a customer got a wrong answer first. By the third bug of the same kind, it caught the problem before any customer asked, in 2 steps.
Nobody wrote a detection rule for that. The agent wrote it itself, stored it in Hindsight, and ran it on its own. This post covers how that works, what our A/B test showed, and where it fell short.
Context: an agent that debugs another agent's memory
We built Memory SRE, an agent that investigates why a memory-backed support agent gave a wrong answer. A rep flags a bad answer. An investigator agent searches memory with 10 tools, finds the memory the wrong claim came from, classifies the failure (identity split, stale fact, recall miss, and so on), applies a reversible fix through Hindsight's curation API, then re-asks every affected question to verify the fix.
The failure we saw most often is an identity split: the CRM calls a customer "Saffron Retail", billing calls the same company "Mehta Brothers Trading LLP", and tag-scoped recall never connects the two. So the agent confidently answers from half the facts.
Fixing one split is useful. But every fix taught the investigator something, and then that knowledge disappeared. The next split cost exactly as much as the first. That's what an agent without real agent memory looks like: it's smart at every step and learns nothing between them.
Two banks: one for customers, one for lessons
The design choice that made everything else possible was giving the agent a second Hindsight bank just for what it learns. The support agent's memories live in the main bank. Everything Memory SRE learns lives in a lessons bank, as three kinds of memory:
- Playbooks (procedural memory): the symptom, the tool path that worked, the root cause, the fix and the verified outcome.
- Detection rules that the agent writes itself, in general terms.
- Exceptions: patterns a human reviewer rejected.
Tags keep them separate (sre:playbook, sre:rule, sre:exception), and Hindsight's consolidation turns repeated lessons into observations with proof counts, so a rule confirmed three times carries more weight than a rule seen once.
The post-mortem: where a fix becomes a rule
After every verified fix (never an unverified one), a post-mortem runs. The LLM gets the full incident: question, wrong answer, correction, root cause, the exact tool path, and the evidence that linked the two identities. It returns a symptom signature and, if the incident generalizes, a rule. Both get retained:
rule = pm.get("rule") if isinstance(pm.get("rule"), dict) else {}
signal = str(rule.get("signal_type") or "none").strip().lower()
if signal in RULE_SIGNALS:
_retain(f"Detection rule rule-{iid} ({rule.get('name') or ftype}), learned from {iid}. Signal type: {signal}. "
f"How to check: {rule.get('how_to_check') or '-'} Confirmation: {rule.get('confirmation') or '-'}",
f"rule-{iid}", ["sre:rule", f"rule:{signal}", f"type:{ftype.lower()}"], "Memory SRE detection rule")
The key constraint is RULE_SIGNALS. The model can write any rule it wants in prose, but the rule has to name a signal type we can actually execute: an email domain, email address, phone number or account id. After the Saffron incident, the agent wrote roughly this: "If two customer records share a company email domain (not a free-mail provider), check whether they are the same customer." That's the general lesson behind "Saffron and Mehta Brothers both use mehtabros.co.in".
Patrol and Watch run only what was learned
Two autonomy loops use those rules:
- Patrol runs every learned rule across every customer record in the bank.
- Watch runs after every live answer, as cheap text checks on what was just recalled, and calls the LLM only if a rule fires.
This is the part I'm proudest of: neither loop has a hard-coded detection strategy. Their checks come only from rules in the lessons bank:
def learned_rules() -> list[dict]:
"""Detection rules the post-mortems wrote, from Hindsight: [{id, signal_type, text}]."""
if not config.MEMORY_ENABLED:
return []
rules = []
for d in _docs("sre:rule"):
signal = next((t.split(":", 1)[1] for t in d["tags"] if t.startswith("rule:")), None)
if signal in RULE_SIGNALS:
rules.append({"id": d["id"], "signal_type": signal, "text": d["text"]})
return rules
With an empty lessons bank, Patrol runs zero checks. Everything it catches, it catches because of something the agent learned from its own past incidents. When a rule fires, the investigator still checks the candidate pair with compare_records before anything is linked, and an autonomy policy decides whether the fix applies automatically (confidence ≥ 0.85, reversible) or waits for approval.
When the human says no
Rules over-generalize. In our data, two unrelated companies, Pinecrest Hospitals and Vanadium Energy, both have tickets from nimbusit.in, because they use the same managed-IT vendor. The email-domain rule flags them. The agent dismisses the pair, and a reviewer can click Reject pattern so it never comes up again. That rejection is also memory:
_retain(f"Rejected identity proposal: '{catalog.customer_name(a)}' and '{catalog.customer_name(b)}' are NOT the "
f"same customer. The shared {signal} '{value}' is not linking evidence. Reviewer's reason: {reason}.",
doc_id, ["sre:exception", f"signal:{signal}", f"value:{value}", f"pair:{pair_key(a, b)}"],
"Memory SRE reviewer feedback")
Patrol skips excepted values, and the investigator sees them in its case file. Feedback changes behavior instead of disappearing into a log.
All of this is also summarized in a Hindsight mental model, the "Memory SRE Playbook", which refreshes after every post-mortem. The investigator can read it with its get_playbook tool, and the UI shows its version history, so you can read how the agent's understanding changed over time. The Hindsight docs cover mental models well. They're an underrated feature.
The A/B test: memory OFF vs ON
Claims like "it learns" are cheap, so we measured it. Our benchmark runs the same scenario twice, in separate processes with separate Hindsight banks: four identity splits (Kestrel, Saffron, Monsoon, Neelgiri), in a fixed order. In each episode a customer asks a question that touches the split. One arm has the lessons bank disabled.
| Memory OFF | Memory ON | |
|---|---|---|
| Wrong answers reaching customers | 4 | 2 |
| Investigation steps | 27 | 18 |
| LLM calls | 43 | 34 |
| Wall time | 827 s | 958 s |
| False identity links | 0 | 0 |
With memory OFF, the agent diagnosed all four splits correctly, but every split reached a customer first and each cost about 7 steps. With memory ON, episodes 3 and 4 had no wrong answer at all: Patrol had already linked Monsoon and Neelgiri, at 2 steps each.
The honest part
The ON arm didn't go smoothly. In episode 1, the investigator settled for a verified but shallow fix. It called the failure MISSING_KNOWLEDGE and retained the correction instead of linking the identities. The answer was fixed, but the post-mortem had no general signal to learn from, so no rule was written.
In episode 2 (Saffron), it correctly diagnosed an identity split and wrote the email-domain rule. The Patrol that followed ran 12 rule checks, found 3 candidates and prevented all 3, including going back and properly linking Kestrel, repairing the root cause that episode 1 had missed. I didn't expect that, and it's the most convincing moment of the whole project: learned memory fixed a mistake the agent made before it learned.
We closed the loophole behind episode 1 afterwards: if a search shows another record sharing a signal with the customer, the agent must now compare the two before concluding a fact is missing. And wall time went up with memory, because post-mortems, patrols and playbook refreshes aren't free. What memory buys is fewer wrong answers and cheaper later cases, not speed.
Lessons learned
1. Store lessons separately from the data they're about. A dedicated lessons bank with clear tags made playbooks, rules and exceptions easy to recall, list and audit.
2. Let the model write rules, but only in a vocabulary you can execute. Free-text lessons are nice to read. Rules bound to executable signal types are what actually change behavior.
3. Only learn from verified outcomes. A post-mortem that runs only after re-ask verification means the agent never learns from a fix that didn't work.
4. Negative feedback is memory too. Storing "this pattern is not evidence" stopped the same false positive from coming back.
5. Measure against a no-memory baseline. Without the A/B arms, I would've credited the model for improvements that came from memory, or missed that episode 1 learned nothing.
The code, benchmark script and raw results are at github.com/079Sathya/hindsight-agent. If you're building on Hindsight, try giving your agent a second bank for what it learns. It's a small change that makes a big difference.

Top comments (0)