Building RETRACE, an on-call agent that recalls the past incident that looks like this one.
The first time I asked my incident agent for help, it gave me a textbook troubleshooting checklist: confirm the symptom, scope the impact, check service health, look at the database layer. Useful, generic, and exactly what you could get from any LLM with no context. The second time, after I'd fed it a handful of past incidents, it said something different: the pattern matched earlier checkout-api failures where a traffic spike exhausted the database connection pool, and it pointed back to the fix that had already worked.
That second answer is the whole point of this project. I call it RETRACE. It's an on-call incident agent that doesn't just reason about the problem in front of it. It recalls the specific past incident that looks like this one and reuses what a human already figured out.

The RETRACE UI running locally on 127.0.0.1:8000. Left: a box for the new incident. Right: the agent's response. Top right: the count of demo incidents in memory.
What it does and how it hangs together
RETRACE is small on purpose. There are three pieces: a FastAPI backend, Hindsight as the memory layer, and Groq running the reasoning step. The flow is:
Browser (templates + static)
│ fetch /api/analyze
▼
FastAPI (app.py)
│ hs.recall(bank_id, query) ──► Hindsight (memory)
│ llm.chat.completions.create(...) ──► Groq (reasoning)
▼
Response: historical match, what happened, previous fix, next checks

The same flow as a diagram. The web UI sends incident reports to the backend. The backend queries Hindsight for similar past cases, passes that context to the Groq LLM, and returns grounded recommendations to the UI.
When a new incident comes in, the backend doesn't just hand it to an LLM and hope. It first calls Hindsight's recall against a bank of past incidents, and only then builds a prompt that includes whatever memories came back. The LLM's job isn't to invent a diagnosis. It's to explain why a past incident matches and reuse the root cause and fix already on record. If nothing relevant comes back, it says so instead of guessing.
The project itself is a tiny folder: app.py for the server, data/incidents.json for the demo history, templates/ and static/ for the page, a .env file for keys, and run.ps1 to start everything. After checking that the JSON file exists and parses, I start it with python -m uvicorn app:app --reload.

The project in VS Code. The terminal validates data/incidents.json and then launches the server with uvicorn's --reload flag.
That distinction between retrieval-grounded reasoning and open-ended reasoning is the core technical story here, and it took the most iteration to get right.
## Making the LLM stick to what it's told
My first version of the prompt just said "here are some past incidents, help with this new one." It mostly worked, but it also occasionally made things up: plausible-sounding root causes that weren't grounded in anything I'd retained. An agent that recalls an incident correctly but then reasons past it into fiction is worse than one that admits it doesn't know.
The fix was to be explicit about the boundary. Here's the system prompt for the "warm" (memory-enabled) path:
WARM = ("You are RETRACE, an on-call incident agent with organizational memory. You get a NEW incident and "
"MEMORIES of past incidents. If a memory matches (same service, error signature or failure pattern), cite its "
"incident id, explain WHY it matches, and reuse its recorded root cause and fix. Never invent incidents that are "
"not in the memories. If nothing matches, set match_found=false and say so. " + SCHEMA)
"Never invent incidents that are not in the memories" is doing real work in that sentence. It's the difference between a system that's citing organizational history and one that's impersonating it. I also keep a separate "cold" prompt with no memory access, so the UI can run the same incident both ways. That's the purpose of the two buttons, Run cold and Run with Hindsight.
I also learned not to trust a single model to always return clean JSON. Structured-output failures are common enough with open-weight models under load that I built in a fallback:
MODELS = [os.getenv("GROQ_MODEL", "openai/gpt-oss-120b"), "qwen/qwen3-32b"]
def ask_llm(system: str, user: str) -> dict:
last = None
for model in MODELS:
try:
r = llm.chat.completions.create(model=model, temperature=0.2, ...)
text = re.sub(r".?", "", text, flags=re.S)
return json.loads(re.search(r"{.}", text, re.S).group(0))
except Exception as e:
last = e
raise HTTPException(502, f"LLM failed on all models: {last}")
This is unglamorous code, but it matters more than the clever parts. And because the memory lives in Hindsight rather than in a model's context window, switching models doesn't touch the memory at all.
## Where memory actually lives
It would have been easy to fake "memory" by stuffing recent conversation history into the prompt. That's not what this is. Every past incident is retained into Hindsight as a self-contained memory: service, symptom, root cause, fix, and resolution time in one record:
def to_memory_text(i: dict) -> str:
return (f"Incident {i['id']} on {i['date']} in {i['service']} (handled by {i['engineer']}). "
f"Symptom: {i['symptom']}. Root cause: {i['root_cause']}. "
f"Fix: {i['fix']}. Resolved in {i['minutes']} minutes.")
The demo dataset has five incidents across five very different failure modes, and the UI shows them under "Retained incident history":

The retained history: two checkout-api incidents (502s after a traffic spike, connections at the maximum), an auth-service logout problem after a deploy, a catalog-api latency jump after an import, and a delayed notification-worker queue. Each card shows severity and time to resolve. Below them, the four-step loop: Retain, Recall, Ground, Learn.
Recall works on meaning, not keyword overlap. A new incident described in slightly different words than the original still surfaces the right memory, because Hindsight matches on semantic similarity rather than exact strings. That's what makes this feel less like a search bar and more like something that understands your incident history.
The four-step strip at the bottom of the page is the whole philosophy in miniature: Retain every resolved incident, Recall relevant past failures, Ground the LLM's answer in them, and Learn, meaning the next resolved incident becomes memory too.
## Results: the same incident, run both ways
I typed one incident into the box: "checkout-api is returning 502s after a traffic spike. Database connection are at the maximum and retries are increasing." Then I ran it twice.
Without memory: The agent opened with a "cold-start notice" saying it had no prior context and was giving a generic first-pass checklist. What followed was sensible but generic: confirm the symptom with curl or health checks, scope the impact, check pod CPU/memory/threads, inspect the load balancer, then move on to the database layer.

Cold run. The status badge reads "Cold run complete" and the panel says no incident history was used.
With memory: Hindsight recalled 31 relevant memories, and the response came back in four clear sections: the historical match, what happened before, the previous fix, and recommended next checks. It identified that earlier incidents were also traffic spikes that exhausted the database connection pool and caused 500/502 errors and rising retries. The recorded fix was to increase the connection-pool size, correct connection-handling code so connections are released, and add monitoring of pool usage. It also noted that a service restart combined with fixing the release logic had cleared a similar issue on a different API.

Run with Hindsight. The header reads "Hindsight recalled 31 relevant memories," and the answer is organized around what actually happened last time.
The gap between those two answers is the entire pitch. Neither is wrong, but only one is grounded in what actually happened, and that's the difference between rediscovering a fix and just applying it.
## What still needs work
Looking at the warm response honestly, two things stand out. First, it refers to "Memory 1" and "Memory 2" rather than citing incident IDs like INC-1098, even though the prompt asks it to cite the ID. The recalled text is what the model sees, so the fix is probably to make the ID more prominent in what I pass in. Second, "31 memories" is more than the five incidents in the dataset, which suggests Hindsight stores more than one record per retained incident, or that older test runs are still in the bank. I want to understand which before I trust that number in a demo.
## Lessons learned
Retrieval-then-reason beats reason-then-hope. Once I forced the LLM to work from retrieved memories instead of open-ended context, both quality and trustworthiness went up. It's a less flexible architecture, but flexibility wasn't what I needed.
Telling the model not to invent things is not optional. I assumed grounding would be implicit if I handed over the right memories. It wasn't. The instruction has to be direct and repeated in the schema.
Decoupling memory from the model is worth the extra abstraction. Hindsight owns the memory and the LLM is just a reasoning layer, so I can change models or fail over during an outage without touching a single retained incident.
Specific data beats generic data. Early tests with vague incidents ("service down, users affected") gave mushy matches. Once I wrote incidents with real service names, realistic error strings, and concrete fixes, recall quality jumped.
A "no memory" mode is a feature. Keeping the cold path in the app, not just on a slide, means I can show a teammate exactly what memory adds, on any incident, at any time.
If you're building something similar, the Hindsight GitHub repo is the place to start. The retain and recall primitives are simple enough that the interesting work is in how you structure what you retain and how disciplined you are about keeping the reasoning grounded in it.
## Further reading
Hindsight documentation
What is agent memory? (Vectorize)
Top comments (0)