description: "A suggested fix is evidence for an investigation, not proof the incident was resolved. How I built an incident-memory agent where only confirmed resolutions reach durable memory."
tags: ai, python, programming, webdev
The easiest way to poison an incident memory system is to write every plausible answer back into it. I built this one around a stricter rule: a suggested fix is evidence for an investigation, not evidence that the incident was resolved.
That sounds obvious until the first useful answer arrives.
A model sees db_pool=exhausted, recalls a similar outage, and proposes retry jitter and a larger connection pool. The proposal may be excellent — or exactly wrong for this deploy. Retain it as verified, and the next engineer inherits a confident echo of an untested guess.
The problem: memory that learns from its own guesses
An agent with memory has an obvious superpower: it can recall a past outage instead of starting from scratch. It has an equally obvious failure mode: it treats its own output as evidence.
Most "agent memory" integrations have a single write path. Somewhere in the code there is a saveMemory() call, and it fires on whatever the agent just produced. Do that in an incident tool and the system slowly launders guesses into institutional memory. Next quarter, a fragment of a hallucinated runbook shows up as precedent.
The design decision that shaped this system was to make those states impossible to confuse:
- analysis does not write memory
- a confirmed resolution becomes a durable incident record
- feedback is its own kind of evidence
Hindsight is the memory layer that makes those records useful later, but it does not get to decide what counts as operational truth. We do.
The approach: one job, three writes
The application is an incident-resolution service: a React console over a FastAPI backend. An engineer pastes a title, impact, severity, and a log excerpt. The backend formats that into a query, asks Hindsight for related operational history, then hands the current incident and the recalled records to an Agno-managed model. The reply is two or three proposed resolutions, each with a confidence value and references to the records behind it.
Hindsight handles server-side embedding and fact extraction on retain and recall, so the service runs no separate embedding pipeline. The larger benefit is architectural: the memory bank is a durable record of incidents and decisions, not an opaque transcript. The Hindsight documentation covers the retention primitives, and Vectorize's overview of agent memory is a good framing for why a system needs more than a model's context window.
Here is the whole flow:
flowchart LR
A[Engineer and React console] -->|incident report| B[FastAPI backend]
B -->|incident plus recalled records| C[Agno managed model]
C -->|resolutions and references| B
B -->|recall evidence| M[(Hindsight memory bank)]
B -.->|write resolved only| M
The API deliberately starts with a non-writing operation:
@app.post("/api/incidents/new", response_model=AnalyzeResponse)
def create_incident(incident: IncidentNew, top_k: int = Query(default=3, ge=1, le=10)):
try:
return analyze_incident(incident, top_k=top_k)
except RuntimeError as exc:
raise HTTPException(status_code=500, detail=str(exc))
POST /api/incidents/new returns a short-lived incident ID, suggestions, and the recalled records. It does not persist the incident merely because somebody asked a question. That boundary prevents the system from learning from every half-formed report and every generated answer.
The other two writes are explicit. POST /api/incidents/resolve receives a real resolution after the incident is fixed. POST /api/incidents/feedback receives a helpful or not-helpful verdict for a particular suggestion. The frontend keeps the original incident payload with the response so it can resend it with the verdict; the server needs no hidden session.
Confirmation is a write boundary
The only route that creates a confirmed-fix record calls persist_resolution. It stores enough context to judge whether the precedent applies: title, severity, description, log excerpt, and the resolution.
def persist_resolution(client, bank_id, incident_id, incident, resolution):
content = format_incident_for_memory(incident_id, incident, resolution)
client.retain(
bank_id=bank_id,
content=content,
document_id=f"incident-{incident_id}",
metadata={"status": "resolved", "severity": incident.severity.value},
)
return f"incident-{incident_id}"
The document ID is not cosmetic. incident-{incident_id} makes a second submission of the same confirmation an upsert instead of a duplicate memory. During incident response, retries are normal; idempotency beats an after-the-fact deduplication heuristic.
I also chose a document-oriented archive. Semantic recall surfaces fact fragments, but the incident explorer fetches the original documents — when a suggestion cites a past outage, the engineer reads the full context, not a compacted fact. In operational tooling, the full record is the audit trail.
Hindsight earns its place here: a recall interface over retained operational records, without forcing a local vector database, a chunking worker, and another store for extracted facts into the application. The system still owns the schema of trust, though — status=resolved means something different from everything else in the bank.
Feedback is evidence, not a resolution
Feedback has a different failure mode. It is tempting to record a thumbs-up as proof that a suggestion worked. But a person may have found an answer useful because it narrowed the search, described a symptom accurately, or led elsewhere. A thumbs-down is even less like a resolution, but still valuable: this framing or candidate fix was not useful here.
So feedback is retained under a different document namespace and metadata shape:
content = (
f"User feedback on incident {incident_id}: {verdict} | "
f"Rated suggestion index: {suggestion_index} | "
f"{format_incident_for_query(incident)}"
)
client.retain(
bank_id=bank_id,
content=content,
document_id=f"feedback-{incident_id}",
metadata={"status": "feedback", "helpful": str(helpful).lower()},
)
Notice what is absent: we never promote the generated prose into a confirmed runbook step. We retain the incident context, the verdict, and which ranked suggestion was assessed — a durable positive or counterexample that preserves the difference between "this was helpful" and "this fixed production."
That keeps the incident lifecycle legible: a new incident triggers recall and proposals but writes nothing; feedback retains a verdict about a suggestion; resolution retains the confirmed fix.
I prefer this over a single saveMemory() path with an ever-growing pile of optional fields. Separate operations make the caller state its intent, make authorization easier to reason about, and give us a place to add retention or retrieval policy by record type — without inferring truth from free-form text later.
Retrieval needs an evidence budget, not a magic number
One implementation detail was surprisingly easy to get wrong. Hindsight recall does not expose a native top_k parameter in this integration. It uses a recall budget and a token ceiling to decide how much context to return; the service then limits the records it passes onward to the model and UI.
resp = client.recall(
bank_id=bank_id,
query=query_text,
budget=budget,
max_tokens=max_tokens,
include_chunks=True,
)
results = getattr(resp, "results", []) or []
for record in results[:max(top_k, 0)]:
similar.append(SimilarIncident(
memory_id=getattr(record, "id", "") or "",
text=getattr(record, "text", "") or "",
document_id=getattr(record, "document_id", None),
))
It is not just a parameter-name quirk: recall breadth and prompt cardinality are different controls. The first sets how much evidence Hindsight returns; the second constrains what reaches the generator and the engineer. Conflating them makes latency and answer quality harder to tune.
The prompt makes the evidentiary requirement explicit: every proposed resolution must carry references to the recalled records. If nothing is recalled, the instructions require low confidence rather than pretending generic advice came from local history. The model may synthesize an answer; it may not erase the difference between precedent and conjecture.
What it looks like in practice
An engineer reports: "Checkout requests return 504s after a deploy; payments-api has 200 active database connections and retries are synchronized." The service turns the title, description, severity, and logs into one Hindsight query. A previously retained record about connection-pool saturation and missing retry jitter is a plausible match.
The response may contain a candidate such as: enable jittered backoff, inspect the pool limit, and scale only after confirming connection saturation. Crucially, it also names the recalled record that supports the advice — the console shows that relationship, not just a confidence number.
The deployed console is live at incident-agent-o8vf.vercel.app.
If the suggestion was useful, the engineer records that verdict; if the remediation is validated, someone calls /resolve. Two inspectable records result: feedback-<id> and incident-<id>. A wrong idea leaves a negative verdict as a counterexample — it never masquerades as a fix.
I am not claiming this guarantees faster resolution. That is an outcome to measure in production, not a property conferred by a retrieval call. What the system guarantees by construction is provenance: a generated proposal, a user verdict, and a confirmed resolution occupy different states.
Pitfalls I walked into
- The better the answer, the more it feels savable. The first genuinely useful suggestion is the most dangerous one, because it is the one you want to persist. A useful answer is not a verified answer.
- Persisting on the analysis call is the default mistake. Engineers investigate noisy, incomplete reports. A write on every query turns that noise into durable context and makes later recall less trustworthy.
- A thumbs-up is not a resolution. Recording it as one silently upgrades opinion into fact. Keep the verdict in its own namespace.
-
top_kis not a recall budget. Treating the application-level slice as the retrieval control hides how much evidence the memory layer actually considered. - Hidden session state makes feedback unrecoverable. Sending the original incident payload back with the verdict is simpler to retry and inspect than reconstructing a conversation.
Takeaways
- Retention needs a trust model before it needs clever retrieval. If the write path does not distinguish proposal from confirmation, bad guidance will eventually acquire the appearance of history.
- Make the first analysis non-persistent. Recall freely; write deliberately.
- Store the original incident with the verdict. A stateless feedback endpoint is easier to recover and retry.
- Treat citations as part of the response contract. A reference is the path an on-call engineer uses to decide whether the precedent applies.
- Tune recall and prompting separately. Hindsight's recall budget sets evidence breadth; the application-level slice sets how much reaches the model and the engineer. Configure and observe them as separate levers.
Conclusion
I did not build this to turn incident response into a conversation with a model. I built it to make prior operational experience easier to retrieve without laundering guesses into institutional memory.
Hindsight supplies the durable recall layer; explicit write boundaries keep that memory worth trusting.
Any system that learns from its own outputs has to answer one question first: what am I allowed to believe? In an incident tool the honest answer is narrow. Recall everything worth reading. Write back only what a human has confirmed. And keep every proposal, verdict, and resolution in a state where the next engineer can see exactly which one they are looking at.
Built with Hindsight for agent memory, FastAPI, React, and Agno.



Top comments (0)