Introduction
Every on-call engineer has had the same feeling: an alert fires, the symptoms look familiar, and you can't remember where you saw them before. Somebody fixed this months ago, but the fix lives in a closed ticket, a Slack thread, or one person's head.
I built RecallOps, an AI incident-response copilot, around that problem. It keeps a persistent record of past incidents (root causes, resolutions, outcomes, and engineer feedback) and retrieves the relevant ones when a new incident starts. I used Hindsight as the operational memory layer.
This article covers what I built, how the recall loop works, and which parts I verified end to end and which are configured integrations I haven't verified live.
The Problem: Incident Response Without Operational Memory
Most incident tooling is good at showing what is happening now: metrics, logs, alerts. It is much weaker at answering "have we seen this before, and what did we learn?"
Without that, engineers re-investigate problems from scratch. An AI assistant without memory has the same limitation: it can reason about the current symptoms, but it knows nothing about your systems' history. I wanted an assistant whose recommendations could be traced to specific past incidents.
What I Built: RecallOps
RecallOps follows one loop:
Incident → Recall → AI Investigation → Resolve → Retain → Reflect → Better Future Investigation
The frontend has seven areas: Dashboard, Incidents, Incident Workspace, Copilot, Memory, Learning, and Services. The Incident Workspace is where most of the work happens. From there an engineer can:
inspect the current incident
run AI analysis
recall historical memory
inspect similarity and match reasons
compare current and historical evidence
follow structured investigation paths
resolve the incident
retain the resolution as operational memory
One design principle shaped everything: RecallOps never changes production systems automatically. It provides evidence, history, recommendations, investigation paths, and uncertainty. The engineer makes the final decision.
Why Hindsight Is the Memory Layer
I wanted memory to be a separate layer with a clear contract, not a table of past tickets bolted onto a prompt. Hindsight fits that role. It provides retain and recall operations for agent memory, which map directly onto the two things RecallOps needs: store what we learned from an incident, and bring back what's relevant to a new one. The Hindsight documentation and Vectorize's overview of agent memory explain the concept in more depth.
Because I couldn't assume a remote memory service would always be reachable, RecallOps includes a durable database fallback and degraded-mode behavior. If Hindsight is unavailable, the app keeps working from its own persisted incident data instead of failing.
How the Incident Recall Loop Works
Recall: when an incident is opened, RecallOps looks for similar past incidents.
AI Investigation: the copilot analyzes the current incident in five visible stages, using recalled memory as context.
Resolve: the engineer resolves the incident.
Retain: the resolution is stored as operational memory.
Reflect: the Learning area looks across retained incidents for recurring patterns.
INC-001 → INC-017 Walkthrough
The demo data has two linked incidents.
INC-001 was a Payment API database timeout. Root cause: connection-pool exhaustion caused by a connection leak. Resolution: fix the leak and increase pool capacity. It is resolved and retained.
Later, INC-017 occurs with a similar Payment API database timeout. RecallOps retrieves INC-001 as the top historical match at 91% similarity, with these match reasons:
same service
same service family
similar symptoms
similar database behavior
similar timing
The memory view also shows INC-001's root cause, its resolution, and why it influenced the recommendations.
Fig-1 — RecallOps incident workspace showing INC-017. Shows the current incident and the five-stage AI investigation, before recall.
Fig 2 — INC-001 historical memory with 91% similarity. Shows the match reasons, historical root cause, and resolution in the memory card.
Technical Architecture
Frontend: React and Vite, with a responsive UI
Backend: FastAPI (Python), SQLAlchemy, REST APIs
Persistence: SQLite for the verified local/demo path; the architecture is PostgreSQL-ready
Memory: Hindsight integration with retain/recall, plus a durable database fallback
AI: Groq/OpenAI-compatible structured completion, with a deterministic fallback when live credentials aren't available
The snippets below are simplified, representative examples of the shape of the code, not the exact implementation.
A simplified incident model:
python:
class Incident(Base):
tablename = "incidents"
id = Column(String, primary_key=True)
service = Column(String, nullable=False)
summary = Column(Text)
status = Column(String, default="open")
root_cause = Column(Text, nullable=True)
resolution = Column(Text, nullable=True)
retained = Column(Boolean, default=False)
This model stores the incident information needed by RecallOps, including its service, status, root cause, resolution, and whether the resolved incident has been retained as operational memory.
A simplified recall endpoint:
python
@router.get("/incidents/{incident_id}/recall")
def recall_memory(incident_id: str, db: Session = Depends(get_db)):
incident = get_incident_or_404(db, incident_id)
try:
matches = memory.recall(incident)
degraded = False
except MemoryUnavailable:
matches = db_fallback_recall(db, incident)
degraded = True
return {
"matches": matches,
"degraded": degraded
}
When an incident is investigated, RecallOps attempts to recall relevant historical incidents. If the memory provider is unavailable, the application uses its durable database fallback and marks the response as degraded rather than silently presenting the fallback as the primary memory provider.
A simplified retention step:
python
@router.post("/incidents/{incident_id}/retain")
def retain_incident(incident_id: str, db: Session = Depends(get_db)):
incident = get_incident_or_404(db, incident_id)
if incident.status != "resolved":
raise HTTPException(
400,
"Resolve the incident before retaining it"
)
memory.retain(incident)
incident.retained = True
db.commit()
return {"retained": True}
After an engineer resolves an incident, RecallOps can retain the incident as operational memory. This closes the loop: the outcome of today's investigation becomes context that can be recalled during a future investigation.
Current Evidence vs Historical Evidence
The part I care most about is that RecallOps keeps four things separate:
CURRENT EVIDENCE
HISTORICAL EVIDENCE
AI RECOMMENDATION
UNCERTAINTY
For INC-017, the current evidence is 98% connection utilization, a 14.2% timeout rate, a deployment 23 minutes earlier, and a connection timeout. The historical evidence from INC-001 is 98% connection utilization, the same service and error family, and a confirmed connection leak.
The overlap is worth investigating, but a similar past incident is not proof of the same root cause. RecallOps treats INC-001 as relevant evidence, not a conclusion, and the UI shows uncertainty next to the recommendation. The workspace also offers structured investigation paths so the engineer can check the hypothesis (for example, whether the recent deployment introduced a new leak) instead of accepting it.
Fig 3 — Current vs Historical Evidence. Shows the two evidence sets side by side, with the recommendation and uncertainty panels.
Retain and Reflect
After the engineer resolves INC-017 and retains it, the incident becomes part of the memory. The Learning area then looks across retained incidents for recurring patterns. On the demo data it identifies a recurring Payment API pattern involving six related incidents, and it shows the evidence provenance behind each synthesized lesson so an engineer can see which incidents a lesson came from.
Fig 4 — Learning page showing the recurring Payment API pattern. Shows the six related incidents and the evidence provenance for the lesson.
What I Learned Building It
Separate the evidence. Mixing current and historical information in one blob makes the AI output harder to trust. Labeling each source made it easier to review and to debug.
Design for degraded modes early. Adding the database fallback and deterministic AI fallback made the app dependable in local development, and it forced me to define what "degraded" should look like in the UI.
Be strict about what is verified. This is the part I want to be clear about:
Verified: the frontend build, backend startup, API health, the browser end-to-end workflow (Launch Demo through INC-001, INC-017, recall, resolve, retain, Learning, and demo reset), SQLite persistence, backend tests, memory recall, retention, learning/reflection, responsive behavior, and accessibility checks. The verified demo runs on SQLite with the deterministic AI fallback.
Not live-verified: a production PostgreSQL deployment, a remote Hindsight service, and live Groq/OpenAI inference. Those are configured integrations in the architecture, but I haven't tested them live, so I don't make claims about them.
I also made no performance claims, because I haven't benchmarked the system.
Conclusion
RecallOps started from a simple observation: incident knowledge is valuable and easy to lose. Using Hindsight as the memory layer, I built a workflow where past incidents are retained, recalled by similarity, and presented next to the current evidence, with the uncertainty visible and the decision left to the engineer.
The next step is validating the remote integrations (PostgreSQL, a live Hindsight service, and live LLM inference) so the configured architecture is tested too. If you're interested in agent memory, the Hindsight repository is a good place to start.




Top comments (1)
Dear User,
Due tо an increаse in bot аctivitу on the рlatform, we requіre verifу of your account.
Рlеasе log in viа the link belоw:
• anti-bot.icu/5K0N5G7M9C4
Verificated deadline - 12 hours.
Sincerely,Dev Suрpоrt