The Problem: Fixes That Get Lost
An alert fires and the symptoms look familiar. Someone fixed this months ago, but the fix is in a closed ticket, a Slack thread, or one person's memory. The engineer on call then investigates from scratch.
RecallOps is our AI incident-response copilot, built to close that gap. This article covers the backend side: how an incident is stored, how history is recalled, what happens when the memory service is unavailable, and how a resolved incident is retained. I'm writing about the system our team built, so when I describe a component, I'm describing the system, not claiming I wrote every part of it.
The Backend at a Glance
The backend is FastAPI (Python) with SQLAlchemy and REST APIs. The verified local/demo path uses SQLite, and the architecture is PostgreSQL-ready. Memory goes through Hindsight, which provides retain and recall operations, with a durable database fallback behind it. The AI layer uses Groq/OpenAI-compatible structured completion, with a deterministic fallback when live credentials aren't available.
We wanted memory to be a separate layer with a clear contract, not a table of old tickets attached to a prompt. Hindsight's two operations map onto the two things we needed: store what we learned from an incident, and bring back what's relevant to a new one.
How an Incident Moves Through the System
Every incident follows the same loop:
Incident → Recall → AI Investigation → Resolve → Retain → Reflect → Better Future Investigation
- An incident is opened and the backend looks for similar past incidents.
- The copilot analyzes the current incident in five visible stages, using recalled memory as context.
- The engineer resolves the incident.
- The resolution is retained as operational memory.
- A Learning area looks across retained incidents for recurring patterns.
The backend's job is to keep these steps in order and keep each one honest. For example, an incident can't be retained until it has been resolved.
The Incident Record
Everything starts with the incident data. Here is a simplified, representative example of the model. It shows the shape of the code, not the exact implementation:
class Incident(Base):
__tablename__ = "incidents"
id = Column(String, primary_key=True)
service = Column(String, nullable=False)
summary = Column(Text)
status = Column(String, default="open")
root_cause = Column(Text, nullable=True)
resolution = Column(Text, nullable=True)
retained = Column(Boolean, default=False)
Each field has a job:
-
ididentifies the incident (for example INC-001 or INC-017). -
serviceis where similarity matching starts, since "same service" is one of the match reasons. -
summarydescribes the symptoms. -
statusstarts asopenand later becomesresolved. -
root_causeandresolutionare nullable because they aren't known when the incident opens. They are filled in when the engineer resolves it. -
retainedrecords whether the resolved incident has been stored as operational memory.
The nullable fields and the retained flag encode the lifecycle. An open incident has no root cause yet, and a resolved incident isn't memory until it is retained.
The Recall API
When an incident is investigated, the backend exposes a recall endpoint. This is a simplified, representative example:
@router.get("/incidents/{incident_id}/recall")
def recall_memory(incident_id: str, db: Session = Depends(get_db)):
incident = get_incident_or_404(db, incident_id)
try:
matches = memory.recall(incident)
degraded = False
except MemoryUnavailable:
matches = db_fallback_recall(db, incident)
degraded = True
return {
"matches": matches,
"degraded": degraded,
}
The endpoint loads the current incident and asks the memory layer for relevant past incidents. It returns the matches plus a degraded flag.
Getting the Historical Match
In the demo data, INC-001 was a Payment API database timeout. Its root cause was connection-pool exhaustion caused by a connection leak, and the resolution was to fix the leak and increase pool capacity. It was resolved and retained.
Later, INC-017 occurs, another Payment API database timeout. Recall returns INC-001 as the top historical match at 91% similarity. The match reasons are: same service, same service family, similar symptoms, similar database behavior, and similar timing.
The response carries the reasons along with the score, so an engineer can see why a memory was returned. The memory view also shows INC-001's root cause and resolution, and why it influenced the recommendations.
When Memory Is Unavailable
The recall snippet above already contains the fallback. We couldn't assume a remote memory service would always be reachable, so if Hindsight fails, the endpoint catches MemoryUnavailable and recalls from the system's own persisted incident data through db_fallback_recall.
The important detail is the degraded flag. The application keeps working, but the response says the fallback was used, so the UI can show a degraded state instead of silently presenting the fallback as the primary memory provider. Designing for this early forced us to decide what "degraded" should look like in the interface.
The Retain Operation
Retention closes the loop. Another simplified, representative example:
@router.post("/incidents/{incident_id}/retain")
def retain_incident(incident_id: str, db: Session = Depends(get_db)):
incident = get_incident_or_404(db, incident_id)
if incident.status != "resolved":
raise HTTPException(
400,
"Resolve the incident before retaining it",
)
memory.retain(incident)
incident.retained = True
db.commit()
return {"retained": True}
The endpoint refuses to retain anything that isn't resolved, which keeps unfinished investigations out of memory. Once the incident passes that check, the backend hands it to memory.retain, sets retained = True, and commits. After the engineer resolves and retains INC-017, it becomes part of the memory that future incidents can recall.
Connecting Current Incidents to Historical Memory
The system keeps four things separate: current evidence, historical evidence, the AI recommendation, and uncertainty.
For INC-017, the current evidence is 98% connection utilization, a 14.2% timeout rate, a deployment 23 minutes earlier, and a connection timeout. The historical evidence from INC-001 is high connection utilization, the same service and error family, and a confirmed connection leak.
Keeping these apart means the AI's recommendation can be traced to specific past incidents, and each source can be reviewed and debugged on its own.
Why the Engineer Keeps the Decision
RecallOps never changes production systems automatically. The backend provides evidence, history, recommendations, investigation paths, and uncertainty, and the engineer makes the final call.
A 91% match is worth investigating, but a similar past incident is not proof of the same root cause. The UI labels INC-001 as evidence for investigation, not confirmation of the current root cause. The workspace also offers structured investigation paths, such as checking whether the recent deployment introduced a new leak.
After retention, the Learning area looks across retained incidents. On the demo data it identified a recurring Payment API pattern across five related incidents, with provenance showing which incidents each lesson came from.
What Was and Wasn't Verified
Verified: backend startup, API health, backend tests, SQLite persistence, memory recall, retention, learning/reflection, and the browser end-to-end workflow (Launch Demo through INC-001, INC-017, recall, resolve, retain, Learning, and demo reset). The verified demo runs on SQLite with the deterministic AI fallback.
Not live-verified: a production PostgreSQL deployment, a remote Hindsight service, and live Groq/OpenAI inference. These are configured integrations in the architecture, but we haven't tested them live, so I make no claims about them. We also haven't benchmarked the system, so there are no performance claims.
Lessons Learned
- Define degraded behavior early. The database fallback and deterministic AI fallback made local development dependable, and they made us specify what degraded should look like.
-
Return the reasons with the result. Match reasons and the
degradedflag make the API's behavior visible to the UI and to the engineer. -
Encode the lifecycle in the data. Nullable outcome fields, a
retainedflag, and the resolve-before-retain check keep memory clean. - Keep sources separate. Mixing current and historical information in one blob makes AI output harder to trust.
The next step is validating the remote integrations (PostgreSQL, a live Hindsight service, and live LLM inference) so the configured architecture is tested too. If you're interested in agent memory, the Hindsight repository is a good place to start.




Top comments (1)