DEV Community

Pooja.b
Pooja.b

Posted on

Retaining and Recalling Incidents with Hindsight: A Backend View

The Problem: Fixes That Get Lost

An alert fires and the symptoms look familiar. Someone fixed this months ago, but the fix is in a closed ticket, a Slack thread, or one person's memory. The engineer on call then investigates from scratch.

RecallOps is our AI incident-response copilot, built to close that gap. This article covers the backend side: how an incident is stored, how history is recalled, what happens when the memory service is unavailable, and how a resolved incident is retained. I'm writing about the system our team built, so when I describe a component, I'm describing the system, not claiming I wrote every part of it.

The Backend at a Glance

The backend is FastAPI (Python) with SQLAlchemy and REST APIs. The verified local/demo path uses SQLite, and the architecture is PostgreSQL-ready. Memory goes through Hindsight, which provides retain and recall operations, with a durable database fallback behind it. The AI layer uses Groq/OpenAI-compatible structured completion, with a deterministic fallback when live credentials aren't available.

We wanted memory to be a separate layer with a clear contract, not a table of old tickets attached to a prompt. Hindsight's two operations map onto the two things we needed: store what we learned from an incident, and bring back what's relevant to a new one.

How an Incident Moves Through the System

Every incident follows the same loop:

Incident → Recall → AI Investigation → Resolve → Retain → Reflect → Better Future Investigation

  1. An incident is opened and the backend looks for similar past incidents.
  2. The copilot analyzes the current incident in five visible stages, using recalled memory as context.
  3. The engineer resolves the incident.
  4. The resolution is retained as operational memory.
  5. A Learning area looks across retained incidents for recurring patterns.

The backend's job is to keep these steps in order and keep each one honest. For example, an incident can't be retained until it has been resolved.

The Incident Record

Everything starts with the incident data. Here is a simplified, representative example of the model. It shows the shape of the code, not the exact implementation:

class Incident(Base):
    __tablename__ = "incidents"

    id = Column(String, primary_key=True)
    service = Column(String, nullable=False)
    summary = Column(Text)
    status = Column(String, default="open")
    root_cause = Column(Text, nullable=True)
    resolution = Column(Text, nullable=True)
    retained = Column(Boolean, default=False)
Enter fullscreen mode Exit fullscreen mode

Each field has a job:

  • id identifies the incident (for example INC-001 or INC-017).
  • service is where similarity matching starts, since "same service" is one of the match reasons.
  • summary describes the symptoms.
  • status starts as open and later becomes resolved.
  • root_cause and resolution are nullable because they aren't known when the incident opens. They are filled in when the engineer resolves it.
  • retained records whether the resolved incident has been stored as operational memory.

The nullable fields and the retained flag encode the lifecycle. An open incident has no root cause yet, and a resolved incident isn't memory until it is retained.

Fig 1 — RecallOps incident workspace showing INC-017 during investigation,<br>
including the recalled historical memory and current evidence.

The Recall API

When an incident is investigated, the backend exposes a recall endpoint. This is a simplified, representative example:

@router.get("/incidents/{incident_id}/recall")
def recall_memory(incident_id: str, db: Session = Depends(get_db)):
    incident = get_incident_or_404(db, incident_id)

    try:
        matches = memory.recall(incident)
        degraded = False
    except MemoryUnavailable:
        matches = db_fallback_recall(db, incident)
        degraded = True

    return {
        "matches": matches,
        "degraded": degraded,
    }
Enter fullscreen mode Exit fullscreen mode

The endpoint loads the current incident and asks the memory layer for relevant past incidents. It returns the matches plus a degraded flag.

Getting the Historical Match

In the demo data, INC-001 was a Payment API database timeout. Its root cause was connection-pool exhaustion caused by a connection leak, and the resolution was to fix the leak and increase pool capacity. It was resolved and retained.

Later, INC-017 occurs, another Payment API database timeout. Recall returns INC-001 as the top historical match at 91% similarity. The match reasons are: same service, same service family, similar symptoms, similar database behavior, and similar timing.

The response carries the reasons along with the score, so an engineer can see why a memory was returned. The memory view also shows INC-001's root cause and resolution, and why it influenced the recommendations.

Fig 2 — INC-001 historical memory with 91% similarity. Shows the match reasons,<br>
historical root cause, and resolution in the memory card

When Memory Is Unavailable

The recall snippet above already contains the fallback. We couldn't assume a remote memory service would always be reachable, so if Hindsight fails, the endpoint catches MemoryUnavailable and recalls from the system's own persisted incident data through db_fallback_recall.

The important detail is the degraded flag. The application keeps working, but the response says the fallback was used, so the UI can show a degraded state instead of silently presenting the fallback as the primary memory provider. Designing for this early forced us to decide what "degraded" should look like in the interface.

The Retain Operation

Retention closes the loop. Another simplified, representative example:

@router.post("/incidents/{incident_id}/retain")
def retain_incident(incident_id: str, db: Session = Depends(get_db)):
    incident = get_incident_or_404(db, incident_id)

    if incident.status != "resolved":
        raise HTTPException(
            400,
            "Resolve the incident before retaining it",
        )

    memory.retain(incident)
    incident.retained = True
    db.commit()

    return {"retained": True}
Enter fullscreen mode Exit fullscreen mode

The endpoint refuses to retain anything that isn't resolved, which keeps unfinished investigations out of memory. Once the incident passes that check, the backend hands it to memory.retain, sets retained = True, and commits. After the engineer resolves and retains INC-017, it becomes part of the memory that future incidents can recall.

Connecting Current Incidents to Historical Memory

The system keeps four things separate: current evidence, historical evidence, the AI recommendation, and uncertainty.

For INC-017, the current evidence is 98% connection utilization, a 14.2% timeout rate, a deployment 23 minutes earlier, and a connection timeout. The historical evidence from INC-001 is high connection utilization, the same service and error family, and a confirmed connection leak.

Keeping these apart means the AI's recommendation can be traced to specific past incidents, and each source can be reviewed and debugged on its own.

fig-3

Why the Engineer Keeps the Decision

RecallOps never changes production systems automatically. The backend provides evidence, history, recommendations, investigation paths, and uncertainty, and the engineer makes the final call.

A 91% match is worth investigating, but a similar past incident is not proof of the same root cause. The UI labels INC-001 as evidence for investigation, not confirmation of the current root cause. The workspace also offers structured investigation paths, such as checking whether the recent deployment introduced a new leak.

After retention, the Learning area looks across retained incidents. On the demo data it identified a recurring Payment API pattern across five related incidents, with provenance showing which incidents each lesson came from.

fig-4

What Was and Wasn't Verified

Verified: backend startup, API health, backend tests, SQLite persistence, memory recall, retention, learning/reflection, and the browser end-to-end workflow (Launch Demo through INC-001, INC-017, recall, resolve, retain, Learning, and demo reset). The verified demo runs on SQLite with the deterministic AI fallback.

Not live-verified: a production PostgreSQL deployment, a remote Hindsight service, and live Groq/OpenAI inference. These are configured integrations in the architecture, but we haven't tested them live, so I make no claims about them. We also haven't benchmarked the system, so there are no performance claims.

Lessons Learned

  • Define degraded behavior early. The database fallback and deterministic AI fallback made local development dependable, and they made us specify what degraded should look like.
  • Return the reasons with the result. Match reasons and the degraded flag make the API's behavior visible to the UI and to the engineer.
  • Encode the lifecycle in the data. Nullable outcome fields, a retained flag, and the resolve-before-retain check keep memory clean.
  • Keep sources separate. Mixing current and historical information in one blob makes AI output harder to trust.

The next step is validating the remote integrations (PostgreSQL, a live Hindsight service, and live LLM inference) so the configured architecture is tested too. If you're interested in agent memory, the Hindsight repository is a good place to start.

Top comments (1)