DEV Community

Kasi Thanmayee Anjana
Kasi Thanmayee Anjana

Posted on

Hindsight in RecallOps: Turning Resolved Incidents into Memory

Using Hindsight to Turn Resolved Incidents Into Memory

Most teams fix an incident and then lose what they learned. The fix ends up in a closed ticket, a Slack thread, or one person's head. Months later, a similar alert fires, and the engineer on call starts from zero.

RecallOps, the AI incident-response copilot our team of five built, is designed around this problem. This article is about one part of it: how a resolved incident becomes reusable operational knowledge. I'm describing the system our team built, not claiming I wrote every piece of it.

Three Numbers, Three Meanings

The demo data has three counts that are easy to mix up:

  • 6 stored incidents: every incident record in the database.
  • 5 retained memories: incidents that have been resolved and then explicitly retained as operational memory.
  • 5 related incidents: the group behind the recurring Payment API pattern that the Learning page found.

These are different measures. A stored incident is just a record. A retained memory is a record with a resolution that was deliberately saved for future recall. A related incident is one that shares a pattern with others. The 5 related incidents are not "the other 5 out of 6", and there are not 6 related incidents.

Why Retention Is a Separate Step

Hindsight, the memory layer we use, gives us two operations: retain and recall. We kept retention as its own step instead of saving every incident automatically.

An open incident has a summary and symptoms, but no confirmed root cause and no resolution. If we recalled it later, we would be recalling a guess. So the data model encodes the lifecycle:

class Incident(Base):
    __tablename__ = "incidents"

    id = Column(String, primary_key=True)
    service = Column(String, nullable=False)
    summary = Column(Text)
    status = Column(String, default="open")
    root_cause = Column(Text, nullable=True)
    resolution = Column(Text, nullable=True)
    retained = Column(Boolean, default=False)
Enter fullscreen mode Exit fullscreen mode

(Simplified and representative, not the exact implementation.)

root_cause and resolution are nullable because they aren't known when an incident opens. The engineer fills them in when resolving it. retained records whether the resolved incident has been stored as memory. Resolved and retained are two separate facts.

The Retain Step

The retain endpoint enforces that order:

@router.post("/incidents/{incident_id}/retain")
def retain_incident(incident_id: str, db: Session = Depends(get_db)):
    incident = get_incident_or_404(db, incident_id)

    if incident.status != "resolved":
        raise HTTPException(
            400,
            "Resolve the incident before retaining it",
        )

    memory.retain(incident)
    incident.retained = True
    db.commit()

    return {"retained": True}
Enter fullscreen mode Exit fullscreen mode

(Simplified and representative.)

If the incident isn't resolved, the request is rejected. Only incidents with meaningful resolution information can become memory. That is why the counts differ: stored means "exists", retained means "worth recalling".

The Recall Step

Retention only matters if it changes what happens next time. When a new incident is investigated, the backend asks the memory layer for relevant past incidents:

@router.get("/incidents/{incident_id}/recall")
def recall_memory(incident_id: str, db: Session = Depends(get_db)):
    incident = get_incident_or_404(db, incident_id)

    try:
        matches = memory.recall(incident)
        degraded = False
    except MemoryUnavailable:
        matches = db_fallback_recall(db, incident)
        degraded = True

    return {"matches": matches, "degraded": degraded}
Enter fullscreen mode Exit fullscreen mode

(Simplified and representative.)

If Hindsight is unavailable, the endpoint falls back to the system's own persisted incident data and sets degraded to true. The app keeps working, and the UI can show that the fallback was used instead of hiding it.

The INC-001 → INC-017 Example

In the demo data, INC-001 was a Payment API database timeout. Its root cause was connection-pool exhaustion caused by a connection leak. The resolution was to fix the leak and increase pool capacity. It was resolved and retained.

Later, INC-017 appears: another Payment API database timeout. Recall returns INC-001 as the top match at 91% similarity. The match reasons are:

  • same service
  • same service family
  • similar symptoms
  • similar database behavior
  • similar timing

The reasons come back with the score, so the engineer can see why the memory was returned. The current evidence for INC-017 is 98% connection utilization, a 14.2% timeout rate, a deployment 23 minutes earlier, and a connection timeout. We keep that separate from the historical evidence from INC-001: high connection utilization, the same service and error family, and a confirmed connection leak.

Where Reflection Comes In

Retaining single incidents gives us recall. Looking across retained incidents gives us something more: patterns. After retention, the Learning page reviews the retained incidents together. In the demo data it found a recurring Payment API pattern across five related incidents. Each lesson shows provenance, meaning which incidents it came from, so an engineer can check the claim against the source instead of trusting a summary.

This is the difference between a searchable archive and operational knowledge. An archive answers "has this happened before?" Reflection answers "does this keep happening, and why?" One timeout can be bad luck. Five related ones point to something structural, which is worth a different kind of fix.

Memory Supports the Engineer

RecallOps never changes production systems automatically. It provides evidence, history, recommendations, investigation paths, and uncertainty. The engineer decides.

A 91% match is a good lead, not proof. The UI labels INC-001 as evidence for investigation, not confirmation of the current root cause. The workspace also suggests structured investigation paths, such as checking whether the recent deployment introduced a new leak. The memory shortens the search, but the engineer still confirms the cause.

What Was and Wasn't Verified

We verified backend startup, API health, backend tests, SQLite persistence, recall, retention, learning/reflection, and the browser workflow from Launch Demo through INC-001, INC-017, recall, resolve, retain, Learning, and demo reset. That demo runs on SQLite with the deterministic AI fallback.

We have not live-verified a remote Hindsight service, a production PostgreSQL deployment, or live Groq/OpenAI inference. They are configured integrations in the architecture, but I make no claims about how they behave live. We also haven't benchmarked the system, so there are no performance claims.

What I Learned

  • Resolution gates memory. Requiring resolved before retain keeps unfinished investigations out.
  • Count things precisely. Stored, retained, and related are different numbers, and mixing them makes the system look more capable than it is.
  • Explain every match. Reasons and the degraded flag make the system's behavior visible.
  • Patterns need accumulation. Reflection only becomes useful once several incidents have been retained.

Lesson: operational memory is not saved history. It is resolved knowledge, retained on purpose, recalled with reasons, and checked by an engineer.

Top comments (0)