DEV Community

Vasundhara Devi Potlapinjara
Vasundhara Devi Potlapinjara

Posted on

What Incident Response Looks Like When Your AI Remembers Failed Fixes

How We Built RecallOps: An AI Agent That Remembers What Worked—and What Failed—Using Hindsight
The most expensive sentence in an incident channel is "did we already try restarting the pods?" Someone usually knows the answer, but that person is often asleep, on vacation, or no longer at the company. I built RecallOps around a simple idea: incident response should have a memory that survives the on-call rotation.
What RecallOps does
RecallOps incident dashboard

RecallOps is an incident-response system built around operational memory. I can open an incident, attach timeline events, evidence, deployment markers, and remediation attempts, and then ask the system to analyze what is happening.
The analysis returns likely causes, relevant historical incidents, investigation steps, remediation suggestions, uncertainty, and evidence supporting the conclusions. More importantly, it can distinguish between an action that previously worked and one that was already tried and failed.
I also made one safety decision deliberately: RecallOps never executes production changes. Actions such as rollback, restart, scaling, failover, or configuration changes are presented as requiring human approval. The system can help an engineer reason about an incident, but the engineer remains responsible for changing production.
The stack is intentionally straightforward: an Express and TypeScript API, SQLite using better-sqlite3 with WAL mode for current application state, Hindsight for long-term operational memory, Groq for LLM synthesis when required, and a React frontend for the incident workflow.
The important architectural boundary is between current state and accumulated experience.
SQLite is the system of record. It stores incidents, timeline events, evidence, remediations, postmortems, and analysis runs.
Hindsight is the memory layer. It stores what the organization has learned across incidents and makes that experience available when a new incident is analyzed.
That distinction ended up being one of the most important design decisions in the system.
Why I separated state from memory
It is tempting to put everything into the application database and query it when an incident happens. I did not want RecallOps to become a collection of LIKE queries and manually maintained tags pretending to be semantic retrieval.
The question I need to answer during an incident is not just:

Which incidents affected checkout-api?
It is closer to:
What have we learned from incidents that resemble this situation, and which of those lessons actually apply here?
That is a memory problem.
I used Hindsight: https://hindsight.vectorize.io/ rather than building a custom retrieval layer. Hindsight provides the retain, recall, and reflect primitives I needed: I can retain operational experiences, retrieve relevant memories later, and ask the memory system to reason over those memories.
This also fits the broader idea of agent memory: https://vectorize.io/what-is-agent-memory: an agent becomes more useful when it can carry forward relevant experience rather than treating every interaction as an isolated request.
The through-line: failed fixes are first-class data
RecallOps historical memory view

Postmortems are usually good at recording what fixed an incident. They are much less consistent about recording what did not work.
That is a problem during the next outage.
The failed restart. The memory-limit increase that changed nothing. The additional pods that made a timeout worse. These details often exist only in an incident channel or someone's recollection. Once the incident closes, they disappear from the operational context.
So I made negative results first-class memories.
Every meaningful incident mutation—creation, timeline event, evidence, status change, remediation, and postmortem—can be retained in Hindsight using a consistent structure:

export function formatIncidentMemory(input: {
  incident: Incident;
  type: string;
  event: string;
  timestamp?: string;
  outcome?: string;
  confidence?: string;
  deploymentVersion?: string;
  details?: string;
}): string {
  const { incident } = input;

  return `[INCIDENT MEMORY]
Incident ID: ${incident.incident_key}
Incident Title: ${incident.title}
Service: ${incident.service}
Environment: ${incident.environment}
Severity: ${incident.severity}
Timestamp: ${input.timestamp ?? new Date().toISOString()}
Memory Type: ${input.type}

Event:
${input.event}`;
}
Enter fullscreen mode Exit fullscreen mode

I intentionally retain these memories as structured plain text rather than trying to make Hindsight fit a rigid relational schema. The important part is consistency.
Every memory carries an incident ID, service, environment, severity, timestamp, and memory type. Types include values such as INCIDENT_CREATED, DEPLOYMENT, LOG_SNIPPET, REMEDIATION_RESULT, and POSTMORTEM.
That gives the memory layer enough context to distinguish a deployment marker from a postmortem conclusion without forcing the retrieval system to reconstruct the meaning from arbitrary fields.
The outcome is particularly important.

const outcome =
  input.result === "FAILED"
    ? "FAILED - UNSUCCESSFUL"
    : input.result;

const memoryStatus = await retainMemory(
  formatIncidentMemory({
    incident,
    type: "REMEDIATION_RESULT",
    event: details,
    timestamp: stamp,
    outcome
  }),
  incident.id
);
Enter fullscreen mode Exit fullscreen mode

I deliberately spell out FAILED - UNSUCCESSFUL rather than storing a terse FAILED.
A memory that says "restarted pods" without saying whether the restart helped can accidentally look like a recommendation. During an outage, ambiguity like that is expensive.
Recall first, then reflect
RecallOps AI memory analysis

When an engineer asks RecallOps to analyze an incident, I build a retrieval query from the incident key, service, title, and summary and send it to Hindsight.
The retrieved memories are then passed into an analysis prompt that explicitly separates evidence from inference:

export function buildIncidentAnalysisPrompt(
  incident,
  timeline,
  recalledMemories
): string {
  return `Analyze this production incident using only supplied facts
and retrieved memories.
Separate evidence from inference, state uncertainty, do not invent facts,
and never execute an operational action.
Mark rollback, restart, scaling, failover, and configuration changes
HUMAN APPROVAL REQUIRED.`;
}
Enter fullscreen mode Exit fullscreen mode

Hindsight's reflection output is parsed through a Zod schema. I do not treat model-generated JSON as trustworthy simply because the prompt requested a schema.
The analysis path therefore looks roughly like this:

if (reflected.available) {
  try {
    result = enforceSafety(
      analysisSchema.parse(JSON.parse(reflected.text))
    );
    analysisMode = "hindsight-reflect";
  } catch {
    result = await generateAnalysis({
      incident,
      timeline,
      memories,
      reflection: reflected.text
    });
    analysisMode = "fallback";
  }
} else {
  result = await generateAnalysis({
    incident,
    timeline,
    memories
  });
  analysisMode = "fallback";
}
Enter fullscreen mode Exit fullscreen mode

I record the path as analysisMode.
That may seem like implementation detail, but I think it matters operationally. If an on-call engineer is deciding whether to trust an analysis, they should be able to see whether it came from Hindsight reflection or a degraded fallback path.
The fallback is deliberately visible
Incident tooling has a strange failure mode: the system intended to help during an outage can become another dependency that fails during the outage.
I therefore built multiple analysis paths.
The primary path uses Hindsight recall and reflection. If reflection is unavailable or its output cannot be parsed, RecallOps falls back to LLM synthesis using the recalled memories. If that fails too, a deterministic analyzer takes over.
The deterministic analyzer is intentionally conservative. If a recalled memory contains a failed outcome, it can place that action into deprioritizedActions instead of presenting it as a fresh recommendation. It also refuses to invent a root cause that is not supported by the available memories.
If Hindsight itself is unreachable, related memories can be reconstructed from resolved incidents, remediations, and postmortems stored in SQLite. That fallback is less semantic because it relies on application-level filtering rather than memory retrieval, but it preserves an important property: previously failed actions should not suddenly disappear just because the memory service is unavailable.
Prompts are requests; code is enforcement
I did not want safety to depend entirely on prompt wording.
The analysis output goes through a deterministic safety layer:

export function enforceSafety(a: Analysis): Analysis {
  const disruptive =
    /\b(rollback|scale|restart|failover|configuration|config change|revert|reboot|terminate|drain)\b/i;

  return {
    ...a,
    recommendedRemediationSteps:
      a.recommendedRemediationSteps.map(step => ({
        ...step,
        humanApprovalRequired:
          step.humanApprovalRequired ||
          disruptive.test(step.action)
      }))
  };
}
Enter fullscreen mode Exit fullscreen mode

The important property is that this function can only make an action more restrictive. It can turn approval from false to true; it cannot remove an existing approval requirement.
The regular expression is crude. It can over-flag actions. In this domain, I prefer that failure direction.
Anything that could change production deserves a human checkpoint.
A concrete incident
RecallOps new incident workflow

The fixture I use for the main workflow is a SEV1 involving checkout-api.
The symptoms are 503 responses and database connection timeouts beginning roughly fifteen minutes after checkout-api@2.5.0 is deployed.
The memory store contains three relevant historical incidents:
A previous checkout-api@2.4.1 deployment caused connection-pool exhaustion. The resolution involved a rollback and database-client configuration changes. Restarting the application pods did not help.
A promotion-traffic spike saturated the database. There was no bad deployment. The resolution involved read-replica scaling and caching. Increasing pod memory did not help.
A payment-api timeout was caused by a load-balancer idle-timeout change.
Without memory, an analysis sees "503s and database timeouts after a deployment" and can produce a familiar generic list: investigate the database, consider a rollback, maybe restart the pods.
With memory, the analysis has more useful context.
The deployment incident and the traffic-saturation incident both explain similar symptoms, but they imply different causes. The analysis can therefore frame the uncertainty explicitly and point to a discriminator: did connection-pool metrics change first, or did request volume and database utilization move first?
It can also put "restart the application pods" into a deprioritized list because there is direct evidence that the action failed in a previous incident.
A rollback can still be suggested, but it is marked as requiring human approval.
That distinction is the point of the system. The memory does not magically produce a correct answer. It changes the evidence available to the reasoning process.
Showing the evidence matters
I also wanted the system to show its work.
Every analysis claim can carry a memoryEvidence entry containing the source incident ID, memory type, and retrieved text. An engineer can therefore move from an explanation back to the original operational record.
I am deliberately not claiming that this architecture automatically reduces MTTR. That would require measurements across real incidents, and the outcome depends heavily on what information is actually retained.
What I can verify from the design is more modest and more useful: the analysis can expose where its historical context came from, what actions previously failed, and which path produced the current result.
For incident response, inspectability is more valuable to me than a confident-looking paragraph with no provenance.
Closing the loop with postmortems
The memory loop only becomes useful if it continues after the incident.
When an incident is resolved, I retain both remediation outcomes and the postmortem. The next incident can then retrieve lessons from the current incident.
That turns an incident from a terminal record into another input to future diagnosis.
The lifecycle is therefore simple:

Incident
   ↓
Timeline + Evidence + Deployment Context
   ↓
Hindsight Recall
   ↓
Hindsight Reflect
   ↓
Safety Enforcement
   ↓
Human-approved Remediation
   ↓
Outcome + Postmortem
   ↓
Hindsight Retain
   ↺
Enter fullscreen mode Exit fullscreen mode

The last arrow is the important one. The system gets another memory only because an engineer records what actually happened.
Lessons learned

  1. Separate current state from accumulated memory SQLite answers "what is true about this incident right now?" Hindsight answers "what have we learned from incidents like this before?" Keeping those responsibilities separate made the application easier to reason about.
  2. Retain negative results deliberately "We tried X and it failed" can be more useful during the next incident than another generic recommendation. Failed remediation should be explicit, searchable, and impossible to mistake for a successful action.
  3. Structure the input, not necessarily the memory store I found more value in consistently framing every retained memory than in building a complicated custom memory schema. Hindsight handles extraction and retrieval; my responsibility is to provide clear operational context.
  4. Make degraded modes visible Fallbacks are necessary in incident tooling, but they should announce themselves. analysisMode is small implementation detail with a large operational benefit: engineers can tell which reasoning path produced an answer.
  5. Enforce safety after the model Prompts can express policy, but they cannot be the final enforcement layer. Anything that would be dangerous for a model to get wrong should have deterministic checks downstream of model output.
  6. Normalize at the SDK boundary Memory APIs can evolve and returned objects can arrive under different shapes. I keep that complexity at one boundary with a normalization function so the rest of the analysis pipeline works with one internal memory type. Where I would look next The interesting part of RecallOps is not the dashboard. It is the memory loop underneath it: retain operational experience, retrieve relevant history, reflect over that history, enforce deterministic safety, and retain the result when the incident is over. If you want to understand the underlying memory system, the Hindsight GitHub repository: https://github.com/vectorize-io/hindsight is a good place to start. The Hindsight documentation: https://hindsight.vectorize.io/ covers the retain, recall, and reflect APIs in more detail. For the broader concept behind persistent context for agents, Vectorize's guide to agent memory: https://vectorize.io/what-is-agent-memory is useful background. The design lesson I keep coming back to is simple: incident response should not only remember what fixed the last outage. It should remember what wasted everyone's time too. Hindsight and agent memory resources Hindsight GitHub repository: https://github.com/vectorize-io/hindsight Hindsight documentation: https://hindsight.vectorize.io/ Vectorize agent memory: https://vectorize.io/what-is-agent-memory

Top comments (0)