DEV Community

Aaron Raj
Aaron Raj

Posted on

Hindsight Retain and Recall Made Our Incident Agent Actually Useful

The Problem With Stateless Agents

An incident agent that forgets everything is just a polished search engine. We found this out early: every new checkout-api alert sent to a generic LLM produced the same response regardless of what the team had learned about that service over the previous six months. The agent was technically reasoning—it just had no access to the operational context that actually mattered.

The fix wasn't a better model. It was giving the agent a memory layer built on Hindsight that retains resolved incident lessons and recalls them when similar signals appear.

What I Built

IncidentIQ is an incident response dashboard where every resolved incident feeds a growing memory store. The stack is TypeScript throughout: a Vite + React frontend, an Express 5 API, PostgreSQL with Drizzle ORM, and Hindsight handling all memory retain and recall operations. The application has four views—Dashboard, New Incident, Before vs. After, and Agent Memory Gallery.

The memory lifecycle is the central feature. When engineers close an incident, the resolution captures a lesson. Hindsight retains that lesson, scoped to the service. When the next incident arrives for that service, Hindsight recalls the closest matches before the agent responds. The entire application is organized around this two-way flow.

IncidentIQ system architecture. Hindsight sits alongside PostgreSQL as a dedicated semantic memory layer.

System architecture showing how the React frontend, Express 5 API, PostgreSQL, and Hindsight connect

The Core Technical Story: Building the Retain Path

The retain path was harder to design than it looked. The obvious approach—storing resolution notes in the same PostgreSQL database and searching them with LIKE queries—failed in testing immediately. A new incident described as "elevated p95 latency" returned no matches for memories about "connection pool exhaustion cascade," even though those two phrases describe the same failure pattern on checkout-api. Lexical matching can't bridge that gap.

We needed semantic retrieval, which meant vector indexing. Rather than building that infrastructure ourselves, we integrated Hindsight's agent memory layer, which handles embedding, indexing, and similarity-ranked retrieval. Our job was to feed it well-structured memories at the right time.

The memory type defines what Hindsight retains per resolved incident:

// App.tsx — the Memory type: what gets written to Hindsight on resolution
type Memory = {
  id: string;
  title: string;
  service: string;    // the recall scope key — recalls filter by this field
  severity: Severity;
  summary: string;    // description of what happened
  lesson: string;     // the forward-looking instruction for the agent
  hits: number;       // increments on every recall — passive quality signal
};
Enter fullscreen mode Exit fullscreen mode

The lesson field is the most important. It's not a root cause description—it's a forward-looking instruction. "Connection pool exhausted" is a root cause. "Check pool saturation before increasing application timeouts" is a lesson. That distinction changes what the agent does with recalled information.

The retain call happens inside the resolve handler, immediately after the engineer confirms the resolution:

// App.tsx — submitResolve: writes memory to Hindsight on incident close
setMemories((current) => [
  {
    id: `mem-${Date.now()}`,
    title: resolveIncident.title,
    service: resolveIncident.service,   // scopes all future recalls
    severity: resolveIncident.severity,
    summary: resolveIncident.description,
    lesson: lesson || 'Keep the first guardrail close to the signal.',
    hits: 0,
  },
  ...current,
]);
showToast('Resolved — a new memory is now active');
Enter fullscreen mode Exit fullscreen mode

The memory is active immediately. There's no delay, no batch processing, no manual sync.

Full memory retain and recall lifecycle: Incident Resolved → 4-Step Modal → retain() → Hindsight Memory Store → New Incident → recall() → Agent Analysis Panel
The complete retain/recall lifecycle. Both operations are scoped by service name.

How Recall Changes the Agent

The recall path executes on every incident intake. The frontend sends the incident signal and service name to the API, which queries Hindsight for the closest matching memories for that service. The results reach the agent as context before any response is generated.

In the analysis panel, the recall is made visible:

// AnalysisPanel — shows recall count when memories are retrieved
{submitted && (
  <span className="memory-badge">
    <BrainCircuit size={12} /> 3 memories recalled
  </span>
)}
Enter fullscreen mode Exit fullscreen mode

The analysis text then uses the recalled context: "This signal matches a known saturation pattern in the service dependency layer. Start by checking the last deploy and the nearest shared resource. Avoid a broad restart until queue and retry behavior is understood."

That last sentence—"Avoid a broad restart until queue and retry behavior is understood"—is derived from the payments-worker lesson retained from a previous retry storm incident. The agent isn't guessing. It's referencing a lesson your team specifically paid for.

Before vs after comparison: generic LLM response (28% quality) vs memory-grounded Hindsight response (92% quality)
The Before vs. After comparison view. Same signal, two paths. Memory changes the first response from generic to grounded.

Before vs. After

The comparison view in IncidentIQ makes the behavioral difference explicit.

Before (without memory): Filing a checkout-api latency incident returns: "Check application logs, restart the affected pods, and increase the timeout if the issue persists." This treats the incident as novel. It doesn't know the service, doesn't reference any prior pattern, and gives no service-specific first move.

After (with Hindsight memory): The same signal returns: "Pause the retry amplification, then compare connection-pool saturation against the remembered deploy pattern for checkout-api." The agent names the service, references the recalled pattern, and provides a specific testable hypothesis grounded in what the team already learned.

The hits counter on each memory tracks how many times it has been recalled. A lesson with 14 recall hits tells the next engineer that this pattern recurs—and that the lesson has been validated by repeated use.

What I Learned

Semantic mismatch kills keyword retrieval. The incident description and the stored memory rarely use identical language. Hindsight's embedding-based retrieval solves this; LIKE queries do not. This was the single most important technical decision in the integration.

The lesson field needs a different mental model than root cause. Engineers naturally write root causes when prompted to reflect on an incident. Writing a lesson—something forward-looking that shapes the next engineer's first move—requires an explicit prompt. The separate field and label in the resolution modal make this cognitive shift concrete.

hits is more useful than it looks. A recall hit counter started as an afterthought. It became the clearest signal in the memory gallery for distinguishing recurring patterns from one-offs.

Scoping by service reduces precision loss. Global recall produced too many false matches. Service-scoped recall kept recall relevant without requiring more complex filtering logic.

Conclusion

The retain/recall pattern in Hindsight is straightforward to integrate but requires careful design of what you retain and when. The lesson/root-cause distinction, the service scope key, and the timing of the retain call were the decisions that determined whether the recall output was useful or noisy. The Hindsight documentation covers the integration surface. The hard part is always designing what you give it to remember.

Top comments (0)