DEV Community

Aakash Sai Ram
Aakash Sai Ram

Posted on

How IncidentIQ Puts Hindsight at the Center of Incident Response

When the Architecture Is the Product

I've worked on systems where memory was an afterthought—a localStorage key, a session variable, a note appended to a ticket. In every case, the agent's usefulness capped out at the boundaries of a single conversation. The moment someone new opened a session, the slate was clean. IncidentIQ is an attempt to build something different: a system where the agent's memory is a first-class architectural component, not an afterthought, and where the entire application is organized around how information moves from past incidents into future ones.

This is the story of how we designed that system.

What I Built

IncidentIQ is an incident response command center backed by a pnpm monorepo: a Vite + React frontend, an Express 5 API, PostgreSQL with Drizzle ORM for persistence, and Hindsight as the memory layer. The application has four views: a Dashboard showing live incident state and agent activity; a New Incident form where engineers file signals and the agent immediately surfaces relevant historical patterns; a Before vs. After comparison screen showing what changes when memory is present; and an Agent Memory gallery where the full memory store is browsable and searchable.

Every part of the application connects to the memory layer either on the write path (when incidents close) or the read path (when new incidents arrive). That symmetry was intentional from the start.
System architecture: Engineer → React frontend → Express 5 API → PostgreSQL/Drizzle, with Hindsight connected to the API via retain() and recall()
IncidentIQ system architecture. Hindsight handles memory independently of the primary database.

The Core Architectural Decision

The critical design question was where the memory boundary should sit. Options ranged from storing resolution notes in the same PostgreSQL database as incidents (simple, cheap, no semantic search), to passing the entire incident history as context in every LLM prompt (expensive, context-limited), to using a dedicated memory layer that handles indexing, retrieval, and scoping independently of the application database.

We chose the third approach, using Hindsight's agent memory model. The reasons were concrete: keyword search over structured notes fails as soon as the wording differs between the original incident and the new one. A checkout-api latency incident described as "elevated p95 response time" should match a memory titled "connection pool exhaustion cascade"—not because those phrases overlap lexically, but because the semantic context is the same. That kind of retrieval requires vector-indexed storage, which Hindsight provides without us building the infrastructure.

The application database (PostgreSQL via Drizzle) holds structured incident records. Hindsight holds the semantic memory store. They're separate on purpose. Incidents are structured data with known fields—id, service, severity, status, timestamps. Memories are semantic artifacts designed for fuzzy retrieval. Conflating them in the same store would either require building retrieval on top of a relational schema or imposing relational rigidity on semantic content.
The memory retain/recall lifecycle: from incident resolution through the 4-step modal to Hindsight's memory store, and back via recall on new incident intake
Memory lifecycle in IncidentIQ. Retain happens at resolution; recall happens at intake. Both are scoped by service name.

How Information Flows Through the System

The write path is driven by the resolution modal. When an engineer closes an incident, they walk through four steps: root cause, resolution action, lesson learned, and a confirmation to retain that lesson as agent memory. The lesson field is the critical one—it's what reaches Hindsight.

// App.tsx — memory shape: what Hindsight retains per resolved incident
type Memory = {
  id: string;
  title: string;
  service: string;    // scopes future recalls to this service
  severity: Severity;
  summary: string;    // root cause description
  lesson: string;     // forward-looking instruction for the agent
  hits: number;       // how many times Hindsight has recalled this memory
};
Enter fullscreen mode Exit fullscreen mode

The service field is the scope key. Every retain call includes it. Every recall call filters by it. A memory from payments-worker does not surface when the next incident is checkout-api, even if the symptoms overlap.

The read path executes on every incident intake. Before the agent generates analysis, the system runs a recall against the memory store scoped to the service of the incoming incident. The matching memories surface in the analysis panel:

// AnalysisPanel — shows recall badge when memories are retrieved
{submitted && (
  <span className="memory-badge">
    <BrainCircuit size={12} /> 3 memories recalled
  </span>
)}
Enter fullscreen mode Exit fullscreen mode

The analysis text then references the recalled context directly: "This signal matches a known saturation pattern in the service dependency layer. Start by checking the last deploy and the nearest shared resource."
Agent Memory gallery showing searchable memory cards with service pills, severity dots, lesson sections, and recall hit counts
The Agent Memory gallery. Engineers browse and search the full memory store. Recall hit counts show which patterns recur most.

Before vs. After

The comparison view makes the behavioral difference explicit. You give it an incident signal for a specific service and run both paths side by side.

Without memory: "Check application logs, restart the affected pods, and increase the timeout if the issue persists." Generic, non-actionable, treats the incident as novel.

With memory: "Pause the retry amplification, then compare connection-pool saturation against the remembered deploy pattern for checkout-api." Specific, service-grounded, immediately testable.

The comparison view quantifies this as a quality score—28% without memory, 92% with. Those numbers reflect how closely the response anchors to a testable first hypothesis. The specific values come from the comparison UI's scoring logic built into the application.

The Resolution Workflow Drives Everything

The 4-step modal is the engine of the memory layer.

// ResolveModal — four-step memory capture on incident close
const [step, setStep] = useState(1); // 1=RootCause, 2=Resolution, 3=Lesson, 4=Retain

const submitResolve = () => {
  if (step < 4) { setStep(step + 1); return; }
  onSubmit(rootCause, resolution, lesson);
};
Enter fullscreen mode Exit fullscreen mode

Each step is gated. You cannot advance without filling in the current field. The friction is intentional—incomplete lessons produce low-quality recalls. The final step ("Retain Memory") confirms that the lesson will be active for the next incident on that service.
The 4-step resolution workflow: Root Cause → Resolution → Lesson Learned → Retain Memory, feeding into Hindsight's service-scoped memory store

The four-step resolution modal drives all memory writes. The lesson field—separate from root cause—is what reaches Hindsight.

What I Learned

Architecture is memory design. The decision about where memory lives—in the application database or in a dedicated semantic store—determined everything downstream. Once we committed to Hindsight as a first-class layer, the data flow became clearer: structured data in PostgreSQL, semantic memory in Hindsight, no crossover.

Scope every retain call at the most specific reliable key. Service name was the right scope for this system. Broader scoping (e.g., all incidents globally) produced too many false recall matches. Narrower scoping (e.g., by specific endpoint) produced too few. Service name sits at the right granularity.

The four-step modal has a real cost. Engineers pushed back on the extra steps when first encountering the resolution flow. The payback takes a few weeks of incidents before the recalled lessons become noticeably useful. That lag is a genuine UX tradeoff, not a solved problem.

Separate the write schema from the read schema. The PostgreSQL incident schema and the Hindsight memory schema serve different retrieval patterns. Trying to unify them would have produced a system good at neither structured queries nor semantic retrieval.

Conclusion

The architecture of IncidentIQ is an argument: when agent memory is treated as a first-class system component—with its own storage, its own retrieval semantics, and its own write path—the agent becomes genuinely useful rather than generically capable. The Hindsight documentation covers the retain/recall mechanics in detail. The key architectural insight is simpler: information only compounds if the system is designed to collect it.

Top comments (0)