I Used Hindsight to Make the Second Incident Different
The first time a production incident happens, investigation is unavoidable. The second time a similar incident happens, repeating the same investigation is a design failure.
That observation shaped OpsMind. I built it around a simple idea: an incident response system should not only analyze the outage in front of it; it should remember what the organization learned after the outage and use that experience when the next one arrives.
The problem I wanted to solve
Most incident tooling is very good at telling me what is happening now.
I can see latency, error rate, CPU, memory, database connections, deployment history, logs, traces, and alerts. What is harder is answering a different question:
Have we seen this before, and what did we learn when we did?
That knowledge is usually distributed across post-mortems, tickets, runbooks, chat messages, dashboards, and the people who happened to be on call.
OpsMind treats that historical experience as part of the incident itself.
The workflow is:
Current incident
↓
Current telemetry and deployment context
↓
Hindsight Recall
↓
Compare past experience with current evidence
↓
Investigation recommendation
↓
Engineer approves the action
↓
Resolved incident + validated outcome
↓
Hindsight Retain
↓
Better context for the next incident
The important part is the last two steps. If the resolution never becomes reusable knowledge, I have only built another incident assistant.
How the system hangs together
The repository is deliberately split into a React presentation layer and a small Node/Express orchestration layer.
The frontend is responsible for the incident console: incident lists, telemetry, investigation results, memory exploration, learning events, runbooks, and post-mortems.
The backend owns the workflow.
server.ts exposes endpoints such as:
POST /api/incidents/:id/investigate
POST /api/incidents/:id/resolve
POST /api/incidents/:id/learn
GET /api/memory/search
GET /api/memory/list
POST /api/memory/test
The interesting boundary is server/hindsightService.ts. I kept Hindsight behind one service rather than spreading calls to the memory system throughout the application.
That gives the incident orchestrator a simple interface:
const recallResult = await hindsightService.recall(
recallQuery,
{
service: incident.service,
limit: 3
}
);
The rest of OpsMind does not need to know how the memory store performs retrieval.
That separation became important because incident memory has different requirements from ordinary application state. I want to be able to change memory-bank configuration, retrieval behavior, retry policy, or deployment mode without rewriting the incident workflow.
The key design decision: remember outcomes, not conversations
I deliberately did not make Hindsight a transcript archive.
An incident chat can contain hundreds of messages, many of which are speculation. Storing the entire conversation does not necessarily give a future incident responder useful experience.
Instead, OpsMind retains the things I actually want another engineer to know:
what happened
which service was affected
the symptoms
the deployment involved
suspected root cause
actions attempted
actions that failed
the action that worked
the final lesson
The type in the repository reflects that:
export interface HistoricalMemory {
id: string;
incidentId: string;
title: string;
service: string;
similarity: number;
symptoms: string[];
rootCause: string;
attemptedActions: string[];
failedActions: string[];
successfulAction: string;
resolutionTime: string;
lesson: string;
deploymentVersion: string;
environment: string;
timestamp: string;
type: 'experience' | 'observation' | 'world';
tags: string[];
}
That structure is intentional.
A future responder usually does not care that an engineer wrote 37 messages while debugging. They care that three things were tried, two failed, one worked, and the root cause was eventually confirmed.
That is the useful unit of organizational memory.
The second incident is where the architecture proves itself
I built the main workflow around two related Payments API incidents.
The first incident, INC-104, has a familiar production failure signature:
P99 latency: 4.8s
HTTP 5xx rate: 18%
DB connection pool: 96%
Deployment: v4.2.1
Affected checkout users: 31%
The investigation identifies a database connection leak associated with the deployment. Restarting instances and increasing replicas are not useful fixes; rolling back the deployment resolves the incident.
The important operation happens after resolution.
OpsMind converts the validated outcome into a memory:
const retainResult = await hindsightService.retain(
Incident ${incident.id} Resolution: ${rootCause}. +
Successful fix: ${successfulAction}. +
Lesson: ${lessonText},
{
incidentId: incident.id,
service: incident.service,
rootCause,
successfulAction,
failedActions,
deploymentVersion: incident.deploymentVersion,
lesson: lessonText,
tags: [
incident.service.toLowerCase(),
'post-mortem',
'connection-pool',
'regression'
]
}
);
This is where I think persistent memory becomes materially different from simply adding retrieval to an application.
The system is not just storing documentation. It is storing an operational outcome.
Six weeks later, INC-137 shows a similar pattern:
P99 latency: 5.1s
HTTP 5xx rate: 16%
DB connections: 94%
Deployment: v4.2.3
The deployment version is different, so I don't want the agent to mechanically repeat the previous fix.
Instead, the investigation queries organizational memory:
const recallQuery =
Find previous production incidents involving ${incident.service}, +
elevated latency, 5xx errors, database connection saturation, +
and deployment-related regressions.;
const recallResult = await hindsightService.recall(
recallQuery,
{
service: incident.service,
limit: 3
}
);
The result gives the agent a historical reference point.
Now the question changes from:
What could cause this?
to:
Which parts of the previous failure pattern match this incident, and which parts don't?
That distinction matters.
The historical incident tells us that database connection exhaustion and deployment changes are worth investigating. It also tells us that scaling replicas was previously ineffective.
But the current release is v4.2.3, not v4.2.1.
So the recommendation becomes something closer to:
Inspect connection lifecycle changes in the current deployment before immediately applying the historical rollback.
That is the behavior I wanted from memory: reuse experience without blindly copying the past.
Why I put Hindsight behind a service boundary
The Hindsight integration uses the official @vectorize-io/hindsight-client package.
Configuration stays outside the application:
HINDSIGHT_BASE_URL="https://api.hindsight.vectorize.io"
HINDSIGHT_API_KEY="your-hindsight-api-key"
HINDSIGHT_BANK_ID="opsmind-production-memory"
The memory service owns the connection and the three important operations.
Recall searches historical experience.
Retain records validated learning.
Reflect provides a place for reasoning over retrieved memory and current incident context.
The service abstraction also gives me a useful operational property: Hindsight connectivity is observable.
The application exposes memory status, including whether the bank is connected and when the last recall or retain occurred. I don't want an incident engineer wondering whether a recommendation came from live organizational memory or from a degraded path.
The underlying concepts are documented in the Hindsight documentation, and the broader idea of persistent context is described by Vectorize in its explanation of agent memory.
I kept the engineer in the loop
I also made a deliberate decision not to let the incident agent silently change production.
The agent can:
analyze current signals
retrieve historical incidents
identify similarities and differences
propose investigation steps
explain the evidence behind the recommendation
The engineer still decides what actually happens to production.
That separation is visible in the data model too. Investigation results contain a recommendation, evidence, confidence, historical matches, and a reasoning trace rather than an opaque "execute this command" result.
export interface IncidentInvestigation {
summary: string;
findings: string[];
likelyCauses: string[];
recommendedAction: string;
why: string;
evidence: string[];
historicalMatches: HistoricalMemory[];
confidence: 'High' | 'Medium' | 'Low';
confidenceNote?: string;
reasoningSteps: InvestigationReasoning[];
hindsightTrace: {
recalledCount: number;
reflected: boolean;
reflectSummary?: string;
bankId: string;
isLiveHindsight: boolean;
latencyMs: number;
};
}
For incident response, explainability is not a cosmetic feature. If an agent says "do X," I want to know which current signal caused that recommendation and which historical experience influenced it.
What the user actually sees
I did not make the chat window the center of the product.
The primary interface is an incident investigation workspace.
An engineer can see:
the current incident and severity
live-looking telemetry signals
deployment timing
investigation findings
historical memory matches
failed actions from previous incidents
the recommended next step
evidence supporting the recommendation
the resulting resolution
the learning that gets retained afterward
There is also a dedicated Memory Explorer. That matters because persistent memory should not be an invisible black box. Engineers need to inspect what the system remembers.
After a resolved incident is retained, the learning appears as an organizational memory event. The next investigation can then retrieve it.
That creates a visible loop:
Incident → Investigate → Resolve → Learn
↑ ↓
└──────── Hindsight Recall ────┘
What I learned building this
- The useful memory unit is smaller than the incident
My first instinct was to think about retaining incident history as a whole. The more useful abstraction is the validated lesson.
"What happened?" is useful.
"What happened, what did we try, what failed, what worked, and under which conditions?" is much more useful.
- Similarity is not a decision
A high similarity score is a reason to investigate a historical incident, not permission to repeat its remediation.
That is why the second incident compares deployment versions and current telemetry against the historical memory.
Memory should narrow the search space, not eliminate engineering judgment.
- Failed actions are first-class knowledge
Most systems emphasize the successful fix.
I think the failed attempts are just as important.
If an engineer already learned that restarting pods does not solve a connection leak, the next engineer should not have to rediscover that fact at 2 AM.
- Memory needs an observable lifecycle
I want to know when memory was recalled, what was retrieved, and when new knowledge was retained.
That is why OpsMind exposes Hindsight status and includes a hindsightTrace in investigation results.
Memory that cannot be inspected becomes difficult to trust.
- The second incident is a better test than the first
A first incident can demonstrate that an agent can analyze telemetry.
A second related incident tests whether the system actually learned.
That became the central engineering test for OpsMind:
If the organization has already solved this class of problem, does the next investigation behave differently?
If the answer is no, persistent memory is probably just decoration.
Where I would take it next
The current architecture gives me a foundation for connecting real observability and operational systems directly to the incident workflow.
The next step is not simply "more memory."
I would connect:
metrics and traces
deployment systems
incident-management platforms
runbooks
post-mortem repositories
service ownership metadata
change history
Then I would make the memory lifecycle stricter: validated outcomes should be retained with provenance, conditions, and confidence, while contradictory or stale memories should be surfaced rather than silently treated as truth.
That is the part I find most interesting about building incident intelligence around Hindsight.
The goal isn't to make an agent that remembers everything.
The goal is to make an agent that remembers the right operational experience, retrieves it when it matters, and helps an engineer avoid repeating mistakes the organization has already paid to learn.
That is what I wanted the second incident to prove.
Top comments (0)