DEV Community

Cover image for HINDSIGHT INCIDENT COPILOT
Chaitra Beechu
Chaitra Beechu

Posted on

HINDSIGHT INCIDENT COPILOT

## Building an Incident Copilot on Hindsight
What is wired, what is simulated, and what we learned
Hindsight Incident Copilot | AI, DevOps, Python, Open source

An alert at 2 a.m.

PaymentService is returning 502s after a deploy.
Someone on the team has seen this before, but they are asleep, and the answer is buried in a post-mortem from March. The first time, it took 47 minutes to find the fix: a missing REDIS_HOST in the new deployment slot.
Nothing about the problem was especially hard. The hard part was remembering that it had happened.
That is the problem we set out to prototype: an incident-response agent that keeps past incidents, post-mortems, and runbooks in long-term memory, then recalls the closest matches when something breaks. This article describes what we actually built, including which parts talk to a real memory system and which parts are still a scripted demo.
The key distinction. The project includes a real Hindsight integration for retaining and recalling incident text. The deployed front end does not call it.

1. The problem

Incident knowledge is scattered across tickets, post-mortems, runbooks, and people's heads. Under pressure, engineers re-diagnose failure classes the team has already solved. The knowledge exists; retrieval is the bottleneck.
Keyword search helps a little, but reports describe the same failure in different words:
| The alert says | The post-mortem says |
| --- | --- |
| "502s after deploy" | "missing env var in new slot" |
We wanted memory that connects symptoms to root causes to fixes.

2. What we built

The project is called Hindsight Incident Copilot. Its repository contains three pieces:
| Piece | What it does |
| --- | --- |
| Web front end TanStack Start, React, Tailwind | A landing page and an interactive chat demo. A recommendation sits on the left and a "Retrieved evidence" panel (titles, snippets, similarity bars) on the right. |
| Backend stub backend-stub/main.py | A FastAPI service using the hindsight-client package to call recall and retain against a Hindsight server. |
| Seed script and data data/seed-hindsight.py | Meant to load five incidents and three runbooks into a memory bank named org-incidents. |
One boundary needs to be stated plainly: the deployed front end does not call the backend. Its replies come from a hard-coded lookup in src/lib/incident-knowledge.ts. The Hindsight code exists separately; the two are not connected.

3. System architecture

The intended flow, as drawn in the docs and reflected in the backend stub: an engineer describes an alert, the backend recalls similar memories from Hindsight, an LLM turns those memories into a recommendation, and after the incident is resolved the post-mortem is retained back into memory.
The recall and retain steps exist as code. The LLM step does not. The /incident endpoint returns fixed placeholder text after a successful recall, and a source comment says to replace it with a real LLM call. The deployed UI takes a separate path: user messages are matched against keywords and return canned replies.
Figure 1. Two separate paths. Path A is what users see today. Path B is the real Hindsight integration, not yet connected to the UI.

4. How Hindsight provides memory

We used Hindsight, an open-source agent memory system from Vectorize. It runs as a server, started locally by scripts/start-hindsight.sh using the ghcr.io/vectorize-io/hindsight:latest Docker image on port 8888. The script passes an LLM API key for Hindsight's own processing. Our integration uses two client calls, both present in the repository.

Retain: write incident knowledge

The seed script flattens each incident into a text block and retains it. Runbooks are retained the same way, as numbered steps. We build no graph or embeddings ourselves; we pass prose to retain.
data/seed-hindsight.py

content = f"""Incident {inc['id']} ({inc['date']}) — Service: {inc['service']}
Symptoms: {inc['symptoms']}
Root cause: {inc['root_cause']}
Resolution: {inc['resolution']}
Tags: {', '.join(inc['tags'])}
Related runbook: {inc.get('runbook', 'N/A')}
"""
client.retain(bank_id=BANK_ID, content=content)
Enter fullscreen mode Exit fullscreen mode

Recall: retrieve relevant memories

The backend queries the bank and maps each result into the shape the UI expects. A code comment acknowledges that the result shape "depends on the Hindsight client version," so the mapping is a guess until checked against a real response.
backend-stub/main.py

client = Hindsight(base_url=HINDSIGHT_URL)
results = client.recall(bank_id=BANK_ID, query=query)

for i, r in enumerate(results[:5] if isinstance(results, list) else []):
    evidence.append({
        "title": getattr(r, "title", None) or f"Memory #{i+1}",
        "similarity": float(getattr(r, "score", 0.8) or 0.8),
        "snippet": str(getattr(r, "content", r))[:280],
    })
Enter fullscreen mode Exit fullscreen mode

Retain after resolution: add the post-mortem

A POST /retain endpoint accepts a post-mortem and calls client.retain(...). This is the write path for learning. The docs and UI copy also describe Hindsight's reflect operation, but no code calls it.

5. Before vs. after

The implemented before-and-after is at the UI level. Before, the evidence panel says, "Memories appear here after you describe an incident." After clicking "PaymentService returning 502s after latest deploy," the demo displays three evidence cards beside a four-step recommendation that starts with checking the exception stack and verifying environment variables in the current slot.
INC-2847*94%*
PaymentService 5xx runbook*88%*
INC-3102 (connection pool exhaustion)76%

The content is faithful to the sample data. INC-2847's recorded root cause is the missing REDIS_HOST, and INC-3102's is pool exhaustion, with the pool maximum raised from 20 to 50. The similarity figures, however, are typed into a TypeScript file; they are not produced by retrieval. We did not run a controlled comparison of resolution time with and without memory, and we do not report one.

6. How the agent learns over time

The learning design is simple: every resolved incident becomes new text in the bank via retain, so the next recall can surface it. The /retain endpoint and the seed script are the implementation. Several parts remain unimplemented:
| Gap | Detail |
| --- | --- |
| No memory on/off toggle | The front end has no switch to compare behavior with and without memory. |
| No learning curve | No measurements of recall quality as the bank grows, and no chart or test showing improvement. |
| No automatic capture | The docs list automatic retain() after resolution, but the only trigger in the code is a manual call to /retain. |
| No live bank counts | The panel shows 1,247 incidents, 892 post-mortems, and 156 runbooks. These are literals; the sample data has five incidents and three runbooks. |
The UI and landing page also show a "40–70% target MTTR reduction" and a "94.6% Hindsight LongMemEval" figure. Neither was measured by this project, so we do not rely on them here.

7. Technical implementation

The front end is a TanStack Start app with a single IncidentDemo component holding message, typing, and evidence state. Its current behavior is a keyword lookup:
src/lib/incident-knowledge.ts

export function matchIncident(text: string): AgentReply {
  const l = text.toLowerCase();

  if (l.includes("database") || l.includes("latency") || l.includes("orders")) {
    return REPLIES.database;
  }
  if (l.includes("deploy") || l.includes("canary") || l.includes("health")) {
    return REPLIES.deploy;
  }
  return REPLIES.default;
}
Enter fullscreen mode Exit fullscreen mode

The demo has three canned replies with evidence cards and a 900 ms timer imitating latency. It is a UI prototype of what a recall-backed answer should look like.

The backend stub is a small FastAPI service with GET /health, POST /incident, and POST /retain. The /incident route returns a source field of "hindsight" or "simulated", and falls back to a canned PaymentService answer if the client is missing or recall returns nothing.

Documentation and UI copy also mention Azure OpenAI, Azure Monitor, Teams, and an agent framework provider. No code in the repository implements any of them, so we treat them as a design sketch.

8. What we learned

A mocked front end can hide a missing integration

Because the canned replies looked right, it was easy to feel further along than we were. The backend's source flag should be shown in the UI.

Seed data must match its script

In sample-incidents.json, the symptom text is stored under the key severity. The seed script reads inc['symptoms']. By inspection, that raises a KeyError on the first incident, so the seed step as written would fail. We found this by reading the files side by side and did not run the script. One end-to-end run would have caught it.

Unit mismatches leak across the boundary

The front end expects similarity as 0 to 100 (width: ${e.similarity}%). The backend stub's simulated data uses 0.94. Wiring them together without a conversion would render one-pixel bars.

Don't guess third-party response shapes

The recall mapper uses getattr with fallbacks and defaults the score to 0.8 when none is present. That would display a plausible-looking but invented similarity. Until we inspect a real recall response, the evidence panel cannot be trusted.

Write the "not implemented" list early

Separating what the code does from what the docs promise gives us a next-steps checklist:

  • Connect the UI to /incident
  • Add an LLM step that uses recalled text
  • Verify the recall schema
  • Decide whether reflect belongs in the flow

Conclusion

Hindsight Incident Copilot is a working demonstration of the interaction we want (an alert goes in, and history-backed evidence comes out), plus a small, real Hindsight integration for retaining and recalling incident text. The front end is still scripted, the LLM step is a placeholder, and the learning behavior is designed but unmeasured.

The most valuable next step is an end-to-end run against a live Hindsight bank, followed by an honest comparison of answers with and without memory.


Resources: Hindsight GitHub, Hindsight Documentation, Vectorize Agent Memory.


Top comments (0)