DEV Community

V Chaitanya
V Chaitanya

Posted on

How I Built Incident Memory with Hindsight

How I Built Incident Memory with Hindsight

Production incidents rarely happen in isolation.

A service can fail for a different reason, a different service can be involved, and the logs can look completely different from the previous incident. But often, there is useful experience hidden in what happened before.

That led me to build MemoryOS — an AI incident-response copilot designed to connect current production incidents with previous investigations.

The idea is simple:

Investigate once. Remember the resolution. Use that experience when the next incident happens.

MemoryOS uses Hindsight as its persistent memory layer. It can recall relevant previous incident experiences before reasoning about a new incident, then retain the investigation so that future incidents have more context.

The problem with stateless incident analysis

A typical AI incident assistant can take an incident, inspect its symptoms and logs, and generate a diagnosis.

That is useful, but there is an important limitation.

Suppose an engineering team investigates an incident where an API starts returning timeouts because a downstream database connection pool becomes exhausted.

The team identifies:

  • Increased database latency
  • Connection pool exhaustion
  • Requests waiting for connections
  • Upstream timeouts
  • Cascading API failures

The incident is resolved.

Later, another service experiences a similar problem.

The service name is different. The wording of the incident is different. The logs are different.

A stateless assistant may still treat the new investigation as completely independent.

The previous investigation contained valuable operational knowledge, but that knowledge isn't automatically available.

That is the problem MemoryOS is designed to address.

Turning incident response into a memory loop

Instead of treating every incident as an isolated interaction, MemoryOS connects investigation and memory.

The workflow is:


text
New Incident
     ↓
Recall
     ↓
Relevant Experience
     ↓
Reason
     ↓
Structured Investigation
     ↓
Resolution
     ↓
Retain
     ↓
Future Incident

The important part is that memory is not simply an extra page in the application.

It is part of the investigation workflow.

When an incident arrives, MemoryOS captures the service, symptoms, and logs. It then uses Hindsight to retrieve relevant previous experiences.

The retrieved context is combined with the current incident before the investigation is generated.

After the investigation, useful information is retained for future incidents.

The MemoryOS dashboard

The main dashboard is designed around the idea of an incident workspace rather than a generic AI chat interface.

It provides an overview of investigations, retained memories, recalled matches, and services involved in previous investigations.

The dashboard also makes the Hindsight connection visible without making it the entire interface.

The key information shown in the dashboard includes:

Investigations recorded in the workspace
Memory retained through Hindsight
Memories recalled for the latest investigation
Services investigated
Current incidents
Hindsight connection status

During testing, the system retrieved 26 memories for the latest investigation.

That number isn't treated as the answer itself.

The important part is what those memories contain and how they influence the investigation.

![ ](https://dev-to-uploads.s3.us-east-2.amazonaws.com/uploads/articles/v0gxs24yyzp5w6txdlwv.png)

How Hindsight fits into the architecture

MemoryOS uses two important components during investigation:

                 MemoryOS
                    │
             Incident API
                    │
          ┌─────────┴─────────┐
          │                   │
          ▼                   ▼
      Hindsight              Groq
    Memory Layer              LLM
          │                   │
          └─────────┬─────────┘
                    ▼
           Structured Analysis

The frontend is built with:

Next.js
React
TypeScript
Tailwind CSS
Framer Motion
Lucide React

The backend uses Next.js API routes and TypeScript.

Groq handles the reasoning step, while Hindsight provides persistent memory that can be recalled across investigations.

Recalling previous incidents

The first important operation is recall.

MemoryOS builds a context from the current incident and sends it to Hindsight.

A simplified version of the recall request looks like this:

const response = await fetch(
  `${HINDSIGHT_URL}/v1/default/banks/${BANK_ID}/memories/recall`,
  {
    method: "POST",
    headers: {
      "Content-Type": "application/json",
    },
    body: JSON.stringify({
      query: incidentContext,
      types: ["experience", "observation"],
      budget: "mid",
      max_tokens: 3000,
    }),
  }
);

The returned memories become historical context for the current investigation.

The reasoning process therefore becomes:

Current Incident
       +
Historical Experience
       ↓
AI Investigation
       ↓
Root Cause + Evidence + Actions

This is the part that makes persistent memory useful.

The LLM is not expected to magically remember previous incidents.

Instead, the application retrieves relevant experience and gives the model that context when it matters.

A real investigation example

One of the incidents in my testing environment was:

Product Search Degradation

Service:

catalog-api

The incident involved slow product searches and intermittent failures.

The investigation produced a strong historical match.

The system identified a previous incident involving the inventory-api service and connected it with the current catalog degradation.

The investigation identified:

Root cause

Database connection pool exhaustion in the inventory-api service was causing upstream timeouts and search latency spikes.

Evidence

The current incident contained upstream timeout information and requests waiting for inventory responses.

Historical incident context provided another important signal: a previous degradation had been associated with database connection pool exhaustion.

Immediate actions

The generated investigation recommended actions such as:

Check inventory connection-pool metrics.
Temporarily increase pool capacity when appropriate.
Restart the affected service if required.
Add circuit-breaker logic for upstream calls.
Investigate and optimize slow database queries.

The important behavior here isn't that the AI generated a list of possible fixes.

It is that previous incident experience became part of the current investigation context.

The Hindsight memory view

MemoryOS also provides a dedicated memory view.

It shows the experiences retrieved from Hindsight instead of hiding memory completely behind the AI response.

During the latest investigation, the interface showed:

26 memories recalled

and multiple retrieved experiences.

For example, one memory described a recommendation API incident where customer-facing latency increased significantly during high traffic.

Another described the previous Product Search Degradation incident.

Another contained immediate mitigation steps from that investigation.

This makes the memory system inspectable.

Instead of simply saying:

"The AI remembers previous incidents."

MemoryOS can show the actual historical context being retrieved.

![ ](https://dev-to-uploads.s3.us-east-2.amazonaws.com/uploads/articles/nys9ohnk0486nkulgjmv.png)

From incident to reusable knowledge

The investigation workflow inside MemoryOS is intentionally simple:

01 — Incident
Capture symptoms, logs and service context.

02 — Recall
Search Hindsight for relevant experience.

03 — Reason
Correlate current signals with prior outcomes.

04 — Learn
Retain the resolution for future incidents.

This creates a continuous loop:

Incident
   ↓
Recall
   ↓
Reason
   ↓
Resolve
   ↓
Retain
   ↓
Future Incident

The goal is to turn individual incident investigations into reusable operational knowledge.

Investigation history

MemoryOS also keeps a history of investigations.

The history view shows incidents such as:

Product Search Degradation
Inventory Database Saturation
Homepage Personalization Timeout

Each investigation can be opened to inspect the information associated with that incident.

The interface also shows how many memories were associated with the investigation.

This creates a simple progression:

Incident
   ↓
Investigation
   ↓
Memory
   ↓
Future Investigation

Instead of allowing previous work to disappear after an incident is resolved, the system keeps it available as context.

![ ](https://dev-to-uploads.s3.us-east-2.amazonaws.com/uploads/articles/uadsib1bfqelt8hglhvk.png)

Retaining useful information

One of the more important design decisions was deciding what should actually become memory.

More stored information doesn't necessarily mean better memory.

MemoryOS focuses on information that can be useful during future incident investigations:

Root cause
Evidence
Immediate actions
Memory matches
Confidence
Lessons for future incidents

The investigation can then be retained as operational experience.

That means the memory isn't just a collection of raw logs.

It represents what was learned from the investigation.

Before vs. after persistent memory

The behavioral difference can be summarized simply.

Without persistent memory   With Hindsight
Each incident starts independently  Previous experiences can be recalled
Current logs are the primary context    Current logs + historical context
Previous investigation knowledge can be lost    Useful investigation knowledge can be retained
Similar failures may require repeated analysis  Historical patterns can inform the next investigation
AI sees the current incident    AI can reason with previous incident experience

The objective isn't to make the AI automatically correct.

The objective is to make the investigation better informed.

What I learned while building MemoryOS

The most interesting part of the project wasn't simply connecting an LLM to a memory API.

It was figuring out what memory should actually do.

A memory system can store a huge amount of information.

But storage alone isn't useful.

The important question is:

What previous experience will actually help with the incident happening now?

That changed how I designed the workflow.

Instead of thinking about Hindsight as a database, I started thinking about it as part of the investigation process.

Remember
   ↓
Use
   ↓
Resolve
   ↓
Learn
   ↓
Remember again

That is the behavior I wanted MemoryOS to demonstrate.

Limitations

MemoryOS is currently a focused incident-response project rather than a complete enterprise incident-management platform.

There are several limitations.

Production integrations

Incident information is currently entered through the application rather than being automatically collected from production monitoring systems.

Memory quality

The usefulness of an investigation depends partly on the relevance of the memories retrieved from Hindsight.

Human verification

A historical match should be treated as context, not absolute truth.

Engineers still need to verify the evidence before taking production action.

No automatic production changes

MemoryOS generates investigation findings and suggested actions, but it does not automatically modify production infrastructure.

Development environment

The current implementation is configured as a development project using locally running services.

These limitations are important because an incident-response assistant should assist engineering judgment rather than pretend to replace it.

What's next?

There are several directions I would like to take MemoryOS next:

Connect observability platforms directly
Integrate application and infrastructure logs
Add monitoring alerts
Build service dependency graphs
Retrieve relevant runbooks
Add Slack or Microsoft Teams workflows
Generate post-incident reviews
Detect recurring failure patterns
Track incident resolution time
Add role-based access control
Deploy the system as production services

The long-term idea is to make the memory layer useful across the entire incident lifecycle.

Closing

MemoryOS started from a simple observation:

Production teams already have experience. The challenge is making that experience available when the next incident happens.

An incident-response system shouldn't only answer:

"What is happening right now?"

It should also be able to ask:

"Have we seen something like this before, and what did we learn?"

That is where Hindsight becomes important.

By combining an incident-response workflow with persistent memory, MemoryOS can retrieve previous production experiences, use them as context during investigation, and retain useful outcomes for future incidents.

The result isn't an AI that magically knows every production problem.

It's something more practical:

An incident-response system that can remember what happened before — and use that experience when the next incident arrives.

Built with

Next.js · React · TypeScript · Tailwind CSS · Framer Motion · Groq · Hindsight

GitHub: https://github.com/Chaitanya-CSP/memoryos

Hindsight: https://github.com/vectorize-io/hindsight

## Learn more

- [Hindsight GitHub](https://github.com/vectorize-io/hindsight)
- [Hindsight Documentation](https://hindsight.vectorize.io/)
- [Vectorize Agent Memory](https://vectorize.io/what-is-agent-memory)
## Project

[MemoryOS on GitHub](https://github.com/Chaitanya-CSP/memoryos)
Enter fullscreen mode Exit fullscreen mode

Top comments (0)