DEV Community

Cover image for EchoOps: When Incident Memory Changes the Next Decision
Vegu Priya
Vegu Priya

Posted on

EchoOps: When Incident Memory Changes the Next Decision

EchoOps: When Incident Memory Changes the Next Decision

Production incidents are rarely completely new. A Payment API can fail with the same symptoms, engineers can repeat the same troubleshooting steps, and valuable operational knowledge can disappear between incidents.

We built EchoOps around a simple question: what if an incident-response system could remember what failed, remember what worked, and use that experience when a similar incident happened again?

EchoOps is a focused incident-response prototype that uses Hindsight as its persistent experience-memory layer. Its core loop is:

Incident → Investigate → Recall → Recommend → Act → Observe → Retain → Learn

The important part is not simply storing incident history. It is making previous experience change the system's next decision.

The Problem

Consider a Payment API returning HTTP 503 errors during high traffic.

A response system can inspect the current symptoms and produce a reasonable first action. But without persistent experience, the next similar incident starts from the same point.

For EchoOps, we created a deterministic Payment API incident scenario containing HTTP 503 errors, high traffic, connection errors, and an elevated error rate. The simulator also provides synthetic logs and metrics so the workflow can be reproduced consistently.

The first incident follows a simple path.

The system investigates the incident and initially recommends restarting the Payment Service. The simulated restart fails because it does not relieve the connection exhaustion. Further investigation identifies the database connection pool as the root cause. Increasing the connection-pool capacity succeeds and the simulated error rate returns to baseline.

That experience contains more than a successful resolution. It also contains a failed action.

That distinction becomes important during the second incident.

What EchoOps Does

EchoOps has a React and Vite console connected to a FastAPI backend.

React + Vite Console
        |
        v
     FastAPI
        |
        +-- Incident Simulator
        |
        +-- Response Planner
        |
        +-- HindsightService
                |
                v
          Hindsight Cloud
Enter fullscreen mode Exit fullscreen mode

The frontend uses React, TypeScript, Vite, CSS, and Lucide React. The backend uses Python, FastAPI, Pydantic, and the official hindsight-client Python SDK.

The Hindsight integration is isolated inside backend/app/hindsight_service.py. This gives the application a clear boundary between incident-response logic and the external memory layer.

The incident and simulator state remains in process memory, while Hindsight provides the persistent experience memory. The repository explicitly avoids a local-memory fallback.

Hindsight as the Experience Layer

EchoOps uses two core Hindsight operations: Recall and Retain.

Recall happens before a response recommendation. The current incident is converted into a query containing the service, symptoms, HTTP 503 condition, traffic, connection errors, and the type of historical experience being searched for.

The actual integration uses the official Hindsight Python client:

result = await self._get_client().arecall(
    bank_id=self.bank_id,
    query=query,
    budget="mid",
)
Enter fullscreen mode Exit fullscreen mode

This call is implemented in HindsightService.recall_similar_incidents().

The result is converted into structured matches containing the memory ID, text, type, and metadata. EchoOps then passes those actual recalled results into its response-planning logic.

Hindsight's memory model is particularly useful here because the system is not just looking for an identical incident description. It is using stored experience to retrieve information about previous actions, outcomes, root causes, lessons, and strategies. Hindsight provides the persistent memory layer that makes this possible.

Hindsight documentation

From Recall to a Different Decision

This is the central technical behavior of EchoOps.

After Recall, the response planner examines the actual returned memory text. It checks whether the memory is relevant to the current Payment API 503 scenario and whether it contains evidence of both the previous failed restart and successful pool remediation.

The relevant decision logic is implemented directly in backend/app/main.py:

recalled_text = " ".join(
    match["text"] for match in incident["recall"]["matches"]
).casefold()

relevant = bool(incident["recall"]["matches"]) and any(
    term in recalled_text
    for term in ("payment", "503", "connection pool")
)

previous_failure = relevant and "restart" in recalled_text and any(
    term in recalled_text
    for term in ("fail", "unsuccessful", "did not resolve", "did not fix")
)
Enter fullscreen mode Exit fullscreen mode

The planner then determines whether the previous experience contains both a failed and successful outcome.

If it does, the recommendation changes from restarting the service to inspecting the connection pool first.

"recommendation": (
    "Inspect the database connection pool first"
    if learned_outcomes
    else "Restart Payment Service"
)
Enter fullscreen mode Exit fullscreen mode

This is the key distinction between remembering information and using memory.

The retrieved experience changes the response strategy.

Retaining the Outcome

After an incident is resolved, EchoOps sends the complete experience back to Hindsight.

The Retain operation includes the incident, service, symptoms, evidence, hypothesis, investigation, attempted actions, outcomes, root cause, resolution, lesson, future strategy, actions to avoid, and applicability conditions.

The actual implementation uses:

result = await self._get_client().aretain(
    bank_id=self.bank_id,
    content=content,
    context="Resolved SRE incident and its verified action outcomes",
    metadata={
        "incident_id": str(experience["incident"]),
        "service": str(experience["service"]),
        "kind": "incident_experience",
    },
)
Enter fullscreen mode Exit fullscreen mode

This makes the learning loop explicit:

Recall past experience → make a decision → observe the outcome → Retain the new experience.

The failed restart is deliberately retained alongside the successful pool remediation. A future incident can therefore learn not only what worked, but also what should be avoided.

Hindsight GitHub repository

The Before-and-After Behavior

The strongest demonstration comes from running two similar incidents.

Without relevant experience With relevant Hindsight experience
Payment API returns 503 Similar Payment API 503 occurs
Generic first response is generated Previous incident is recalled
Restart Payment Service Previous restart failure is recognized
Restart fails Connection pool is investigated first
Investigation continues from scratch Previous successful remediation informs the plan
No persistent learning Resolved experience is retained for future incidents

For Incident #1, the simulator starts without relevant historical experience. The generic recommendation is to restart the service. That action fails.

After the connection pool is inspected and capacity is increased, the incident resolves. EchoOps retains that experience in Hindsight.

For Incident #2, the system recalls the previous incident. The planner recognizes that restarting previously failed while increasing connection-pool capacity succeeded.

The recommendation therefore changes to:

Inspect the database connection pool first.

The UI also marks the recommendation as memory-influenced.

That is the behavior EchoOps was designed to demonstrate.

Testing the Learning Loop

The repository includes focused tests for the core memory workflow.

The tests verify that EchoOps can:

  • investigate an incident
  • call Hindsight Recall with an incident-specific query
  • record failed and successful actions
  • retain a resolved experience
  • adapt a second incident using recalled experience
  • handle missing Hindsight configuration
  • preserve Hindsight memories when the simulator is reset
  • ignore unrelated memories when generating a plan
  • avoid falsely claiming Hindsight connectivity

One of the most important tests creates the first incident, resolves it, retains the experience, and then starts Incident #2. It verifies that the second incident recalls Incident #1 and changes the recommendation to inspect the database connection pool first.

The repository contains nine focused learning-loop tests using a Hindsight SDK test double. These tests validate the integration boundary without claiming that they replace testing against a live Hindsight account.

What We Learned

1. Memory matters when it changes behavior

A memory panel alone is not enough. The useful signal is a change in the next decision.

2. Failed actions are valuable experience

The failed restart becomes useful knowledge for the next incident. Knowing what not to repeat can be as valuable as knowing what worked.

3. Retrieval relevance matters

Not every historical memory should influence an incident. EchoOps checks the recalled content against the current scenario before marking a decision as memory-influenced.

How Hindsight Fits Into the Learning Loop

EchoOps uses two important Hindsight operations:

Recall retrieves relevant previous experience before generating a response plan.

Retain stores the completed incident experience after the incident has been resolved.

Hindsight's documentation describes Retain as processing incoming information, extracting structured memories, and indexing them for retrieval. Recall retrieves relevant memories using multiple retrieval strategies.

EchoOps does not use a local JSON memory fallback. The repository explicitly requires a configured Hindsight bank and reports an error when Hindsight credentials are missing.

4. The memory boundary should be explicit

Keeping Hindsight operations inside HindsightService makes the integration easier to reason about and test.

5. A simulator is useful, but it is not production validation

EchoOps uses synthetic telemetry and deterministic actions. It does not contact production infrastructure. A production version would need real observability integrations, stronger safety controls, approval workflows, and carefully controlled remediation.

Limitations and Next Steps

The current prototype intentionally focuses on one workflow: a recurring Payment API 503 incident.

Incident and simulator state is held in process memory and resets when the backend restarts. Hindsight remains the persistent experience layer.

The response planner is also deliberately deterministic and explainable. The current repository does not make an external LLM call. Instead, it uses transparent rules over actual Hindsight Recall results.

Future versions could connect the same learning loop to real logs, metrics, alerts, deployment systems, and controlled remediation tools.

The architecture could remain simple:

Observe → Recall → Decide → Act → Observe → Retain

Conclusion

EchoOps started with a common operational problem: similar incidents often require similar investigations, but the experience from previous incidents is not always available when it matters.

The prototype demonstrates another approach.

The first incident creates experience. Hindsight stores that experience. A similar second incident recalls it. The response planner recognizes the previous failed action and successful remediation, then changes its recommendation.

The important result is therefore not simply that the system can remember the past.

It is that the past can change what the system does next.

Every incident leaves an echo. Every response gets smarter.

Tech Stack

Layer Technology
Frontend React + TypeScript + Vite
UI CSS + Lucide React
Backend Python + FastAPI
Data Models Pydantic
Agent / Response Logic Deterministic Python response planner
Persistent Memory Hindsight Cloud
Hindsight Integration Official hindsight-client Python SDK
API REST / FastAPI
Testing pytest + FastAPI TestClient
Deployment Configuration Render + Vercel

Project

EchoOps GitHub:
https://github.com/mohankumard18/echoops-incident-learning

Hindsight:
https://hindsight.vectorize.io/

Hindsight GitHub:
https://github.com/vectorize-io/hindsight

Vectorize Agent Memory:
https://vectorize.io/what-is-agent-memory

Top comments (0)