DEV Community

Cover image for I Built a Hindsight Agent That Remembers Failed Fixes
surya
surya

Posted on

I Built a Hindsight Agent That Remembers Failed Fixes

Production incidents are rarely difficult because nobody knows how to restart a service.

They are difficult because engineers have to remember what happened the last time.

A similar incident may have happened weeks or months ago. Someone may have tried a restart, a rollback, or a configuration change. One approach worked. Another looked reasonable but failed. That experience is usually scattered across logs, tickets, documentation, or someone's memory.

I wanted an incident-response agent that could use that experience directly.

That became OpsMind.

OpsMind is an AI incident-response system that investigates live failures, retrieves relevant past experiences using Hindsight, recommends a remediation, verifies the result independently, and stores the outcome for future incidents.

The interesting part is not that the agent can read logs.

The interesting part is that it can remember what happened after an action was taken.

The problem with stateless incident agents

A normal LLM agent can analyze the current incident.

For example, if a gateway starts returning HTTP 502, an agent can inspect service health, read logs, and look at the current deployment configuration.

That helps answer:

“What is happening right now?”

But incident response often needs another question:

“What happened the last time this happened?”

That historical context can change the next action.

A previous incident may show that restarting a backend did not solve the problem, while changing a gateway configuration did. Without memory, the agent starts the investigation again from a blank context.

I wanted memory to become part of the decision-making process rather than another dashboard panel.

Building a real incident environment

Instead of giving the agent a static incident JSON file, I built a small local environment with two real services.

The first is a payment backend.

The second is an API gateway that forwards requests to the payment backend.

The backend exposes health and payment endpoints and writes timestamped logs.

The gateway reads its upstream target from a live configuration file and forwards actual HTTP requests to the backend.

For the incident scenario, the upstream configuration is deliberately changed from the healthy backend port to an invalid local port.

That produces a real failure:


text
Gateway → 502 Bad Gateway
Backend → 200 OK
Logs   → connection refused
Config → upstream changed to invalid port
Enter fullscreen mode Exit fullscreen mode

Top comments (0)