How We Built an Incident Response Agent That Stops Repeating the Same Mistakes Using Hindsight
There's a bug that lives in almost every ops team I've worked with. Not in the codebase — in the process. An engineer gets paged at 2am. The database is melting. They fix it — increase the connection pool from 20 to 50, write a postmortem, close the ticket. Two months later, a different engineer gets the same page. Different shift. Different memory. Same fix. Same 45 minutes wasted.
We built OpsMind to kill that cycle.
What the System Does
OpsMind is an incident response agent backed by Hindsight — a biomimetic memory layer for AI agents. When a production incident arrives, OpsMind doesn't just suggest generic fixes. It searches organizational memory for what has been tried before, compares the historical environment state against the current one, and — crucially — refuses to recommend a fix that's already been applied.
The full flow looks like this:
An incident is ingested (title, description, service, logs, current environment config)
An LLM extracts structured attributes: symptoms, error type, severity, relevant technical entities
Hindsight RECALL retrieves the most relevant past incidents from the memory bank
A differential reasoning engine compares historical state vs current state
Targeted actions are generated and queued for human approval
After resolution, Hindsight RETAIN stores the structured outcome so future agents learn from it
The stack is FastAPI on the backend, React + Tailwind on the frontend, PostgreSQL for transactional records, and Hindsight Cloud as the organizational memory layer.
The Core Technical Story: Differential Reasoning Over Memory
Most teams reaching for RAG-style systems in DevOps fall into the same trap: they build a system that finds the nearest historical incident and regurgitates the past solution verbatim. That's not intelligence — that's a fancier way to copy-paste a runbook.
The core insight behind OpsMind is that knowing a historical fix exists is only half the job. The other half is knowing whether that fix is already applied in the current environment.
Here's the scenario that motivated the design:
Incident #1024 (2 months ago): payment-api returning HTTP 500. Root cause: DB connection pool at 20, exhausted under load. Fix: increase pool to 50. Result: SUCCESS.
Incident #1026 (today): payment-api returning HTTP 500. DB timeouts in logs. Current config: db_pool_size = 50.
A naive agent looks at this, finds 91% similarity to Incident #1024, and tells you to increase the pool to 50. Which is already 50. Problem not solved. Confidence misplaced.
OpsMind catches this in _perform_differential_reasoning:
python
Scenario: DB Pool was 20 in past, fixed to 50. Current environment is ALREADY 50!
if current_pool_size and current_pool_size >= 50 and ("pool" in hist_action or "50" in hist_action):
already_applied.append(
f"Connection pool already scaled to {current_pool_size} (matching previous fix)"
)
novel_factors.append("Current environment already incorporates historical pool fix")
explanation = (
f"CRITICAL INSIGHT: The current environment already has db_pool_size = {current_pool_size}.\n"
f"Therefore, blindly repeating the historical pool resize will NOT solve this incident.\n"
f"OpsMind pivots to investigate root causes that starve the enlarged pool "
f"(such as slow unindexed queries and deployment changes)."
)
When the fix is already applied, the agent pivots. Instead of recommending the pool resize, it generates a new action plan:
Check pg_stat_activity for long-running unindexed queries
Diff the latest deployment (v2.4.2) for query regressions
Review application logs for the specific route exhausting connections
This is the difference between an agent that regurgitates memory and one that reasons over it.
How Hindsight Is Wired In
Agent memory is the piece most teams treat as an afterthought. They store embeddings somewhere, do a cosine similarity lookup, and call it a day. Hindsight's approach is different — it implements three distinct operations that map cleanly to the lifecycle of operational knowledge.
The HindsightService class wraps all three:
RETAIN — called at incident resolution. Stores a structured experience footprint: service, symptoms, root cause, action taken, outcome (SUCCESS or FAILED), resolution time, and the config state at the time of the fix.
python
content_text = (
f"Incident: {payload.get('incident_title')} | "
f"Service: {payload.get('service')} | "
f"Root Cause: {payload.get('root_cause')} | "
f"Action: {payload.get('action_taken')} | "
f"Result: {payload.get('result')} | "
f"Symptoms: {', '.join(payload.get('symptoms', []))} | "
f"Config: {json.dumps(payload.get('context_config', {}))}"
)
await client.post(
f"{self.base_url}/v1/default/banks/{self.bank_id}/memories/retain",
json={"items": [{"content": content_text, "document_id": memory_id}]}
)
RECALL — called when a new incident arrives. The query combines service name, extracted symptoms, and error signature. A multi-strategy scorer combines exact service match (+40 pts), symptom token overlap (+30 pts), and keyword match over titles and root causes (+25 pts). Remote Hindsight Cloud recall supplements this when available.
REFLECT — called from the Analytics and Memory Bank pages. Synthesizes organizational patterns: which actions consistently succeed, which fail, which services are repeat offenders. This is where the system stops being reactive and starts being institutional.
The local fallback design is deliberate. If Hindsight Cloud is unreachable, the service reads from a local hindsight_bank.json. Every resolved incident writes to both the cloud and the local bank. The system is never blind.
The LLM Layer
The LLM sits upstream of the memory — its job is to extract structure, not to generate solutions. Given an incident title, description, and log snippet, it returns a JSON payload with extracted symptoms, error type, severity, and relevant technical entities.
We built three tiers into LLMService:
Gemini (via direct REST call to generativelanguage.googleapis.com)
OpenAI (GPT-4o-mini via chat completions)
Built-in heuristic — a keyword-based parser that catches HTTP 500s, connection pool mentions, Redis failures, authentication errors, and gateway timeouts with high precision
python
if self.api_key and self.provider == "gemini":
try:
return await self._call_gemini_analysis(title, description, service, logs)
except Exception as e:
logger.warning(f"Gemini API call failed ({e}). Falling back to heuristic analyzer.")
Built-in DevOps heuristic analyzer
return self._heuristic_analysis(title, description, service, logs)
The heuristic exists because LLM API availability shouldn't be a single point of failure in an incident response tool. In practice, it correctly classifies the most common production failure modes without ever making a network call.
Human-in-the-Loop Actions
The agent generates action recommendations but does not execute them automatically. Every proposed action — restart a service, scale a pool, inspect a deployment — goes into a pending queue and requires explicit human approval through the UI.
This is intentional. The approval modal shows the action type, reasoning, risk level, and relevance source:
memory_differential — emerged from comparing current vs historical state
memory_historical_success — directly mirrors a previously successful resolution
heuristic — generated from pattern matching with no historical precedent
That distinction matters when you're deciding how much trust to extend to a recommendation. Failed or rejected actions are logged, and that outcome feeds back into Hindsight on the next resolution.
Concrete Example: What the Agent Actually Produces
For Incident #1026 with db_pool_size = 50 already set:
Memory Recall: Retrieved Incident #1024, 91% relevance — "DB pool exhaustion, pool=20, fixed to pool=50, SUCCESS"
Differential Verdict: "The current environment already has db_pool_size = 50. The previous fix is incorporated. Pivoting to investigate root causes that starve the enlarged pool."
Generated Actions:
Check database query latency (pg_stat_activity) — risk: low — source: memory_differential
Inspect recent deployment v2.4.2 — risk: low — source: memory_differential
Review application timeout logs — risk: low — source: memory_differential
After engineers investigate and find a missing composite index introduced in v2.4.2, they resolve the incident. They fill in root cause, resolution, and click "Retain Experience in Hindsight." The next time this pattern recurs, the memory bank now contains both the original pool fix and the index regression pattern — and OpsMind knows which one applies based on current config state.
Lessons Learned
The problem isn't finding relevant history — it's knowing when history is stale. Vector similarity is a solved problem. The harder problem is what to do when you find a high-confidence match that no longer applies. Differential reasoning over config state is where the real value lives.
Store FAILED outcomes in Hindsight, not just successes. The Reflect operation surfaces patterns like "blind service restarts fail when queue poison pills are present." That kind of negative knowledge prevents engineers from retrying dead ends. A memory bank that only records successes is half-blind.
Build your memory operations to be dual-write from day one. The local fallback to hindsight_bank.json has saved us during development more times than expected. Cloud availability should never be a prerequisite for operational memory.
LLM calls should extract structure, not generate runbooks. The LLM's job here is entity extraction — symptoms, error types, technical keywords. The reasoning happens downstream, in code, where it's deterministic and testable. Asking an LLM to generate an action plan directly produces suggestions that have no awareness of historical context or current environment state.
Human-in-the-loop isn't a compromise — it's the correct default. An agent that acts autonomously in production might be impressive in a demo. An agent that shows its reasoning, surfaces its evidence, and waits for a human to confirm is the one that gets deployed and stays deployed.
OpsMind is live with the backend on Railway and frontend on Vercel. The Hindsight documentation and GitHub repository cover the full API surface if you want to wire memory into your own agent. For a deeper look at what agent memory actually means in production systems, the Vectorize team has written a solid overview of agent memory architecture.
The operational memory problem isn't going away as organizations run more agents in production. Systems that can learn from what happened and reason about whether that learning still applies are the ones that will actually reduce toil. Everything else is just a fancier runbook.
Top comments (1)
the incident 1024 vs 1026 example is the bit that carries this. 91% similarity pointing at a fix that's already in production is exactly the failure mode nobody demos, because their demo data always matches. catching it before recommending is the difference between an agent people trust and one they route around.
the other line i'd underline is the one about storing failed outcomes, not just successes. most memory-layer writeups only keep wins. the failures are where the actual signal is.