DEV Community

Kawsik M
Kawsik M

Posted on

Built an AI Incident Response Agent That Remembers What Worked

How do you know an agent's memory is helping? "The answers feel better" is the usual evidence, and it's terrible. I wanted something I could point at, so I built the evaluation into the product: one endpoint that answers the same incident twice, once with memory and once without.

It ended up being the most useful debugging tool in the whole system.

The system

OnCall Memory is an incident co-pilot. An engineer submits an incident (title, service, description, symptoms, logs). The backend recalls similar past incidents from a long-term memory bank, grades their relevance to the service in question, and has an LLM produce a structured diagnosis and runbook. Resolved incidents get retained back into the bank.

The stack: FastAPI on Python 3.11, a React 19 + TypeScript frontend, Groq serving openai/gpt-oss-120b, and Hindsight as the memory engine, holding a bank called oncall-memory. Two containers under Docker Compose; storage is embedded Postgres with pgvector inside the Hindsight container, on a named volume.

The through-line: make the counterfactual cheap

When you add memory to an agent, the question that matters is what the agent would have said without it. That's the counterfactual, and normally it's expensive: you'd need to rerun old inputs against an old configuration and squint at the diff.

So POST /api/incidents/compare does this in one request. The same incident goes through two paths:

With memory: recall from Hindsight, grade relevance, reason over the results.
Without memory: the cold-start baseline. Same model, same schema, no recalled context.

Here's a real-shaped response for a connection-pool incident on order-service:

json
{
"with_memory": {
"likely_root_cause": "HikariCP connection leak in OrderPaymentProcessor",
"confidence": "high",
"recommended_fix": "Set leakDetectionThreshold and add circuit breaker. Pod restart previously FAILED."
},
"without_memory": {
"likely_root_cause": "Generic database connection saturation",
"confidence": "medium",
"recommended_fix": "Restart the order-service pod and increase database max_connections.",
"runbook": [ "Restart application pod", "Scale database replica" ]
},
"contrast_summary": "Without memory, the agent suggested restarting the pod (which historically failed). With Hindsight memory, it recalled the specific leak in OrderPaymentProcessor and warned against restarting."
}

The cold-start answer isn't a strawman. "Restart the pod and raise max_connections" is what a reasonable engineer, or model, says when it knows nothing about your system. It's also the remedy that failed on 4 May, saturating again in eight minutes. The comparison shows the value of memory as a specific wrong action avoided.

Why the baseline uses the same model

I could have made the "without" path dumber, and it would have made the memory look better. I refused, for a reason. If both paths run openai/gpt-oss-120b at temperature 0.2 against the same Pydantic output contract, then any difference between them is attributable to one variable: the recalled context. That's a controlled comparison as far as a single request can be.

The contract is shared:

typescript
interface StructuredRecommendation {
likely_root_cause: string;
confidence: "high" | "medium" | "low";
historical_evidence: string[];
recommended_fix: string;
citing: string;
runbook: string[];
}

With no memory, historical_evidence is empty and confidence lands at medium or lower, and the diagnosis has nothing to cite. That's the right behavior. An agent that's equally confident with and without evidence is broken.

Where memory comes from: Hindsight

The "with memory" side depends on recall quality. Hindsight does multi-strategy retrieval (Temporal, Entity, Semantic, Recency), and I get all four from one call. My application adds a relevance layer on top of that, grading each recalled memory as direct_match, analogy or low, so a memory from a different service can't be presented as proof about this one.

The comparison view exposes those grades. I can see when the with-memory answer is better because of a direct precedent versus when it's leaning on an analogy, and I can see when recall pulled in noise. For background on why this layer matters, Vectorize's overview of agent memory is a decent primer.

Using the comparison to change things safely

Here's how I actually use it. When I change a prompt, a relevance rule or the memory format, I run the same handful of presets through /compare and read the diffs.

Did with-memory still cite the right dates?
Did without-memory stay generic, or did it start leaking specifics?
Did an analogy get promoted to a direct match?

That last one is a real regression class. A change to the service-matching logic that loosens is_same_service makes cross-service memories look like evidence, and the with-memory answer starts sounding more confident. In the comparison view, that looks like a suspiciously good answer for a service with no history.

The same scenarios exist as tests under backend/tests/: a cold start on a service without history (must not hallucinate another service's components, confidence medium or low), closed-loop learning (save a WORKED fix, see it cited with high confidence next time), and a failed-remedy warning. The tests assert; the comparison view lets me look.

What it doesn't tell you

I want to be straight about limits. A single-incident comparison is an existence proof, not a benchmark. It shows that memory can change the answer for the better on a given input. It doesn't tell you the rate across your incident history. If you want that number, you'd replay a corpus of historical incidents through both paths and have engineers grade the results, which I haven't claimed to do here.

It also can't tell you when memory is wrong. If a stored root cause was mistaken, the with-memory answer will be confidently mistaken too. That's why resolved incidents are written through a form that echoes back exactly what will be retained:

json
{
"title": "Postgres Connection Pool Saturation",
"service": "order-service",
"root_cause": "HikariCP connection leak in OrderPaymentProcessor",
"fix": "Set leakDetectionThreshold=2000ms and wrapped external calls in resilience4j circuit breaker",
"outcome": "WORKED"
}
Lessons

Build the counterfactual into the product. If "with vs. without" is one request, you'll actually run it. If it's a notebook, you won't.

Keep the baseline honest. Same model, same schema, same temperature. Only the recalled context varies.

Read diffs when you change anything. Prompt edits, ranking tweaks and format changes all shift behavior. The side-by-side is a regression check a unit test can't be.

Suspiciously confident is a bug signal. A high-confidence answer on a service with no history means the relevance layer is leaking.

Be clear about what a demo of value proves. One input shows possibility. Rates need a corpus and human grading.

If you want to try this pattern with your own agent, the Hindsight docs cover retaining and recalling, and the comparison endpoint is about thirty lines on top of that.

Top comments (0)