From HTTP 500 to Connection-Pool Exhaustion: How an AI Agent Builds an Incident Investigation Hypothesis
A burst of HTTP 500s is a symptom, not an explanation. When payment requests begin failing just after a traffic spike, “the database is overloaded” may sound plausible—but an engineer still needs to know what to check, what else could explain the errors, and how much confidence to put in an early guess.
That is the narrow problem IncidentMind tries to solve. It takes a small, explicit description of a current incident, searches Hindsight for related incident memories, and gives those memories to an LLM as evidence for a structured investigation. The interesting engineering work is in the boundaries: what goes into the context, how prior cases are compared, and how the system is instructed to keep a hypothesis from turning into a claimed fact.
Start with a specific incident, not a vague question
IncidentMind’s Streamlit form gathers four inputs: service, environment, error, and observed symptoms. For example, its defaults describe a Payment API in Production returning HTTP 500 errors during payment processing, with unusually high database connections shortly after a traffic spike.
Those fields are the investigation context. The app combines them into a Hindsight query and asks for previous incidents involving similar services, errors, symptoms, root causes, triggers, and resolutions. It also explicitly asks that the current incident not be returned as historical experience. In the application, the request is built along these lines:
query = f"""
Current production incident:
Service: {service}
Error: {error}
Symptoms: {symptoms}
Find PREVIOUS and HISTORICAL incidents involving similar:
- service
- error
- symptoms
- root causes
- triggers
- resolutions
IMPORTANT:
Do not return the current incident as historical experience.
"""
The environment is passed to the analysis stage as well, even though this particular recall query is assembled from service, error, and symptoms. Separating retrieval context from reasoning context is useful: Hindsight is asked to find related experience, while the LLM receives all four current incident fields when forming its analysis.
Why “ask an LLM what caused it?” is not enough
A model prompted only with “Payment API, HTTP 500, high database connections” can produce a technically plausible answer. But plausibility is not incident evidence. It may jump to a familiar database explanation, overlook a gateway problem, or recommend a fix before anyone has checked what the current system is doing.
IncidentMind changes the prompt’s inputs. It passes both the current incident and text returned from Hindsight, labeling the latter as historical experience. The model is then asked to explain which old events are relevant, what matches, what differs, and why those similarities matter. This makes the answer a comparison task rather than an unsupported diagnosis.
The app retrieves up to ten memories, filters possible current-incident entries and duplicates, and displays up to five unique memories. The filter matters because a result can look highly relevant precisely because it is the live event being investigated. The interface also displays the retrieved memory text so an engineer can inspect the evidence instead of seeing only the model’s summary of it.
A previous payment outage becomes a lead
The repository’s sample data includes INC-009, a Payment API incident dated 2026-08-27. It records HTTP 500 errors with unusually high database connections, a database connection-pool exhaustion root cause after a traffic spike, and a resolution that increased connection-pool capacity and scaled application workers.
Now suppose today’s incident has the same visible signals: Payment API, HTTP 500 errors, high database connections, and a traffic spike. Hindsight can return INC-009 because the service and symptoms align. That match changes the investigation in a useful way: connection-pool exhaustion becomes a concrete hypothesis, and the old resolution suggests operational areas worth examining—pool capacity and worker load.
But the old incident is not a diagnosis of the new one. The historical event may share surface symptoms while differing in an important cause. Perhaps this time the pool is healthy and a downstream payment gateway is slow; perhaps the traffic spike is coincidental. IncidentMind’s prompt tells the model to discuss both similarity and difference, and clearly labels the likely root cause as a hypothesis.
The important distinction is between four kinds of information:
- Current facts: what the operator entered about this incident, such as HTTP 500s and high connections.
-
Historical evidence: what the retained record says happened during
INC-009. - Hypothesis: a possible explanation for today’s symptoms, such as connection-pool exhaustion.
- Confirmed cause: a conclusion supported by current evidence after investigation.
Only the first two are directly available to the model in this example. Confirmation has to come from an engineer validating against the running system.
Prompting for an investigation, not a verdict
IncidentMind asks the LLM to produce six sections: incident summary, relevant past experience, likely root cause, investigation steps, recommended action, and confidence. The relevant-experience section must identify what happened previously, compare similarities and differences, and explain why the old event matters. The likely-cause section must remain a hypothesis. The investigation section must contain exactly three concrete checks.
That structure is a practical way to guide the model toward useful work. For the Payment API scenario, checks might focus on current pool utilization and wait queues, whether request latency tracks database wait time, and whether application workers or the downstream gateway are constrained. These are examples of the kind of checks the prompt requests; the repository does not hard-code those exact three checks as a deterministic result. The LLM generates them from the current description and recalled context.
The recommended-action section asks for the safest immediate mitigation the SRE should consider, preferably reversible. INC-009 provides a prior resolution to reference, but the model must still consider whether it fits the present evidence. The application explicitly tells the LLM not to perform production changes automatically. It is an assistant that presents recommendations for human judgment, not an operations executor.
Confidence is included because an investigation needs to communicate uncertainty, not just a list of ideas. The model must choose Low, Medium, or High and explain why that level fits the available evidence. This is a qualitative model-generated assessment, not a calibrated probability or a score derived from measured outcomes. A relevant memory can strengthen a lead while missing telemetry should constrain certainty.
The hard boundary: no live telemetry
The repository describes IncidentMind as not directly accessing live production infrastructure. It does not query Prometheus or Grafana, inspect traces, or read logs. The model reasons from the incident details entered in the form and the text Hindsight recalls. Even a good historical match cannot establish that today’s connection pool is exhausted, that workers are saturated, or that a mitigation will be safe in the current environment.
That limitation is not a footnote to the reasoning; it defines the reasoning. The output is a prioritized investigation aid. An engineer must compare the hypothesis with current dashboards, logs, traces, and operational knowledge before acting. Recommendations should be evaluated in context, even when a prior incident was resolved by a similar change.
What this design teaches us
Building this flow reinforced a useful lesson: memory helps an AI investigation most when the system makes the source and status of each claim visible. Hindsight supplies relevant historical records; the prompt asks for comparison; the UI exposes the recalled text; and the response separates a likely cause from confirmation. None of those steps guarantees a correct diagnosis, but together they give an engineer a more reviewable starting point than a free-form guess.
For IncidentMind, INC-009 is valuable because it tells the agent where to look first—not because it tells the agent what must be true. The investigation still belongs to the engineer, and current evidence still has the final word. That is the difference between an AI that remembers an outage and one that lets yesterday’s outage dictate today’s answer.
Top comments (0)