Building an Incident Assistant That Recommends Mitigations but Leaves Production Control to Engineers
A production incident is exactly when a confident-sounding suggestion can do the most damage. If payment requests are failing, an agent that can restart workers or change connection-pool settings might appear helpful—but acting on an unverified diagnosis can turn an outage into a larger one.
IncidentMind takes a more restrained role. It helps an engineer organize the incident, recall related experience, and form a plan for what to investigate and what mitigation to consider. It does not execute remediation against production. That boundary is a design choice visible in the prompts and workflow: the AI can recommend, while the engineer remains responsible for checking the evidence and deciding what to do.
Autonomous remediation is a different problem
To safely make a production change, an agent would need more than a plausible explanation. It would need reliable access to current system state, a way to assess the blast radius of an action, clear authorization, and safeguards around execution and rollback. IncidentMind’s repository does not implement those capabilities. There is no connection to production metrics, logs, traces, or infrastructure controls, and there is no remediation executor.
Instead, the app accepts an engineer’s description of the current incident and returns an investigation. Its prompt specifically tells the model not to pretend to have real-time telemetry and not to automatically perform production changes. This keeps the generated response in the realm of decision support rather than control.
From observation to a hypothesis
The Streamlit form captures service, environment, error, and symptoms. Those are observations as supplied by the user. For the sample Payment API scenario, the report might say that HTTP 500 errors are occurring during payment processing, database connections are unusually high, and the issue began after a traffic spike.
IncidentMind searches Hindsight for related historical incidents and includes returned memories in the LLM context. The model is asked to produce six sections, starting with a concise incident summary and a discussion of relevant past experience. That past experience is supporting context, not a statement about what is happening right now.
The distinction is important. The sample history includes INC-009: a Payment API incident with HTTP 500s and high database connections after a traffic spike. Its recorded cause was connection-pool exhaustion, and its resolution involved increasing pool capacity and scaling application workers. That is historical evidence. It makes a similar failure mode worth considering, but it does not establish the current root cause.
IncidentMind’s prompt says this directly:
- Historical incidents are evidence, not proof.
- Do NOT claim the root cause is confirmed unless evidence
supports it.
- Clearly distinguish historical facts from your current
hypothesis.
- Explain similarities and differences.
- Do not pretend to have real-time production telemetry.
The “Likely Root Cause” section is therefore framed as a hypothesis. In the example, connection-pool exhaustion might be a reasonable lead. A downstream payment-gateway delay or another database issue could still explain similar symptoms. The engineer needs current evidence to decide which explanation holds.
Turn a guess into checks
Rather than stopping at a cause label, the prompt asks for exactly three practical investigation steps. This matters because an investigation should create a route from observation toward verification. For the Payment API case, an LLM might suggest checking pool utilization and wait queues, comparing request latency with database wait time, or inspecting worker saturation and downstream gateway latency.
Those are illustrative examples, not fixed outputs guaranteed by the code. The model generates the checks from the incident description and recalled memories. The repository does not query the systems where an engineer would perform them. They are prompts for human investigation, and their usefulness depends on the inputs and model response.
Recommendations are proposals, not commands
The “Recommended Action” section asks for the safest immediate mitigation an SRE should consider, with reversible actions preferred where possible. In the Payment API scenario, the previous incident’s pool-capacity increase and worker scaling may inform what to evaluate. They should not be copied blindly into production: the current pool could be healthy, and scaling workers could increase database pressure if connection usage is already the bottleneck.
The language “should consider” is deliberate. An AI-generated recommendation is a proposal that needs operational context: current load, service limits, deployment state, rollback options, and the team’s incident procedures. IncidentMind does not apply the recommendation. The engineer decides whether it is appropriate, obtains whatever approvals their environment requires, and performs any action through their existing tools.
This division keeps five concepts separate:
- Observation: what the user reports, such as HTTP 500s and high database connections.
-
Historical evidence: what a prior retained incident records, such as
INC-009. - Hypothesis: a possible current cause, such as pool exhaustion.
- Recommendation: a mitigation the engineer may consider after checking evidence.
- Confirmed resolution: what actually fixed the incident after the team validates the outcome.
Confusing those categories is a common path to overconfidence. The AI has access to the first two in this workflow and generates the next two. The last one requires real-world investigation and confirmation.
Confidence should communicate uncertainty
IncidentMind’s final section asks the model to choose Low, Medium, or High and explain why. A confidence statement can help the engineer understand whether the analysis rests on a close historical match, limited symptom detail, or uncertain evidence.
It is qualitative, though. The application does not calculate a calibrated probability, compare predictions with a labeled dataset, or tie confidence to measured retrieval quality. It is a model-generated explanation of certainty, not a guarantee. A High label cannot replace checking telemetry, and a Low label does not necessarily mean the investigation steps are useless.
“Teach IncidentMind” still has a human boundary
Once the analysis is displayed, the interface invites the engineer to store it in Hindsight under a “Teach IncidentMind” heading. Its copy says: “Once this incident has been reviewed and resolved, store it in Hindsight so a future incident can learn from this experience.” The button then stores the entered incident details and AI investigation text.
That wording describes the intended human workflow, but the current code does not enforce a confirmed resolution field or require the engineer to edit the AI analysis before storing. It tracks whether the record was stored in the current Streamlit session; it does not verify that the incident is actually resolved. This is an important distinction: human review is encouraged by the interface, but validated outcome capture is not yet a structured gate.
Consequently, a retained record can contain user-provided symptoms plus generated analysis without a separately recorded confirmed cause and resolution. That limits how confidently future investigations should interpret it. A useful next step would be to collect the confirmed root cause, remediation performed, and outcome as separate reviewed fields, clearly distinguished from the model’s original hypothesis.
The Payment API example, with an engineer in control
Imagine a new Payment API event: HTTP 500 errors during checkout, database connections at an unusually high level, and a recent traffic spike. IncidentMind can recall INC-009 and tell the model to compare the two incidents. It may hypothesize pool exhaustion, recommend checking connection waits and worker capacity, and suggest a reversible mitigation for the SRE to evaluate. It can also explain its confidence in that assessment.
What it cannot do is inspect the live database pool, verify that the same failure mechanism is present, increase capacity, or confirm that a mitigation worked. Those responsibilities stay with the engineer and the team’s operational systems. The AI contributes a memory-informed starting point; it does not take ownership of the incident.
The engineering lesson
The most important part of this design is not just that an AI can produce a recommendation. It is that the system’s boundaries are visible: historical context is labeled, causes are presented as hypotheses, investigation steps ask for verification, recommendations are phrased for consideration, and the interface leaves production action to people.
There is room to improve the boundary further. IncidentMind has no live telemetry integrations, confidence is not calibrated, the “reviewed and resolved” condition is not enforced by code, and the stored analysis is not separated from a confirmed outcome. Those limitations should guide how the tool is used and what future workflow changes matter most.
Incident response needs useful suggestions under pressure, but it also needs someone accountable for deciding what happens next. IncidentMind is built to help engineers think with the team’s incident history while leaving production control where it belongs: with engineers who can see and verify the system in front of them.



Top comments (0)