DEV Community

NASANI RAGAMALA
NASANI RAGAMALA

Posted on

I Put a Hard Boundary Between Investigation and Remediation

The most important backend decision in an incident-response system is deciding what the agent is allowed to do.

I did not want an investigation endpoint that could quietly turn into a production execution endpoint. The backend needs to support detailed analysis while keeping operational authority behind an explicit boundary.

That led me to separate four concerns: incident data, investigation, memory, and remediation.

Start with the API Contract

The frontend interacts with the backend through distinct operations such as:

GET /api/dashboard/stats
GET /api/incidents
GET /api/memory
GET /api/runbooks

Each operation has a clear responsibility. Investigation is responsible for analysis, similar incidents provide historical context, memory exposes persistent knowledge, and runbooks provide reviewed operational procedures. Resolution records the verified outcome of an incident.

The important distinction is that none of these operations means that the model can execute arbitrary production commands.

The backend can help an engineer understand an incident and recommend what to do without giving the model direct operational authority.

Investigation Should Be Read-Heavy

An investigation needs to bring together evidence from several sources.

The investigation contract reflects this by returning information such as a root-cause synthesis, supporting evidence, similar incidents, a recommended action, and a relevant runbook.

This gives the backend a clear implementation target while keeping the frontend independent of how the analysis is actually performed.

The frontend does not need to know whether the backend retrieved historical incidents from a database, used Hindsight, or called a particular model provider. It only needs a structured investigation result.

Hindsight Belongs Behind the Memory Boundary

I wanted Hindsight to be an infrastructure dependency rather than something every API handler understands directly.

For example, the frontend interacts with memory through a simple API function:

export async function getMemoryItems(params = {}) {
const response = await apiClient.get('/api/memory', { params });
return response.data;
}

The frontend sees memory as a resource. The backend decides how that resource maps to Hindsight operations.

That separation keeps the application from becoming tightly coupled to the implementation details of a particular memory provider.

Hindsight provides the underlying memory capabilities, while the application treats memory as a service boundary.

Resolution Is Where the Learning Loop Closes

The most interesting endpoint is probably not investigation. It is resolution.

When an engineer resolves an incident, the client sends structured outcome information to the backend:

export async function resolveIncident(id, data = {}) {
const response = await apiClient.post(
/api/incidents/${id}/resolve,
data
);

return response.data;
}

The backend can then close the incident and retain the verified outcome.

The overall process is straightforward: an investigation gathers current evidence and relevant historical context, the reasoning layer produces a recommendation, an engineer reviews that recommendation, and the verified resolution is recorded afterward.

This is cleaner than allowing the LLM to write arbitrary memory whenever it generates an answer.

The engineer's verified outcome is stronger memory because it represents what actually happened rather than what the model predicted.

Why I Separated Runbooks from Memory

A historical incident is evidence. A runbook is a procedure. They are not interchangeable.

Suppose Hindsight returns an earlier incident where a configuration change helped resolve connection saturation. That information can be useful during the investigation because it provides historical context.

It does not automatically mean that the same command should be executed on the current system.

The backend can retrieve a runbook separately through an endpoint such as:

GET /api/runbooks/{id}

The investigation can then use the historical incident as evidence and the runbook as the reviewed procedure.

This allows the system to make a distinction between learning from what happened before and following an approved operational process.

That distinction is important because retrieved memory should inform a recommendation, not become an executable instruction.

The LLM Should Also Sit Behind a Service Boundary

The application configuration separates the AI provider and model:

export const APP_CONFIG = {
aiProvider: 'Groq',
defaultAiModel: 'openai/gpt-oss-120b',
confidenceThreshold: 80,
};

The broader settings model also accounts for providers such as Groq, OpenAI, and Anthropic.

That means the backend should treat the model provider as replaceable.

The investigation service should receive structured context and return structured analysis. Incident routes should not need to understand provider-specific implementation details.

This separation becomes useful when the model changes, a provider becomes unavailable, or the system needs to support another provider later.

Failure Should Be Explicit

Another important backend decision is making failures visible.

An investigation can fail for several different reasons. The incident might not exist, telemetry might be incomplete, memory retrieval might not provide useful context, the LLM provider might fail, a runbook might be unavailable, or the model might return an unusable result.

These cases should not all become a generic "AI failed" message.

The backend should preserve enough state for the frontend to tell the engineer which stage of the investigation failed.

This is one reason the six-stage investigation model is useful. It provides a structure for tracking where an investigation currently stands and where a failure occurred.

Human Approval Belongs Outside the Model Call

The backend can generate a recommendation, but approval should remain a separate state transition.

The frontend explicitly presents the result as an AI recommendation that requires human review. The backend can represent this distinction through states such as investigating, recommendation ready, awaiting approval, approved, remediated, and resolved.

The exact persistence model can evolve as the system grows, but the boundary should remain clear.

An LLM should not be able to move an incident directly from investigation to resolution.

The model provides analysis. The engineer makes the operational decision.

Hindsight Makes the Backend Stateful in a Useful Way

A conventional incident API mainly manages the current state of an incident.

Adding Hindsight introduces another dimension: historical knowledge that can influence future investigations.

A resolved incident can therefore contribute information to a later investigation. Previous root causes, successful resolutions, failed approaches, and engineer notes can provide additional context when a similar problem occurs again.

This creates a practical learning loop. The system investigates the current incident, records the verified outcome, and makes that experience available as historical context for future investigations.

The result is more than an incident database. It is a backend where the outcome of one investigation can become useful input for the next one.

That is the boundary I wanted from the beginning: the AI can investigate aggressively, retrieve relevant experience, and recommend an action, while production remediation and verified outcomes remain explicit backend operations controlled by the engineering workflow.

Top comments (0)