Why Hindsight Needs a Human Approval Gate
An AI agent can recommend a remediation without being allowed to execute it.
That distinction is central to OpsMind.
The system was designed to investigate SRE incidents, use Hindsight to retrieve relevant operational experience, generate a diagnosis, recommend a runbook, and then wait for human approval before remediation.
Persistent memory makes the agent more informed.
It does not make the agent autonomous by default.
Diagnosis Is Not Execution
OpsMind separates the incident-response workflow into two major stages:
Reasoning
What is likely wrong, what evidence supports it, and what should be done?
Action
Should the recommended remediation actually be executed?
The first stage is performed by the AI SRE agent.
The second requires explicit human approval.
Figure 1 — The human approval gate sits between AI diagnosis and remediation.
This separation is particularly important when historical memory is involved.
A previous incident may have been successfully resolved using a particular action, but that does not automatically mean the same action is safe or appropriate for the current incident.
Hindsight Provides Context, Not Permission
Hindsight can retrieve previous incident experiences.
For example, suppose previous incidents show that increasing database connection pool capacity helped resolve similar Payment API failures.
That historical experience can be useful.
But it should not be interpreted as:
“The system previously did this, therefore do it now.”
Instead, it means:
“A similar operational situation previously had this successful outcome. Check whether the current evidence supports considering the same approach.”
This distinction is fundamental to the OpsMind design.
Memory improves the information available to the agent.
The human remains responsible for approving the resulting action.
INC-008
INC-008 provides an example of the workflow.
The Payment API showed:
- 6.1-second latency
- 26% HTTP 500 errors
- 97% database connection utilization
- Requests waiting for database connections
The AI agent analyzed the current evidence and historical context and identified database connection pool exhaustion as the likely root cause.
It then generated recommended remediation steps.
Figure 2 — The AI diagnosis and recommended remediation are presented before execution.
At this stage, OpsMind had not changed a production system.
The diagnosis was an explanation.
The runbook was a recommendation.
The next step was the approval gate.
The Human Approval Gate
OpsMind explicitly waits for approval before resolving the incident.
The backend separates the analysis and resolution operations.
The analysis endpoint produces a diagnosis and sets the incident to an approval-required state.
The resolution operation only proceeds after the workflow receives approval.
This creates a clear operational boundary:
AI recommends → Human approves → Remediation proceeds
That boundary is simple, but it changes the risk model substantially.
An engineer can inspect:
- The root cause
- Evidence
- Historical incidents
- Recommended actions
- Confidence
- Runbook steps
before deciding whether to continue.
Simulated Remediation
The current OpsMind workflow uses simulated remediation.
The runbook is explicitly marked as simulation-only, and the system does not make changes to production infrastructure.
This allows the complete incident lifecycle to be demonstrated without introducing operational risk.
The resolution process can still show:
- Which runbook actions would be performed
- Whether the simulated actions succeeded
- What outcome was produced
- Whether the experience should be retained
The architecture therefore preserves the same logical lifecycle while keeping execution controlled.
Resolution Creates New Knowledge
The approval gate is not the end of the workflow.
After successful remediation, OpsMind records the outcome as a learning event.
The retained information includes:
hindsight_client.retain(
bank_id=BANK_ID,
content=learning_record,
context="OpsMind SRE incident learning",
metadata={
"incident_id": str(incident_id),
"service": str(service),
"type": "incident_outcome",
"successful": str(successful).lower(),
},
)
The result is a feedback loop:
Diagnose → Approve → Resolve → Retain
The retained experience can then become historical context for a future incident.
Figure 3 — The successful resolution is converted into organizational memory after remediation.
Why This Matters With Memory
Persistent memory changes the behavior of an agent over time.
Suppose INC-008 establishes that a particular combination of database connection symptoms was successfully resolved through a specific remediation.
Later, INC-007 is investigated.
Hindsight can return INC-008 as historical context.
Figure 4 — A future investigation can retrieve the previous resolution as historical context.
This makes the human approval gate even more important.
The agent now has more historical experience available to it, but the increased amount of information does not automatically justify autonomous action.
More memory should mean better-informed decisions, not fewer controls.
Designing for Inspectability
Another goal of the dashboard is to make the agent's reasoning inspectable.
Instead of showing only:
Root cause: database connection pool exhaustion.
OpsMind exposes the surrounding reasoning context.
The engineer can see:
Evidence
What signals support the diagnosis?
Historical Context
Which previous incidents were retrieved?
Recommended Actions
What does the agent propose?
Confidence
How strongly does the available evidence support the diagnosis?
Runbook
What remediation steps are configured?
This makes the system easier to evaluate before approval.
What We Learned
1. AI recommendations and execution should be separate
Generating an action and executing an action are different responsibilities.
2. Memory should not bypass controls
A successful historical remediation does not automatically become an instruction for the next incident.
3. Human approval is part of the architecture
The approval gate should not be treated as a cosmetic interface element. It is an explicit control point between reasoning and action.
4. Simulated execution is useful during development
Simulation makes it possible to test the complete incident lifecycle without connecting the prototype to real production infrastructure.
5. Successful remediation can become future context
Once a resolution is validated, it can be retained and made available to future investigations.
Conclusion
The purpose of OpsMind is not to remove engineers from incident response.
It is to give engineers better context when they have to make decisions under pressure.
Hindsight helps the agent remember previous operational experiences.
Current telemetry tells it what is happening now.
The AI combines those inputs into a diagnosis and recommended runbook.
Then the workflow stops and asks a human to approve the action.
That creates a useful division of responsibility:
Memory informs.
Evidence grounds.
AI reasons.
Humans approve.
The system learns from successful outcomes.
That pattern allows persistent agent memory to improve an incident-response workflow without turning historical experience into unchecked automation.
Hindsight on GitHub
Hindsight Documentation
OpsMind — GitHub Repository




Top comments (0)