<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Chakali Madhav</title>
    <description>The latest articles on DEV Community by Chakali Madhav (@chakali_madhav).</description>
    <link>https://dev.to/chakali_madhav</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4146984%2F7c3adba1-a6d8-4baa-988f-20e1d0e8e72c.png</url>
      <title>DEV Community: Chakali Madhav</title>
      <link>https://dev.to/chakali_madhav</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/chakali_madhav"/>
    <language>en</language>
    <item>
      <title>Why Hindsight Needs a Human Approval Gate</title>
      <dc:creator>Chakali Madhav</dc:creator>
      <pubDate>Mon, 28 Sep 2026 14:13:45 +0000</pubDate>
      <link>https://dev.to/chakali_madhav/why-hindsight-needs-a-human-approval-gate-1ln</link>
      <guid>https://dev.to/chakali_madhav/why-hindsight-needs-a-human-approval-gate-1ln</guid>
      <description>&lt;h1&gt;
  
  
  Why Hindsight Needs a Human Approval Gate
&lt;/h1&gt;

&lt;p&gt;An AI agent can recommend a remediation without being allowed to execute it.&lt;/p&gt;

&lt;p&gt;That distinction is central to OpsMind.&lt;/p&gt;

&lt;p&gt;The system was designed to investigate SRE incidents, use Hindsight to retrieve relevant operational experience, generate a diagnosis, recommend a runbook, and then wait for human approval before remediation.&lt;/p&gt;

&lt;p&gt;Persistent memory makes the agent more informed.&lt;/p&gt;

&lt;p&gt;It does not make the agent autonomous by default.&lt;/p&gt;

&lt;h2&gt;
  
  
  Diagnosis Is Not Execution
&lt;/h2&gt;

&lt;p&gt;OpsMind separates the incident-response workflow into two major stages:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reasoning&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;What is likely wrong, what evidence supports it, and what should be done?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Action&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Should the recommended remediation actually be executed?&lt;/p&gt;

&lt;p&gt;The first stage is performed by the AI SRE agent.&lt;/p&gt;

&lt;p&gt;The second requires explicit human approval.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1jsqrmayqebe7y82m2sb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1jsqrmayqebe7y82m2sb.png" alt=" " width="800" height="336"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 1 — The human approval gate sits between AI diagnosis and remediation.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;This separation is particularly important when historical memory is involved.&lt;/p&gt;

&lt;p&gt;A previous incident may have been successfully resolved using a particular action, but that does not automatically mean the same action is safe or appropriate for the current incident.&lt;/p&gt;

&lt;h2&gt;
  
  
  Hindsight Provides Context, Not Permission
&lt;/h2&gt;

&lt;p&gt;Hindsight can retrieve previous incident experiences.&lt;/p&gt;

&lt;p&gt;For example, suppose previous incidents show that increasing database connection pool capacity helped resolve similar Payment API failures.&lt;/p&gt;

&lt;p&gt;That historical experience can be useful.&lt;/p&gt;

&lt;p&gt;But it should not be interpreted as:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;“The system previously did this, therefore do it now.”&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Instead, it means:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;“A similar operational situation previously had this successful outcome. Check whether the current evidence supports considering the same approach.”&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This distinction is fundamental to the OpsMind design.&lt;/p&gt;

&lt;p&gt;Memory improves the information available to the agent.&lt;/p&gt;

&lt;p&gt;The human remains responsible for approving the resulting action.&lt;/p&gt;

&lt;h2&gt;
  
  
  INC-008
&lt;/h2&gt;

&lt;p&gt;INC-008 provides an example of the workflow.&lt;/p&gt;

&lt;p&gt;The Payment API showed:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;6.1-second latency&lt;/li&gt;
&lt;li&gt;26% HTTP 500 errors&lt;/li&gt;
&lt;li&gt;97% database connection utilization&lt;/li&gt;
&lt;li&gt;Requests waiting for database connections&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The AI agent analyzed the current evidence and historical context and identified database connection pool exhaustion as the likely root cause.&lt;/p&gt;

&lt;p&gt;It then generated recommended remediation steps.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxhx763ahtj5qvulbr1ni.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxhx763ahtj5qvulbr1ni.png" alt=" " width="800" height="404"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 2 — The AI diagnosis and recommended remediation are presented before execution.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;At this stage, OpsMind had &lt;strong&gt;not&lt;/strong&gt; changed a production system.&lt;/p&gt;

&lt;p&gt;The diagnosis was an explanation.&lt;/p&gt;

&lt;p&gt;The runbook was a recommendation.&lt;/p&gt;

&lt;p&gt;The next step was the approval gate.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Human Approval Gate
&lt;/h2&gt;

&lt;p&gt;OpsMind explicitly waits for approval before resolving the incident.&lt;/p&gt;

&lt;p&gt;The backend separates the analysis and resolution operations.&lt;/p&gt;

&lt;p&gt;The analysis endpoint produces a diagnosis and sets the incident to an approval-required state.&lt;/p&gt;

&lt;p&gt;The resolution operation only proceeds after the workflow receives approval.&lt;/p&gt;

&lt;p&gt;This creates a clear operational boundary:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AI recommends → Human approves → Remediation proceeds&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That boundary is simple, but it changes the risk model substantially.&lt;/p&gt;

&lt;p&gt;An engineer can inspect:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The root cause&lt;/li&gt;
&lt;li&gt;Evidence&lt;/li&gt;
&lt;li&gt;Historical incidents&lt;/li&gt;
&lt;li&gt;Recommended actions&lt;/li&gt;
&lt;li&gt;Confidence&lt;/li&gt;
&lt;li&gt;Runbook steps&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;before deciding whether to continue.&lt;/p&gt;

&lt;h2&gt;
  
  
  Simulated Remediation
&lt;/h2&gt;

&lt;p&gt;The current OpsMind workflow uses simulated remediation.&lt;/p&gt;

&lt;p&gt;The runbook is explicitly marked as simulation-only, and the system does not make changes to production infrastructure.&lt;/p&gt;

&lt;p&gt;This allows the complete incident lifecycle to be demonstrated without introducing operational risk.&lt;/p&gt;

&lt;p&gt;The resolution process can still show:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Which runbook actions would be performed&lt;/li&gt;
&lt;li&gt;Whether the simulated actions succeeded&lt;/li&gt;
&lt;li&gt;What outcome was produced&lt;/li&gt;
&lt;li&gt;Whether the experience should be retained&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The architecture therefore preserves the same logical lifecycle while keeping execution controlled.&lt;/p&gt;

&lt;h2&gt;
  
  
  Resolution Creates New Knowledge
&lt;/h2&gt;

&lt;p&gt;The approval gate is not the end of the workflow.&lt;/p&gt;

&lt;p&gt;After successful remediation, OpsMind records the outcome as a learning event.&lt;/p&gt;

&lt;p&gt;The retained information includes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;hindsight_client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;retain&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;BANK_ID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;learning_record&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;OpsMind SRE incident learning&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;metadata&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;incident_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;incident_id&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;service&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;service&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;incident_outcome&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;successful&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;successful&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The result is a feedback loop:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Diagnose → Approve → Resolve → Retain&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The retained experience can then become historical context for a future incident.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcgea0co66rcu5fampy99.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcgea0co66rcu5fampy99.png" alt=" " width="800" height="425"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 3 — The successful resolution is converted into organizational memory after remediation.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why This Matters With Memory
&lt;/h2&gt;

&lt;p&gt;Persistent memory changes the behavior of an agent over time.&lt;/p&gt;

&lt;p&gt;Suppose INC-008 establishes that a particular combination of database connection symptoms was successfully resolved through a specific remediation.&lt;/p&gt;

&lt;p&gt;Later, INC-007 is investigated.&lt;/p&gt;

&lt;p&gt;Hindsight can return INC-008 as historical context.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsmslvqd6x25abus7x0zi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsmslvqd6x25abus7x0zi.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 4 — A future investigation can retrieve the previous resolution as historical context.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;This makes the human approval gate even more important.&lt;/p&gt;

&lt;p&gt;The agent now has more historical experience available to it, but the increased amount of information does not automatically justify autonomous action.&lt;/p&gt;

&lt;p&gt;More memory should mean &lt;strong&gt;better-informed decisions&lt;/strong&gt;, not fewer controls.&lt;/p&gt;

&lt;h2&gt;
  
  
  Designing for Inspectability
&lt;/h2&gt;

&lt;p&gt;Another goal of the dashboard is to make the agent's reasoning inspectable.&lt;/p&gt;

&lt;p&gt;Instead of showing only:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Root cause: database connection pool exhaustion.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;OpsMind exposes the surrounding reasoning context.&lt;/p&gt;

&lt;p&gt;The engineer can see:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Evidence&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;What signals support the diagnosis?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Historical Context&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Which previous incidents were retrieved?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Recommended Actions&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;What does the agent propose?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Confidence&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;How strongly does the available evidence support the diagnosis?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Runbook&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;What remediation steps are configured?&lt;/p&gt;

&lt;p&gt;This makes the system easier to evaluate before approval.&lt;/p&gt;

&lt;h2&gt;
  
  
  What We Learned
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. AI recommendations and execution should be separate
&lt;/h3&gt;

&lt;p&gt;Generating an action and executing an action are different responsibilities.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Memory should not bypass controls
&lt;/h3&gt;

&lt;p&gt;A successful historical remediation does not automatically become an instruction for the next incident.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Human approval is part of the architecture
&lt;/h3&gt;

&lt;p&gt;The approval gate should not be treated as a cosmetic interface element. It is an explicit control point between reasoning and action.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Simulated execution is useful during development
&lt;/h3&gt;

&lt;p&gt;Simulation makes it possible to test the complete incident lifecycle without connecting the prototype to real production infrastructure.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Successful remediation can become future context
&lt;/h3&gt;

&lt;p&gt;Once a resolution is validated, it can be retained and made available to future investigations.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;The purpose of OpsMind is not to remove engineers from incident response.&lt;/p&gt;

&lt;p&gt;It is to give engineers better context when they have to make decisions under pressure.&lt;/p&gt;

&lt;p&gt;Hindsight helps the agent remember previous operational experiences.&lt;/p&gt;

&lt;p&gt;Current telemetry tells it what is happening now.&lt;/p&gt;

&lt;p&gt;The AI combines those inputs into a diagnosis and recommended runbook.&lt;/p&gt;

&lt;p&gt;Then the workflow stops and asks a human to approve the action.&lt;/p&gt;

&lt;p&gt;That creates a useful division of responsibility:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Memory informs.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Evidence grounds.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AI reasons.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Humans approve.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The system learns from successful outcomes.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That pattern allows persistent agent memory to improve an incident-response workflow without turning historical experience into unchecked automation.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/vectorize-io/hindsight?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;Hindsight on GitHub&lt;/a&gt;&lt;br&gt;
&lt;a href="https://hindsight.vectorize.io/?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;Hindsight Documentation&lt;/a&gt;&lt;br&gt;
&lt;a href="https://github.com/srivaniyadav174/OpsMind-AI-SRE-Incident-Response-Memory-Agent.git?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;OpsMind — GitHub Repository&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>devops</category>
      <category>sre</category>
    </item>
  </channel>
</rss>
