DEV Community

Cover image for Automating Kubernetes Remediation Workflows Without Human Intervention
Hema Shankar
Hema Shankar

Posted on

Automating Kubernetes Remediation Workflows Without Human Intervention

Diagnosing an incident is one thing. Executing the fix on a live Kubernetes cluster is terrifying.

If your autonomous agent deletes the wrong pod or scales the wrong deployment, you have turned a single incident into a full outage.

We built a recovery pipeline that patches Kubernetes Deployments safely, using hardcoded limits and a policy engine that can veto dangerous actions.

The problem

Our first version let the LLM decide which kubectl command to run.

During testing, it attempted to delete all pods in the kube-system namespace.

We killed the process and immediately added guardrails.

That incident taught us the real lesson:

The AI is not the dangerous part. The lack of guardrails is.

The fix — three layers of defense

Layer 1: The policy engine vetoes critical targets

platform/policy-engine/main.py rejects any action on hardcoded critical infrastructure:

# platform/policy-engine/main.py

critical_targets = [
    "api-gateway",
    "postgres-db",
    "redis-master"
]

if req.target in critical_targets:
    logger.warning(
        "policy_rejected_critical_target",
        target=req.target
    )

    return {
        "authorized": False,
        "reason": "Target is marked as CRITICAL."
    }

return {
    "authorized": True,
    "reason": "Target is non-critical, action approved."
}
Enter fullscreen mode Exit fullscreen mode

The policy engine sits between the AI's proposed action and the Kubernetes execution layer.

The LLM can propose an action, but it does not get to authorize that action.


Layer 2: The execution engine handles Kubernetes failures explicitly

platform/execution-engine/main.py wraps the Kubernetes controller request in a try/except so a local API failure does not leave the incident pipeline hanging:

try:
    k8s_res = await client.post(
        "http://kubernetes-controller.incident-agent-system.svc.cluster.local/execute",
        json={
            "target": target,
            "workflow": plan.get("workflow", [])
        },
    )

    k8s_res.raise_for_status()

    logger.info(
        "k8s_execution_complete",
        incident_id=incident_id
    )

except Exception as e:
    logger.error(
        "k8s_execution_failed",
        incident_id=incident_id,
        error=str(e)
    )

    # Publish a failure event so the incident
    # workflow can continue and update the UI.
    publish_event(
        "RECOVERY_FAILED",
        incident_id=incident_id
    )
Enter fullscreen mode Exit fullscreen mode

The important part here is that a Kubernetes failure does not become a fake recovery success.

The system records the failure and allows the incident workflow to continue.


Layer 3: The Kubernetes controller uses a valid annotation key

platform/kubernetes-controller/main.py patches the Deployment and triggers a rollout using a valid annotation key.

We initially used an invalid annotation prefix containing underscores. Kubernetes rejected the annotation because the prefix must follow DNS-subdomain syntax.

The corrected version uses:

# Valid annotation prefix
deployment.spec.template.metadata.annotations[
    "hindsight-agent.aiops/restartedAt"
] = "now"

# Ensure the deployment has at least one replica
if deployment.spec.replicas == 0:
    deployment.spec.replicas = 1

v1_apps.patch_namespaced_deployment(
    name=cmd.target,
    namespace="hindsight-agent-apps",
    body=deployment
)
Enter fullscreen mode Exit fullscreen mode


This gives the controller a deterministic way to patch the Deployment and trigger the recovery workflow.

The full recovery pipeline

When an incident fires, the execution engine follows a controlled workflow:

  1. Checks the policy engine for authorization
  2. Calls the Kubernetes controller to patch the Deployment
  3. Calls the recovery-verification engine to confirm health
  4. Publishes a RECOVERY_COMPLETED or RECOVERY_FAILED event

The important part is that every stage produces an explicit result.

No silent hangs.

No stuck incidents.

No pretending a failed Kubernetes operation succeeded.

Before vs after

The following results are from our test scenarios:

Metric Before After
Manual intervention time 15 min 0 min
Dangerous actions attempted 3 0
Recovery success rate 72% 98%

These numbers represent our project test results, not universal Kubernetes performance benchmarks.

The honest lesson

Safe autonomous execution requires hardcoded limits and policy-engine fail-safes.

Never let an LLM directly execute unrestricted destructive Kubernetes commands against a production cluster.

The architecture should separate two responsibilities:

The AI decides what to do.

The policy engine decides whether it is allowed.

These are two different jobs.

Keeping them separate gives you a much stronger security and reliability boundary.

The scariest part of autonomous SRE is not necessarily the AI.

It is the permissions you give it.

Why this matters

As the memory bank grows and the agent recalls proven playbooks, the recovery pipeline becomes a deterministic execution layer rather than a generative guess.

Combined with agent memory, every successful recovery can be retained as a validated runbook.

The system learns from previous incidents while the guardrails remain in place.

You can explore the core concepts and implement this yourself by checking out the Hindsight documentation.

The Hindsight GitHub repository also shows how the memory layer can integrate with autonomous agents.

Final takeaway

Autonomous Kubernetes remediation should not mean giving an LLM unrestricted access to your cluster.

A safer architecture looks like this:

Incident
   ↓
AI Agent
   ↓
Proposed Action
   ↓
Policy Engine
   ↓
Authorized?
  ↙     ↘
 No      Yes
 ↓        ↓
Reject   Execution Engine
            ↓
      Kubernetes Controller
            ↓
      Recovery Verification
            ↓
       Recovery Event
            ↓
        Agent Memory
Enter fullscreen mode Exit fullscreen mode

The goal is not to remove control.

The goal is to automate execution without removing the safety boundaries around it.

Top comments (0)