Most Kubernetes AI agents are a rootkit waiting to happen. They demand cluster-admin rights and stream raw stdout to external LLM APIs. SREK3S is a zero-trust, read-only incident response agent. It intercepts pod crashes, scrubs secrets in-memory before network egress, and sandboxes LLM triage in a POSIX-jailed worker. It generates verified GitOps patches with strictly zero cluster write authority and deterministically fails closed to human review.
Features
-
Zero cluster mutations: namespaced
get/list/watchonly. No ClusterRole, no write field on either wire contract, no mounted token on the Agent. - In-memory secret scrubbing: 11 ordered regex rules run before egress. Nothing unmasked reaches a queue, a disk, or a socket.
-
AST YAML validation: a patch survives a YAML AST parse, then
git apply --checkagainst the target's own bytes. - Deterministic Tier-2 escalation: tier, patch, and every flag are computed before a model is…
Most "AI SRE agents" are catastrophic control failures dressed as features. By binding service accounts to cluster-admin and mounting raw credentials, they turn container stdout—which routinely leaks AWS keys and Postgres passwords—into an exfiltration pipeline to third-party LLMs. Worse, granting the model write authority turns a log-line prompt injection into a full cluster mutation requiring zero vulnerabilities in the model itself.
I built SREK3S to take the opposite approach: remove the capability entirely. SREK3S is a zero-trust, read-only AI incident responder. Because it holds no cluster write authority anywhere in its architecture, the question of whether the model would do something dangerous never arises—it physically can't.
What SREK3S Fixes
Instead of relying on prompt hardening, SREK3S enforces strict, structural security boundaries:
- In-Memory Redaction: No secret crosses the network egress boundary. The Go Sentinel uses a hermetic package to scrub secrets in-memory, on the node, before any socket is opened.
- Ordered, Normative Rules: Redaction follows 11 strict rules (targeting PEM blocks, JWTs, AWS keys), executing the cross-line pass first to prevent partial unmasking.
- Zero-Write RBAC: The Sentinel is bound to a namespaced
Rolelimited strictly toget,list, andwatch. There is noClusterRoleor binding. - Credential-Free Agent: The Python Agent runs with
automountServiceAccountToken: falseas UID 10001 with a read-only root filesystem and all capabilities dropped.

Live execution: The Sentinel intercepts an AWS Secret Access Key in a crashing pod and scrubs it in-memory before it ever hits the network.
The Rigorous Engineering (Not Just Another Vibecoded Wrapper)
Instead of blindly piping model hallucinations to kubectl apply, SREK3S treats LLM output with extreme suspicion.
- YAML AST Validation: The
verify_yaml_astfunction parses original and patched documents into Abstract Syntax Trees to prove exactly one semantic field changed, catching type errors simple diffs miss. - True GitOps Failsafes: It runs
git apply --checkin a materialized throwaway repo against the exact bytes provided, refusing to normalize or repair malformed output. - Disposable POSIX Sandboxing: Analysis executes in a child process bound by strict limits (256 MiB memory, 1 CPU-second, 0 core dumps). No state survives the investigation.
- Fail-Closed Law: Incident routing is deterministic. If evidence is ambiguous (e.g., a generic
CrashLoopBackOff), SREK3S refuses to guess, defaulting to a Tier-2 architectural review that dispatches a Markdown war room report without proposing a patch.
I Threw Everything At It: The Verification Gauntlet
I didn't just write happy-path tests; I tried to break this system in every way imaginable. The codebase passes every gate I could throw at it (go vet, gofmt, -race, black, flake8, mypy, pytest), currently sitting at 1087 passing tests (179 in Go).
- The Scrubber Corpus: The memory scrubber is validated against a 46-case corpus spanning 8 groups. This includes 32 maskable secrets and a dedicated
negative_controlsgroup of 6 cases that must survive untouched—because a redaction tool that just blanks out everything is completely useless. - Performance Bounding: The system enforces a throughput budget of ≥20,000 lines/sec/core. This ensures the masking stays well inside the 2-second detection budget, with rules compiled exactly once and never recompiled in the hot path.
- Prose-to-Code Assertions: I built a build gate (
test_scrubber_manifest_spec.py) that parses the normative rule table straight out ofCONTRIBUTING.mdand fails the build if the code's rule IDs, order, or patterns drift from the documentation. - RBAC Enforcement:
TestSentinelRoleGrantsNoMutatingVerbparses the deployment YAML and fails the build if any verb other thanget,list, orwatchever sneaks into the Sentinel's Role. - Catching Real Defects: I document my failures in
lessons-learned.md. Exhaustive testing caught edge cases like aCrashLoopfixture rendered undetectable byrestartPolicy: Never, a Service selector that routed nowhere despite 16 green manifest tests, and aterminationGracePeriodSecondsblock mistakenly placed at the container level—which text-matching tests missed, but the actual API server rejected.
SREK3S extracts the genuinely useful capabilities of LLMs—reading crash evidence and forming hypotheses—without handing over the keys to the cluster.
Try it in 60 Seconds
SREK3S is open-source and MIT-licensed. Because it relies on standard client-go informers, it deploys cleanly to any conformant cluster (k8s, k3s, Minikube, EKS).
You do not need to build Go binaries or compile Python to test it. We publish multi-arch images directly to GHCR.
Clone the repo, apply the quickstart overlay, and detonate the provided memory-leak fixture to watch the in-memory redaction happen live:
Top comments (0)