DEV Community

Cover image for Kubernetes knows who can change production. It doesn't know when nobody should.
houssem kraoua
houssem kraoua

Posted on

Kubernetes knows who can change production. It doesn't know when nobody should.

Monday, 9am. First coffee. First Slack message of the week: "Who changed prod this weekend?"

Silence.

There was a freeze. Everyone agreed on it. It was pinned in the channel. A pipeline ran anyway, because pipelines don't read Slack.
Kubernetes knows who can change production. It has no idea when nobody should. I've spent three years building around that gap.

A freeze that covers one door isn't a freeze

Teams patch the gap with what they have: a pinned message, a calendar, an exit 1 at the top of a CI job, Argo CD sync windows. Each one covers a single path. Sync windows are great if Argo is the only thing touching your cluster, but they only govern Argo's own syncs.

Production has more doors than that. An admin kubeconfig on someone's laptop. A second pipeline nobody remembers. A helm upgrade run by hand "to fix one value".
All of those doors lead to the same room: the API server. Every write passes through admission before it's stored, and that's the only layer where a freeze actually holds.

How Telark enforces a freeze

You don't write policy YAML. You create a protection plan: a scope (applications or namespaces), what to block, a time window, and audit or enforce mode. Nine templates cover the usual suspects, from deletions and image tags to replica scaling, storage, and ConfigMap and Secret changes.
Applications are grouped automatically from labels like app.kubernetes.io/name, so you protect checkout, not seventeen Deployments.

When the window opens, Telark renders the plan into namespaced Kyverno policies in every namespace of the scope. Kyverno refuses matching requests at admission, whether they come from kubectl, CI or a deploy tool:
Replica scaling is blocked by protection plan "<plan-id>".

When the window closes, the policies are removed. Nobody has to remember on Friday.

The admission path, honestly

Put something in production's admission path and it becomes part of production. So here's exactly how it behaves.
Telark isn't in the request path. Kyverno is. Telark deploys the policies and watches them. Every 30 seconds it compares the live policies with what the plan expects, flags the plan as Degraded or Drifted, and redeploys anything that was edited or deleted.

It fails open by default. The bundled Kyverno runs with failurePolicy: Ignore, so if its webhook is down, requests get through, even for enforcing plans. A dead policy engine shouldn't make your workloads undeployable. If you'd rather fail closed, set app.kyverno.failOpen=false and keep Kyverno's admission replicas and disruption budget.

It never blocks what keeps workloads alive. Status writes from kubelet and controllers are never matched, because blocking them is an outage, not a freeze. Pods, ReplicaSets and PVCs created by the controller manager are exempt, so a pod that gets OOM-killed mid-freeze is still replaced.

Subresources are where naive policies leak. A wildcard "block all updates" rule doesn't match subresources, so kubectl scale walks right past it. Telark matches the scale subresources explicitly. The flip side: the replica-scaling template also stops your HPA, because exempting controller identities let kubectl autoscale slip through. Pick that template knowingly.

The emergency exit is cancelling the plan. Deleting the Kyverno policy by hand won't stick: Telark sees drift and puts it back. Cancelling removes the policies.

What you get around the freeze

Audit mode first. Nothing is blocked; the message just says "would be blocked". Telark keeps a ledger per plan and writes a report when the window ends.

Approvals. Plans tagged Production always need approval, and whoever requested a plan can't approve it. Roles support deny rules, so someone can own plans without being allowed to approve them.

Field-by-field history. Every change to an application is recorded field by field, with a snapshot of the manifests from before. Rollback applies a snapshot, and the rollback is itself recorded, so you can undo it too.

Insights. When an app degrades, deterministic rules write a card per affected workload: crash loop, OOM, image pull, stuck rollout, or a regression right after a config change, with the evidence. A small model running in-cluster through Ollama only rewrites the wording, and any rewrite that drops or invents a fact is thrown away. You can switch it off in Settings and nothing else changes.

Limits, and the license

It's not a backup tool and not a GitOps replacement. Keep Velero, keep Argo or Flux; Telark sits underneath them.
v0.1.0 is early. The API is v1alpha1 and may change, so pin the chart version. It's one cluster per install, and the chart bundles Kyverno, Redis, NATS, metrics-server and Ollama, though you can bring your own Kyverno. It needs Kubernetes 1.30+, with 1.33+ as the tested target.
Telark is source-available under the Elastic License 2.0, not OSI open source. You can use, modify and self-host it. You can't offer it to others as a managed service.
Try it

helm install telark oci://ghcr.io/telark/charts/telark -n telark --create-namespace \
--set app.auth.bootstrap.admin=you@example.com

Already running Kyverno? Add --set app.kyverno.enabled=false. The getting started guide takes you to your first plan in about ten minutes.

Then do one thing: put a plan in audit mode on something you care about, leave it for a week, and read the report. You'll either find a change you didn't know was happening, or a place where Telark gets in the way. I want to hear about both, and especially why you'd never put this in your admission path.

telark.io · GitHub

Top comments (0)