💡 Originally published on devtocash.com — where this guide stays updated. I write hands-on DevOps/SRE deep-dives there weekly.
What This Agent Does
A TLS certificate expiry agent finds every certificate in and around your cluster that will expire soon, works out why it has not renewed, and either fixes it or hands a specific owner a specific task. The inventory is deterministic code: it reads cert-manager's metrics, parses every kubernetes.io/tls Secret, and connects to each Ingress host to see which certificate is actually being served. The LLM only sees the shortlist of certificates that are inside the renewal window and still not renewed, and for each one it picks from a small set of diagnoses and remediations. Private keys never leave the process, and the agent cannot create, delete, or edit a certificate by hand. It can re-trigger a cert-manager renewal, restart a workload that is serving a stale cert, or open a pull request.
The reason to build it is that "we run cert-manager" is not the same as "our certificates renew". Every expiry outage I have seen in the last three years happened on a cluster that had cert-manager installed.
Why Certificates Still Expire With cert-manager Installed
cert-manager renews a Certificate at two thirds of its lifetime by default, so a 90-day Let's Encrypt cert enters its renewal window 30 days before expiry. That leaves a long window in which a renewal can be failing quietly. The failures fall into five classes, and each has a different fix.
| Class | How it shows up | Who fixes it |
|---|---|---|
| ACME challenge failing |
Certificate has Ready=False, a Challenge sits in pending with a reason string |
Usually an Ingress or DNS config fix in Git |
| Issuer broken or rate-limited |
Order is errored, message mentions urn:ietf:params:acme:error:rateLimited or the issuer is Ready=False
|
Wait, or change the issuer |
| Unmanaged Secret | A kubernetes.io/tls Secret with no cert-manager.io/certificate-name annotation, created by hand or by a Helm chart, expiring on a date nobody wrote down |
The owning team, ideally by adopting cert-manager |
| Renewed but not reloaded | The Secret holds a fresh cert, the pod is still serving the old one because it read the file at startup | A rollout restart |
| Outside cert-manager's view | Webhook caBundles, kubelet serving certs, a load balancer terminating TLS with its own cert |
Different tooling per case; the agent reports, never touches |
The first two classes are visible in cert-manager's own objects. The third and fourth are not, and they are the ones that page you, because nothing was watching them. A stale cert on an ingress backend also produces the same symptoms as a bad upstream, so if you are chasing intermittent errors at the edge, the ingress-nginx 502 guide covers the non-TLS causes and this agent covers the TLS ones.
Step 1: The Inventory Is Code, Not a Prompt
Three sources, merged by Secret name. The first is cert-manager's metrics, which the read-only Prometheus MCP server can answer:
# Certificates expiring within 14 days
(certmanager_certificate_expiration_timestamp_seconds - time()) / 86400 < 14
# Certificates cert-manager itself says are not ready
certmanager_certificate_ready_status{condition="False"} == 1
Metrics only cover certificates cert-manager manages, so the second source walks every TLS Secret and parses the leaf certificate locally:
# inventory.py — deterministic, read-only, keys never leave this process
import base64, ssl, socket
from datetime import datetime, timezone
from cryptography import x509
from kubernetes import client, config
config.load_incluster_config()
core, net = client.CoreV1Api(), client.NetworkingV1Api()
def parse_leaf(pem: bytes) -> x509.Certificate:
return x509.load_pem_x509_certificate(pem.split(b"-----END CERTIFICATE-----")[0]
+ b"-----END CERTIFICATE-----\n")
def secret_inventory() -> dict:
out = {}
for s in core.list_secret_for_all_namespaces(field_selector="type=kubernetes.io/tls").items:
crt = s.data.get("tls.crt")
if not crt:
continue
leaf = parse_leaf(base64.b64decode(crt))
ann = s.metadata.annotations or {}
key = f"{s.metadata.namespace}/{s.metadata.name}"
out[key] = {
"secret": key,
"managed_by": ann.get("cert-manager.io/certificate-name"), # None = unmanaged
"issuer": ann.get("cert-manager.io/issuer-name"),
"not_after": leaf.not_valid_after_utc.isoformat(),
"days_left": (leaf.not_valid_after_utc - datetime.now(timezone.utc)).days,
"serial": format(leaf.serial_number, "x"),
"sans": [n.value for n in leaf.extensions.get_extension_for_class(
x509.SubjectAlternativeName).value],
}
return out
The tls.key field is deliberately never read. The agent's ServiceAccount needs get and list on Secrets to do this at all, which is the most sensitive permission in this series, so the secrets-management post applies in full: the model receives the dictionary above, never the Secret object.
The third source connects to each Ingress host and compares the served serial to the stored one. This is the only way to catch class four:
def served_serial(host: str, port: int = 443) -> str | None:
ctx = ssl.create_default_context()
ctx.check_hostname, ctx.verify_mode = False, ssl.CERT_NONE # we want the cert, not validation
try:
with socket.create_connection((host, port), timeout=5) as sock:
with ctx.wrap_socket(sock, server_hostname=host) as tls:
der = tls.getpeercert(binary_form=True)
return format(x509.load_der_x509_certificate(der).serial_number, "x")
except (OSError, ssl.SSLError):
return None
def ingress_hosts() -> dict:
hosts = {}
for ing in net.list_ingress_for_all_namespaces().items:
for tls in ing.spec.tls or []:
for h in tls.hosts or []:
hosts[h] = f"{ing.metadata.namespace}/{tls.secret_name}"
return hosts
Merging the three gives a per-certificate record: days left, whether cert-manager owns it, whether cert-manager says it is ready, and whether what is served matches what is stored. Everything with more than 21 days left and a matching served serial is dropped. What remains is the shortlist, and on most clusters it is under ten items.
For cert-manager-managed items on the shortlist, the wrapper also pulls the chain of objects that explains the failure, because the reason string lives at the bottom of it:
kubectl get certificate,certificaterequest,order,challenge -n shop -o wide
NAME READY SECRET AGE
certificate.cert-manager.io/shop-tls False shop-tls 88d
NAME STATE DOMAIN REASON
challenge.acme.cert-manager.io/shop-tls-... pending shop.example Waiting for HTTP-01 challenge propagation: wrong status code '404', expected '200'
That reason string is the single most useful input the model gets, and the wrapper passes it verbatim.
Step 2: The Diagnosis Call
One model call per shortlisted certificate, with a forced tool schema so the answer is structured:
CERT_TOOL = {
"name": "diagnose_certificate",
"description": "Explain why one certificate has not renewed and choose a remediation.",
"input_schema": {
"type": "object",
"properties": {
"diagnosis": {"enum": [
"acme_http01_unreachable", "acme_dns01_not_propagated",
"acme_rate_limited", "issuer_not_ready",
"unmanaged_secret_expiring", "renewed_not_reloaded",
"served_cert_is_external", "unknown"]},
"action": {"enum": [
"retrigger_renewal", "rollout_restart",
"propose_ingress_fix_pr", "propose_adopt_cert_manager_pr",
"wait_for_rate_limit", "ask_owner", "report_only"]},
"workload": {"type": "string",
"description": "namespace/kind/name to restart. Only for rollout_restart."},
"evidence": {"type": "string",
"description": "2-3 sentences quoting the challenge or order reason, "
"the days-left figure, and the served-vs-stored serial."},
"risk": {"type": "string",
"description": "What breaks if this action is wrong, and why the "
"safer alternative was not chosen."},
},
"required": ["diagnosis", "action", "evidence", "risk"],
},
}
SYSTEM = (
"You diagnose TLS certificate renewals for an SRE team. A 404 or connection "
"refused on an HTTP-01 challenge is an Ingress routing problem: the "
"/.well-known/acme-challenge/ path is not reaching the cert-manager solver "
"pod. Choose propose_ingress_fix_pr. A DNS-01 'not yet propagated' reason "
"under two hours old is normal; choose report_only. A rateLimited order is "
"never fixed by retrying: choose wait_for_rate_limit and say when the window "
"clears. If the Secret is fresh but the served serial is old, the pod read the "
"cert at startup: choose rollout_restart and name the workload. An unmanaged "
"Secret is always ask_owner or propose_adopt_cert_manager_pr, never "
"retrigger_renewal. If the served certificate is not in any Secret at all, "
"TLS terminates outside the cluster: served_cert_is_external, report_only. "
"Use retrigger_renewal only when the Certificate is Ready=True, inside its "
"renewal window, and no Order exists for it."
)
The last rule in the prompt matters. cmctl renew is the tempting hammer, and on a certificate with a failing challenge it does nothing except create another failed Order, which counts against the Let's Encrypt limit of five failed validations per hostname per hour. The agent is allowed to swing it only in the one situation where it helps: cert-manager simply has not got round to renewing yet, usually because the controller was down or restarted mid-cycle.
Step 3: Executing With Hard Limits
The RBAC, built the way the least-privilege ServiceAccount post lays out, encodes what the agent cannot do rather than trusting the prompt:
rules:
- apiGroups: [""]
resources: ["secrets"]
verbs: ["get", "list"] # parse tls.crt; tls.key is never read
- apiGroups: ["networking.k8s.io"]
resources: ["ingresses"]
verbs: ["get", "list"]
- apiGroups: ["cert-manager.io"]
resources: ["certificates", "certificaterequests", "issuers", "clusterissuers"]
verbs: ["get", "list"]
- apiGroups: ["cert-manager.io"]
resources: ["certificates/status"]
verbs: ["update"] # what cmctl renew needs, nothing more
- apiGroups: ["acme.cert-manager.io"]
resources: ["orders", "challenges"]
verbs: ["get", "list"]
- apiGroups: ["apps"]
resources: ["deployments", "statefulsets", "daemonsets"]
verbs: ["get", "patch"] # rollout restart annotation only
No create or delete on Secrets, so the agent cannot paste a certificate in by hand or wipe one that cert-manager will regenerate. No write on certificates, issuers, or ingresses, so every configuration change becomes a pull request through the GitOps-for-agents path.
The executor adds three rules RBAC cannot express:
-
rollout_restartis gated by days left. Under 7 days it runs autonomously, because a stale cert that expires at 3 a.m. is worse than a controlled restart at 3 p.m. Above 7 days it posts for approval, with the two serials side by side. -
retrigger_renewalis once per certificate per 24 hours, tracked in the agent's own state. This is the rate-limit protection the prompt cannot guarantee. - All mutations sit behind the circuit breaker. A certificate agent that restarts four deployments during an unrelated incident has made the incident worse, so the SLO burn gate applies here like everywhere else.
The restart itself is the same patch kubectl rollout restart sends, applied only when the executor has re-verified the serial mismatch a second time immediately before acting:
def rollout_restart(ns, kind, name, expected_secret_serial, host):
if served_serial(host) == expected_secret_serial:
return "skipped: served cert already current"
body = {"spec": {"template": {"metadata": {"annotations": {
"certagent.devtocash.com/restartedAt": datetime.now(timezone.utc).isoformat()}}}}}
getattr(apps, f"patch_namespaced_{kind}")(name, ns, body)
A Real Run
A 40-namespace EKS cluster, 61 TLS Secrets, 38 of them managed by cert-manager. The inventory produced a shortlist of six:
shop/shop-tls managed days_left=19 ready=False acme_http01_unreachable -> propose_ingress_fix_pr
api/api-tls managed days_left=27 ready=True renewed_not_reloaded -> rollout_restart (approval)
legacy/portal-tls unmanaged days_left=9 ready=n/a unmanaged_secret_expiring -> ask_owner
admin/grafana-tls managed days_left=3 ready=False acme_rate_limited -> wait_for_rate_limit
edge/www-tls managed days_left=41 ready=True served_cert_is_external -> report_only
ml/jupyter-tls managed days_left=6 ready=True renewed_not_reloaded -> rollout_restart (auto)
The shop failure was a new path-based Ingress rule that shadowed the ACME challenge path with a 404 from the storefront; the PR the agent opened added the /.well-known/acme-challenge/ route back above it. The grafana cert had been recreated by a Helm upgrade six times that week, one Order per upgrade, and had hit the duplicate-certificate limit; the honest answer was to wait two days, and the agent said so with the date. The legacy portal turned out to be a Secret pasted in by a contractor in 2024 with a 2-year lifetime, which nobody knew until the agent asked. The www cert lives on the CloudFront distribution in front of the cluster and was never in scope. Two restarts, one PR, one Slack message, one wait, and nothing expired.
Honest Limits
The served-serial check only works for hosts the agent can reach on 443 from inside the cluster, and only for hosts that resolve to the same place users hit. Split-horizon DNS or an internal-only Ingress class silently drops those hosts from class-four detection, and you should log the count of unreachable hosts as its own metric.
The unmanaged-Secret scan tells you a cert is expiring, not who owns it. Without a service catalog to resolve the namespace to a team, ask_owner degrades into a message in a shared channel, which is where certificate expiries have always gone to die.
And the fifth class stays mostly out of reach. Webhook caBundles, kubelet serving certificates, and service-mesh CA roots each expire on their own schedules, and the upgrade readiness agent is the better place for the webhook check because that is when it bites. This agent earns its keep on the boring majority: the certificates cert-manager was supposed to renew, and quietly did not.
📌 Read the latest version of this guide — plus the full library of DevOps, SRE, Kubernetes, observability & cloud-cost guides — on devtocash.com.
Top comments (0)