DEV Community

Manoj Kumar Kagitha
Manoj Kumar Kagitha

Posted on

Certificates Expire on the Worst Possible Day

Every team I have worked with has a certificate story. The classic version goes like this. A certificate expires on a holiday weekend. The renewal was assigned to someone who left the team months earlier. Production starts serving TLS errors, and the fix takes eleven minutes while the explanation takes three weeks.

Certificates are the most boring part of infrastructure. They are also the part that fails most loudly. A database can degrade slowly. A certificate either works or it does not. One day the handshake succeeds. The next day every client refuses to connect. There is no graceful mode.

This is a walkthrough of how I think about certificate lifecycle management in a large Azure environment: discovery, automation, alerting, and rotation. Nothing here requires exotic tooling. It requires discipline.

Find every certificate first

You cannot manage what you cannot see. Before automation, before policies, before dashboards, you need a list. Every certificate in every environment, with its expiry date, its issuer, and the service that serves it.

In a real enterprise this is never one place. Certificates live in Azure Key Vault, in API Management custom domains, on Application Gateways, on Front Door endpoints, in Kubernetes secrets, and in that one VM nobody documented. In my experience, the first job on certificate work is not renewal. It is inventory.

One Azure CLI query gives you the foundation:

az keyvault certificate list \
  --vault-name $VAULT \
  --query "[].{name:name, expires:attributes.expires}" \
  -o table
Enter fullscreen mode Exit fullscreen mode

Loop that across every vault in every subscription on a schedule and you have expiry visibility. Pipe the output into your monitoring system and you have the start of alerting.

The inventory needs to live somewhere queryable. A spreadsheet dies the day someone forgets to update it. Pull the data into a dashboard or at least a scheduled report. And every certificate needs an owner. Anything marked "unknown" gets one assigned within a week. The unowned certificates are the ones that expire.

Renewal has to be automatic

Manual renewal does not scale, and it does not survive team changes. The person who knows the renewal procedure today will not be on the team in two years. If the process depends on a person, it will fail.

Azure Key Vault can auto-renew certificates issued through integrated CAs. For public CAs, you set up the renewal once and let the vault handle it from there. For Kubernetes workloads, ACME-based renewal with cert-manager removes the human from the loop entirely.

The hard part is not the renewal. It is the deployment. Renewing a certificate in the vault means nothing if the Application Gateway or API Management instance keeps serving the old one. The renewal pipeline has to push the new certificate to every consumer. Key Vault references in App Service and Functions pick up new versions automatically. For gateways and load balancers, you need a pipeline step or an event-driven function that notices the new version and applies it. Test this end to end in a non-production environment first. The renewal ceremony is the whole chain, not the issuance.

Alert long before expiry

Automation fails too. CAs have outages. Permissions drift. Someone changes a firewall rule and the renewal challenge cannot reach the domain. You need alerts that fire early enough to fix the problem by hand.

The schedule I recommend is 60-30-7. Sixty days out: informational, goes to the team channel. Thirty days out: warning, gets a ticket. Seven days out: page someone. Most certificates are fine at sixty days. The point of the early alert is not urgency. It is to catch the broken automation while you still have options.

The alert must go to a team, not a person. Certificates outlive people. Route expiry alerts to a distribution list or an on-call rotation. If the alert goes to an individual inbox, it will eventually go to a departed employee's inbox, which is the same as no alert at all.

Rotate without downtime

Renewing the certificate is step one. Serving it without dropping connections is step two. Most Azure services handle this well now. Application Gateway lets you update the listener certificate without downtime. API Management custom domains can be updated in place. Kubernetes ingress controllers pick up new secrets on reload.

The pattern that keeps you safe: renew early, deploy the new certificate while the old one is still valid, and verify the serving path before the old one expires. Keep the old certificate in the vault for a short overlap period. If something goes wrong with the new one, you roll back in minutes instead of scrambling for a reissue.

Test your rollback. I mean actually test it, in staging, with someone watching the metrics. A rollback plan you have never run is a theory.

Do not forget internal certificates

Public certificates get all the attention because they are customer facing. Internal certificates expire just as hard. Service-to-service mTLS, internal APIs, webhook endpoints between systems: all of them have expiry dates, and they tend to be managed by nobody in particular.

If your organization runs an internal CA, treat its certificates with the same lifecycle as public ones. Same inventory, same alerts, same automation. The internal CA root itself needs a plan. Root rotation is rare and painful, which is exactly why it needs to be written down before it is urgent.

What I would do on day one

If you inherited certificate management tomorrow, here is the order I would work in.

First, build the inventory. Script discovery across every Key Vault and every TLS endpoint you can reach. Assign an owner to each certificate.

Second, set up expiry alerting on the 60-30-7 schedule, routed to a team channel and an on-call rotation.

Third, pick the certificates closest to expiry or most critical, and automate their renewal end to end, including deployment to the serving infrastructure.

Fourth, extend automation to the rest, starting with the public-facing ones.

Fifth, document the rollback procedure and test it once a quarter.

None of this is exciting work. That is the point. Certificates should be boring. The goal is a system where expiry is a non-event: the certificate renews, the pipeline deploys it, the alert never fires, and nobody remembers it happened. Boring is the victory condition.

Top comments (0)