DEV Community

Cover image for Codecov went down today over an expired certificate: how TLS renewal actually fails
Kashif Manzer
Kashif Manzer

Posted on

Codecov went down today over an expired certificate: how TLS renewal actually fails

Codecov went down today over an expired certificate: how TLS renewal actually fails

This morning, Codecov's status page lit up. Frontend and backend, both down. The cause, identified within two minutes of the first update: an expired SSL certificate. The incident timeline, which begins at 06:53 UTC today, goes from "investigating" to "identified: SSL certificate expiration" in two updates.

Yesterday it was Sprinklr. Same failure. An expired certificate knocked part of their production offline for 23 minutes before the team renewed and redeployed it.

Two certificate-expiry outages in 24 hours. And this is not a rare disease. This is the common cold of web operations, and it keeps hospitalizing serious companies.

Your servers are fine. That is the problem.

Here is the strange part of every certificate outage: nothing is actually broken on the server side. The app is running. The database is answering. Every internal health check is green.

The failure happens inside the visitor's browser, in the split second before any page loads. The browser asks for the certificate, sees the expiry date has passed, and refuses to continue. Your monitoring watched the server. The server was never the patient.

This is why the outage always feels so disorienting. The team stares at green dashboards while users stare at a red warning page. Two different realities, and the user's one is the only one that counts.

The hall of fame nobody wants to join

The roll call of serious companies taken down by this is long enough to be embarrassing.

In December 2018, an expired certificate at Ericsson knocked out mobile data for tens of millions of O2 and SoftBank customers across the UK and Japan. O2 users were offline for most of the day, and the damage settlement that followed was reported at around 100 million pounds (source, source).

In February 2020, Microsoft Teams went down for roughly three hours on an expired certificate. Spotify lost about an hour the same year when a TLS certificate lapsed, and the team only found it after a Cloudflare engineer spotted the expired cert (source).

LinkedIn let a certificate on its lnkd.in shortener expire twice, in 2017 and again in 2019. The second time, the renewed certificate had been issued days earlier. It just never got installed on the server (source).

Every one of these teams had smart people and real infrastructure. The certificate still won.

How renewal is supposed to work

A quick primer, because the failure modes only make sense once you see the machinery.

A TLS certificate is a file with an expiry date stamped inside it. Browsers trust it only between its start date and its end date. Public certificates used to last a bit over a year. Now the industry is shrinking that: the CA/Browser Forum cut the maximum validity from 398 days to 200 days starting March 2026, so certificates issued that March started expiring this month (industry reporting). That is the wave we are standing in right now.

Renewal is meant to be automatic. The standard protocol is called ACME. Your server proves it controls the domain, the certificate authority issues a fresh certificate, and a small agent on your machine installs it. Let's Encrypt made this free and ubiquitous. In Kubernetes, a tool called cert-manager does the same job for cluster workloads.

There are two common ways to prove domain control: answer a challenge file over HTTP (called http-01), or create a DNS record (dns-01). Then the agent reloads the web server so the new certificate actually gets served.

That last step is where the bodies are buried.

The five ways renewal silently fails

1. Renewed, but never reloaded. The new certificate lands on disk. The old one stays in the server's memory. Nginx, Apache, or your Go service keeps serving the expired file until someone restarts or reloads the process. This is almost certainly what happened to LinkedIn in 2019: the renewed certificate existed, it just was not the one being served.

2. The copy you forgot. Modern setups terminate TLS in several places: the CDN edge, the load balancer, the origin server. Renewal automation often covers exactly one of these. The edge gets a fresh certificate while the origin's copy quietly ages out, or the other way around. Every layer you did not automate is a timer you are not watching.

3. The DNS token died quietly. DNS-01 challenges need an API token for your DNS provider. Tokens get rotated, scopes get narrowed, the person who created it leaves the company. The renewal job starts failing, and it fails in the one place nobody reads: the output of a cron job. Sixty days of silent failures, then the expiry date arrives.

4. The root you did not know you depended on. Sometimes it is not your certificate that expires but the root it chains to. When Let's Encrypt's old cross-signed root expired, apps like Shopify and Slack saw outages on older devices that did not trust the new root (source). Your certificate can be perfectly fresh and still untrusted, because trust is a chain and chains have more than one link.

5. Monitoring that checks the wrong thing. Most uptime checks hit an HTTP endpoint or a health port. Those pass with an expired certificate. The check that catches this has to complete a TLS handshake from the outside and read the expiry date. Almost nobody's default monitoring does that, so the first alert is always a user.

None of these are exotic. Every one of them is a Tuesday.

A checker that reads the expiry from the outside

The fix is boring, which is why it works. Check every public endpoint from the outside, read the certificate's expiry date, and page someone while there are still weeks left, not minutes.

There is one subtlety that bites people: a normal HTTPS client refuses to talk to a server with an expired certificate, so your checker has to deliberately skip validation. It is not checking identity here, only reading the date. Without that, the handshake fails on an expired cert and you learn nothing.

Here is the whole thing in Go, using only the standard library:

package main

import (
    "crypto/tls"
    "fmt"
    "time"
)

// daysLeft opens a TLS connection the way a browser would,
// reads the certificate's expiry date, and reports how many
// days are left. InsecureSkipVerify is deliberate: we are not
// checking identity here, only reading the date.
func daysLeft(host string) (int, time.Time, error) {
    conn, err := tls.Dial("tcp", host+":443", &tls.Config{
        InsecureSkipVerify: true,
    })
    if err != nil {
        return 0, time.Time{}, err
    }
    defer conn.Close()

    certs := conn.ConnectionState().PeerCertificates
    if len(certs) == 0 {
        return 0, time.Time{}, fmt.Errorf("server presented no certificates")
    }
    expiry := certs[0].NotAfter
    return int(time.Until(expiry).Hours() / 24), expiry, nil
}

func main() {
    domains := []string{
        "example.com",
        "api.example.com",
        "status.example.com",
    }
    for _, d := range domains {
        left, expiry, err := daysLeft(d)
        if err != nil {
            fmt.Printf("%-20s CHECK FAILED: %v\n", d, err)
            continue
        }
        status := "ok"
        if left < 30 {
            status = "RENEW SOON"
        }
        if left < 0 {
            status = "EXPIRED"
        }
        fmt.Printf("%-20s %s: %d days left (expires %s)\n",
            d, status, left, expiry.UTC().Format("2006-01-02"))
    }
}
Enter fullscreen mode Exit fullscreen mode

Put your real domains in the list, not just the marketing site. The API subdomain, the status page, the webhook receiver, the admin panel. The certificate that takes you down is always the one you forgot you had.

Then add the alerting rule that actually works: warn at 30 days, page at 14, treat 7 as an incident. A certificate is free to renew. There is no reason to ever be surprised by its expiry.

The wave is just starting

Back to the timing, because it matters. With maximum validity now at 200 days, every team renews roughly twice as often as they used to. Twice as many renewal windows, twice as many chances for one of the five silent failures above. The teams that treated renewal as "set and forget" are about to learn what "forget" costs, twice a year, on a schedule.

Codecov and Sprinklr, 24 hours apart, might be the first drops of that rain. Or they might be coincidence. Either way, the mechanism does not care about the calendar. It only cares whether someone is watching the expiry date from the outside.

Check yours today. It takes five minutes, and the checker above is most of the work already done.

What is the most embarrassing cause of an outage you have ever seen? I want to hear the one that made the whole team go quiet.

Top comments (0)