Internal certificate authorities fail on lifecycle, not on cryptography
The Monday morning outage that follows a certificate expiry is almost never a cryptographic failure. It is an ownership failure. A RADIUS certificate on a Wi-Fi authentication server expires over a weekend, nobody knows it exists, and the first shift cannot connect.
Why internal PKI drifts
Public certificate rules do not constrain internal PKI. The CA/B Forum timeline that shortens public TLS validity toward 47 days by 2029 applies to publicly trusted certificates, and internal PKI is explicitly outside it. An organisation can therefore set its own validity periods and its own verification policy. That freedom is also where the trouble starts, because nothing forces internal certificate inventories to exist.
Certificate sprawl follows from normal engineering practice. Network engineers own RADIUS and VPN certificates. Identity teams own SAML and OIDC signing certificates. Application teams issue TLS certificates through cloud platforms or CI pipelines. IT administrators deploy device enrolment certificates from an endpoint management system. Each team operates sensibly in isolation, and the result is an estate where nobody holds the complete picture.
An internal CA inherits a second dependency that public CAs do not create: the trust store. Every client that trusts the internal root must be updated when the root changes. That population includes operating systems, browser stores, appliances, embedded devices and third-party services, and it is frequently larger than the team that runs the CA believes.
The failure modes that recur
Three failure modes account for most of the incident reports.
The first is the invisible certificate. The inventory is incomplete, so the first sign of the problem is a service going down rather than a renewal ticket. RADIUS, VPN, SAML signing and device enrolment certificates are the usual candidates because they sit with teams that do not think of themselves as PKI operators.
The second is renewal that stops at the CA. A new certificate is issued and installed on the server, but the service is not reloaded, or the full chain is not deployed, or the private key does not match. The CA reports success and the service stays offline. This is why renewal, rotation and revocation belong in one operational workflow instead of separate tickets. Renewal replaces an expiring certificate, rotation normally generates a new key so the old one does not persist, and revocation handles compromise or policy-driven invalidation.
The third is trust store lag. A root or intermediate change requires every relying client to receive the new chain, and the rollover window has to be long enough for that distribution to complete. Guidance for private CA hierarchies puts subordinate CA validity at two to five times the lifetime of the certificates it issues, and recommends a long root lifetime, precisely to keep these transitions infrequent.
Structure that reduces the risk
Pick validity periods by working backwards from the endpoints. A short leaf lifetime limits exposure if a private key is lost, and it also means more renewals, and a missed renewal is an outage. The published default for private endpoint certificates is around thirteen months, with subordinate CA lifetimes measured in years. Organisations should pick a value deliberately instead of inheriting a default.
Plan root changes as a programme. Replacing a root affects the whole PKI and every relying trust store, so the practical approach is replacement rather than extension: create a successor CA, issue from it, distribute the new root alongside the old, monitor the transition milestones, then disable the predecessor once its certificates have expired. Naming the generations explicitly, such as a G2 suffix, avoids confusion during the overlap.
Automate the leaf, control the root. Automated issuance and renewal for service certificates removes the most common failure mode. Root changes should stay under human approval with a published schedule.
Instrument the inventory as a control. Every certificate record needs an owner, a service dependency list and an expiry date, and the CA should log issuance and revocation through the platform audit trail so that actions are attributable. A monitoring rule that fires at thirty days beats a spreadsheet reviewed quarterly.
Keep revocation honest. An internal CA that has no CRL or OCSP distribution path cannot revoke anything quickly, and OCSP carries its own operational cost because every validation becomes a network call to the responder. Revocation is only meaningful if relying clients actually check.
Limitations
None of this removes the underlying difficulty. Internal PKI governance spans teams with different priorities, and the team running the CA rarely controls the clients that must trust it. Automated renewal reduces missed renewals but introduces automation credentials and pipeline dependencies that then need their own protection. A very short leaf lifetime without automation turns an availability risk into an inevitability.
Defensive implications
Start with an inventory that names an owner for every certificate and every trust store entry, then make expiry a monitored event rather than a documented one, then automate leaf renewal and keep root rotation under a scheduled change process. Treat the trust store population as part of the PKI scope, because it is what makes root rotation expensive, and expensive rotations are the ones that get postponed.
References
- Managing private CA lifecycle: https://docs.aws.amazon.com/privateca/latest/userguide/ca-lifecycle.html
- AWS Private CA best practices: https://docs.aws.amazon.com/privateca/latest/userguide/ca-best-practices.html
- Rotating CA certificates: https://cloud.ibm.com/docs/secrets-manager?topic=secrets-manager-rotating-ca-certificates
- Preparing for shorter public TLS validity: https://www.globalsign.cn/blog_detailed_303
- Practical certificate lifecycle guide: https://www.purple.ai/blogs/zertifikate-richtig-verwalten-ein-praktischer-leitfaden-fur-den-lebenszyklus
Top comments (0)