“We need a 98% uptime SLA” sounds like a straightforward engineering requirement.
On a 30-day month, 98% uptime permits 14 hours and 24 minutes of downtime. For comparison, 99% permits 7 hours and 12 minutes; 99.9% permits 43 minutes and 12 seconds.
The arithmetic is easy. The difficult part is answering: uptime of what, measured how?
Define the promise before designing for it
Is the service “up” when the homepage returns HTTP 200? What if users can load the page but cannot sign in, save a record, or complete a payment?
Before committing to an SLA, define:
- The customer-facing functions covered by it
- The measurement period and method
- What counts as an unsuccessful request
- Whether planned maintenance is excluded
- How an incident starts and ends
A health endpoint that always returns 200 cannot prove your customers can use the product. Measure at least one critical user path from outside the system.
Remove single points of failure
AWS managed services can take substantial infrastructure work off your team. API Gateway and Lambda, for example, do not require you to maintain a single application server. But using managed services does not make the whole application highly available.
Your database configuration, service limits, authentication flow, external dependencies, and deployment process can still take customers down.
For each critical component, ask:
- What happens if this component fails?
- What detects the failure?
- What restores service?
- How long have we observed that recovery taking?
If a database restore is part of the answer, run one. A recovery time written in a document is an estimate until you test it.
Watch the service from outside and inside
An external check tells you whether customers can reach a function. Internal metrics help explain why they cannot.
For an AWS API, I would start with an external HTTPS check and alarms for API Gateway 5xx errors and latency, Lambda errors and throttles, and relevant database health signals. Route alerts to someone responsible for responding.
Avoid paging on every isolated error. Use thresholds and evaluation periods that reflect actual customer impact. Review what happens when a metric stops reporting; missing data can itself hide a failure.
Most of all, make sure an alert leads to an action. If nobody knows what to do when it fires, improve the alert and the runbook together.
Let non-critical features fail separately
A failure should affect as little of the product as possible.
On my site, static pages are served separately from dynamic features such as comments. If the comments API has a problem, readers should still be able to open an article.
That separation does not make the API reliable. It limits the customer impact when the API fails.
Treat deployments as a reliability risk
A healthy AWS region will not protect you from a bad release.
Use a controlled production deploy, check the customer-facing paths immediately afterward, and keep a rollback procedure you have exercised. Record how long detection and rollback take. That time spends part of your downtime allowance.
A 98% target gives you room to recover, but it is not permission to guess. Define the service, measure what customers experience, and test the failures you expect to handle.
I wrote more about the AWS building blocks and recovery choices in the original article.
What does your team currently count as “up”: a responding endpoint or a completed customer action?
Top comments (0)