Uptime monitoring means automatically checking whether your website, API, or service is accessible — and alerting you when it's not. But the full picture is more nuanced than that. Here's what you actually need to know as a developer or DevOps engineer.
The Simple Definition
Uptime monitoring is the practice of sending regular automated requests to your service and verifying that it responds correctly. When it doesn't, you get an alert.
Most uptime monitors work like this:
- Every minute (or every 30 seconds), make an HTTP GET request to your URL
- Check that the response status is 200 (or your expected code)
- Optionally check that the response body contains expected content
- If any check fails, wait for one more confirmation to rule out transient network issues
- Fire an alert via email, Slack, or webhook
Why Uptime Monitoring Matters
The cost of undiscovered downtime
Without uptime monitoring, you find out about downtime one of two ways:
- A user tells you
- You happen to visit your own site
Both are terrible. Users who experience downtime and have to tell you about it are less likely to come back. Every minute of undetected downtime has a measurable cost.
What 99.9% uptime actually means
| Uptime % | Downtime per year | Downtime per month |
|---|---|---|
| 99% | 3.65 days | 7.2 hours |
| 99.9% | 8.7 hours | 43.8 minutes |
| 99.95% | 4.4 hours | 21.9 minutes |
| 99.99% | 52.6 minutes | 4.4 minutes |
| 99.999% | 5.26 minutes | 26 seconds |
Most hosted apps on managed infrastructure (AWS, GCP, Heroku) target 99.9% uptime. To achieve that, you need to know when failures happen.
Types of Uptime Monitoring
1. HTTP/HTTPS Monitoring
The most common type. The monitor sends an HTTP GET request to your URL and checks:
- Status code (usually expecting 200)
- Response time (should be under a threshold, e.g., 3 seconds)
- Response body (optional keyword match)
2. TCP Port Monitoring
For services that are not HTTP — databases, SMTP servers, custom protocol servers. The monitor checks that a TCP connection can be established on a specific port.
3. Ping Monitoring
ICMP ping to a host. Verifies the host is responding at the network layer, independent of any application layer.
4. SSL Certificate Monitoring
Checks your HTTPS certificate's expiry date and alerts you before it expires. Certificate expiry is a top cause of preventable outages.
5. Heartbeat/Cron Monitoring
Inverted monitoring: your process (cron job, background worker) sends a ping to the monitoring service on every successful run. If the service does not receive a ping within the expected window, it alerts you.
Single-Region vs Multi-Region Monitoring
This is the most important concept for production monitoring.
Single-region monitoring checks your service from one location. If that location experiences a network issue, you get a false alert even though your service is fine.
Multi-region monitoring checks from multiple global locations simultaneously. It only alerts when multiple locations agree the service is down. This dramatically reduces false positives.
Vigilmon uses multi-region consensus monitoring: checks run from multiple regions, and an alert only fires when multiple regions confirm the failure.
Setting Up Your First Uptime Monitor
Step 1: Add a health check endpoint
Add a dedicated /health endpoint:
// Express.js
app.get('/health', (req, res) => {
res.json({ status: 'ok' });
});
# Flask
@app.route('/health')
def health():
return {'status': 'ok'}, 200
Step 2: Choose check frequency
- 1 minute — standard for production services
- 30 seconds — for high-availability, latency-sensitive services
- 5 minutes — acceptable for less critical services
Step 3: Configure alerts
At minimum:
- Email — for async notification
- Slack — for immediate team visibility
What to Monitor
| Asset | What to check | Frequency |
|---|---|---|
| Main website | /health returns 200 | 1 min |
| API endpoints | Critical paths return expected status | 1 min |
| Background workers | Heartbeat ping from the worker | Per job cycle |
| SSL certificates | Days until expiry | Daily |
| TCP services | Port accessible | 1 min |
Common Mistakes in Uptime Monitoring
Monitoring from only one region. A single-region failure looks like your service is down when it is not.
Monitoring the homepage instead of a health endpoint. Homepages include third-party scripts, CDN assets, and dynamic queries that can trigger false alerts.
Not testing the alert path. Add a monitor for a URL you control, then take it offline briefly to confirm you actually receive the alert.
Setting check intervals too long. 5-minute intervals mean up to 5 minutes of undetected downtime.
Not monitoring SSL. Expired SSL is the most common preventable outage type.
Getting Started
Vigilmon provides free multi-region uptime monitoring with:
- HTTP, TCP, and ping monitoring
- SSL certificate expiry alerts
- Multi-region consensus (no false alerts from single-region blips)
- Instant alerts via email, Slack, or webhook
- Public status pages
Setup takes about 2 minutes: add your URL, set your alert channel, and you are monitoring.
Try Vigilmon free: vigilmon.online
Top comments (0)