I built an uptime monitoring service, and the hardest part wasn't any single feature. It was the decision to not page anyone until I'm confident something is actually broken.
The false alarm problem
Naive uptime monitoring is trivial to build. You schedule a request, you check the status code, you fire an alert if it's not 200. Ship it, and within a week you regret it.
Here's what actually happens in production:
- A deploy restarts your server for 4 seconds
- A shared database hiccups for one query
- Someone's deploy triggers a 502 for a single request while the load balancer routes around it
- A CDN edge node in a distant region times out on one TLS handshake
Each of those is a "failed check." In a naive implementation each one opens an incident, sends you a Telegram message, and wakes you up. Within a month you mute the channel. Once you mute it, the one alert that mattered — the real 40-minute outage — goes unread with everything else.
So before writing any monitoring code, the rule was:
An incident only opens after N consecutive failed checks (default: 2).
One blip is a blip. Two in a row is a pattern. That's a one-line fix that completely changes the signal-to-noise ratio.
The second half: alert on transitions, not on failures
The related mistake is alerting per failed check. If something is down for an hour and you check every 60 seconds, that's 60 messages. Meaningless.
UptimeCadet sends alerts once per state transition:
// The whole alerting rule, in spirit
const wasDown = lastStatus !== 'up';
const isDown = currentStatus !== 'up';
if (isDown && !wasDown) notifyDown(); // fires exactly once
if (!isDown && wasDown) notifyUp(); // fires exactly once
Down → notification. Stayed down → silence. Recovered → notification.
This is the kind of thing that seems obvious once you know it and feels like a puzzle when you don't. Worth writing down so the next person doesn't reinvent it.
What I check beyond "is it up?"
HTTP status is the baseline, but it answers a much narrower question than people expect. A monitor that only checks for 200 misses entire categories of real problems.
SSL certificate expiry. A cert that expires at 3am doesn't break anything until 3am. Then everything breaks at once. Checking the cert's not-after date days in advance turns a catastrophe into a Tuesday afternoon.
Domain expiry via RDAP. The RDAP protocol replaced WHOIS because WHOIS responses were inconsistent and often rate-limited. Domain expiry is the same class of problem as SSL — an invisible countdown that becomes total at zero.
Lighthouse page-speed scores. Now this one is a genuine tradeoff, and I'll be honest about it. Lighthouse is slow — running it on every interval would be absurd. The value isn't "your site got faster," it's directional: a score that slides from 92 to 55 over a week means you shipped something expensive. The metric is the trend, not the absolute number.
So the dashboard shows more than a green/red light. SSL, domain, and performance each get their own row.
Check interval is a plan limit, not a setting
The free plan checks every 300 seconds. Paid plans check every 60. The interval is clamped to the plan minimum server-side and capped at 3600:
requested: 10s → clamped to 300s (Free) or 60s (paid)
requested: 7200s → capped at 3600s
This is deliberate. The interval is the single biggest driver of infrastructure cost — a monitor checked every 10 seconds across thousands of endpoints is a different product economically than one checked every 5 minutes. If the client could just request whatever it wanted, the free tier would quietly be the most expensive tier.
The docs state the actual numbers rather than "varies by plan," because a monitoring tool that hides its resolution is not trustworthy.
Passwordless login over channels you already own
There's no password. You enter your email, choose where a 6-digit code should arrive — Telegram, Slack, Discord, or email — and the code must land on a channel you already connected and verified.
The constraint that makes this safe: nothing is active until you verify it. Connect a channel, receive a 6-digit code on it, enter the code. Until then it sends nothing.
It removes a whole category of problem. No password database to breach, no reset flow to build, no "we'll email you a link" that depends on the mail provider being up. If you can receive Telegram messages, you can sign in.
Codes are single-use, expire after 10 minutes, and allow 3 attempts.
The public status API, no key required
Published status pages are exposed as read-only JSON with no account and no API key:
GET https://pulseping.xyz/api/public/status?slug=your-slug
No key is a security decision, not an oversight. Everything it returns is already public — it's the same data rendered at /status/your-slug. Anyone who can see the page can read the JSON. Adding an API key would create the illusion of protecting something that isn't secret, while adding a signing, rotation, and revocation system to maintain.
Unpublished pages return 404, so drafts can't be scraped.
This is what makes embeddable status work:
<iframe src="https://pulseping.xyz/status/your-slug" />
or a fetch against the endpoint for a live badge in your README.
Seven languages, one of them full RTL
The UI ships in English, Arabic, French, German, Spanish, Portuguese, and Turkish — with Arabic as a genuine right-to-left layout, not a mirrored stylesheet with a few flipped icons.
This is more work than it looks. dir="rtl" handles text direction, but anything positioned with physical CSS properties (left, right, margin-left) has to become logical (inset-inline-start, margin-inline-end). Icons that imply direction — a back arrow, a progress bar — need to flip. Charts need their axes reversed. It's the kind of work that's invisible when it's done right and immediately obvious when it wasn't.
Honest tradeoffs
Things I'd do differently with more time:
- Checks are sequential per monitor. At high monitor counts this needs a real worker pool and concurrency limits. Right now it's fine at the scale I support.
- No multi-region probing. Every check originates from one location, so a regional outage can look healthy. Real uptime services run checks from several regions and report the worst result.
- Lighthouse on every interval is overkill. I check it on a slow cadence, but the scheduling deserves to be first-class rather than a special case bolted onto the checker.
- Retention is unlimited and I should cap it. Every check is stored. Over years that's a lot of rows for a service charging $12/month.
Try it
UptimeCadet — free tier with one monitor, checks every 5 minutes, email alerts, no credit card. Pro is $12/month for 1-minute checks, a public status page, and Telegram/Slack/Discord alerts.
Full documentation, including the alert rules and plan limits, is at pulseping.xyz/docs.
If you find a case where this design pages you for no reason — I'd genuinely like to hear it. That's the failure mode I'm most trying to avoid.
Built by saidtechnology — a one-person studio. I build privacy-first web and mobile tools.
Top comments (0)