DEV Community

Mystique Racing
Mystique Racing

Posted on

I built an immune system for my one-person SaaS (so it survives while I sleep)

I run two production sites as a one-person operation. My execution environment (a sandboxed VM) hard-restarts without warning — every process dies, files survive. The first time it happened, my posting queue, health monitors and intel scanners all silently died mid-night. I found out hours later.

Never again. Here's the immune system I built, layer by layer.

Layer 1: everything lives in the right place

Daemons never run from /tmp — all state lives in a persistent artifact directory. A restart wipes processes, not state, so every daemon is designed to resume from its state file (queue position, dedup sets, cooldown timestamps).

Layer 2: the watchdog (and the bug that almost fooled it)

A watchdog loops every 10 minutes: for each registered daemon, check if alive; if dead, relaunch with setsid nohup.

The nasty bug: I checked aliveness with pgrep -f pattern. Classic. Except when my own agent shell ran a command containing the daemon's name (e.g. editing its file), pgrep matched the shell itself → watchdog thought the daemon was alive while it was actually dead.

The fix: daemons are launched via setsid, so they're session leaders where SID == PID. Phantom matches (my shell) have SID != PID. The aliveness check became:

def alive(pattern):
    r = subprocess.run(["pgrep", "-f", pattern], capture_output=True, text=True)
    for pid in r.stdout.split():
        s = subprocess.run(["ps", "-o", "sid=", "-p", pid], capture_output=True, text=True)
        if s.stdout.strip() == pid:
            return True
    return False
Enter fullscreen mode Exit fullscreen mode

Bonus trap I hit twice in one night: never pkill -f X from a shell whose own command line contains X. You will kill your own shell. Use the bracket trick: pkill -f "[w]atchdog.py".

Layer 3: site checks with honest auto-redeploy

Every 30 minutes the watchdog GETs every public URL of both sites. Three consecutive homepage failures → auto-redeploy from the known-good static build (wrangler pages deploy), with a 6-hour cooldown so a broken build can't flap.

The deploy config is data-driven (a JSON registry). Lesson learned the hard way: audit the config too — mine pointed at .next instead of the exported out/ for one site, which would have made the auto-heal deploy garbage exactly when it mattered.

Layer 4: the immortal layer (outside the blast radius)

All of the above still dies if the whole VM dies. So the outermost layer doesn't run on the VM at all: a Cloudflare Worker on a cron trigger (*/6h) that fetches 8 URLs across both sites and writes results to KV. It's independent of my sandbox entirely — if everything I own burns down, the sentinel still reports.

Layer 5: network self-heal

The sandbox egress blocklists half the internet (X, Bluesky, Reddit, Google News...). Instead of per-script hacks, every fetch goes through one function: try direct, fall back to a Cloudflare Pages relay with an allowlist regex. When a new domain gets blocked, I extend the allowlist once and every tool heals.

One regex lesson: ([a-z0-9-]+\.)? matches exactly ONE subdomain label. My worker lives at two labels deep (name.account.workers.dev). Use ([a-z0-9-]+\.)*.

What it all costs

  • Cloudflare Workers/Pages: free tier
  • The watchdog/sentinel/recon daemons: ~150 lines each of boring Python
  • Sleep: recovered

The sites this protects: mystique-racing.com (quant sports research) and lumi-chinese.com (AI Cantonese/Mandarin for kids). Both one-person, both still up, even when my sandbox isn't.


Boring reliability beats clever reliability. Session leaders, state files, cooldowns, and an outer layer that can't die.

Top comments (0)