DEV Community

chovy
chovy

Posted on Originally published at dev.profullstack.com

Booking.com's Node savings don't carry over to Bun. Autoheal does.

Booking.com wrote up how they cut Node.js costs by 38% with Watt. They swapped pm2's cluster module for worker threads that accept connections themselves, gave each worker its own port behind nginx, and let Watt restart any worker whose event loop blocks. Result: 30% fewer pods, 20% less memory per pod, lower tail latency.

We run about 230 containers on one box, most of them Bun. So we measured before copying anything.

Consolidation: 13MB per app, not worth it

The appeal is packing many small apps into one process. On Bun the numbers are small:

  • a bare Bun.serve process: 15MB
  • 10 servers as Workers inside one process: 34MB, about 2MB each

Our small apps sit at 50 to 150MB, and that is their own heap, which a shared process would not share. The ceiling across all 111 Bun processes is about 1.4GB. The price is that one crash takes down every app in the process. We skipped it. Booking's gain came from Node's heavier baseline and pm2's IPC supervisor, and we have neither.

The part that did apply: restart the stuck process

We had 183 containers with a Docker HEALTHCHECK. Docker runs it, counts the failures, marks the container unhealthy, and then does nothing. A restart policy only fires when the process exits, and a process with a blocked event loop has not exited. An in-process watchdog can't help either, because its timer never runs.

So @profullstack/watchdog 0.3.0 ships watchdog-autoheal:

npm install -g @profullstack/watchdog
watchdog-autoheal status     # healthy, unhealthy, and who has no healthcheck
watchdog-autoheal            # one pass: restart what Docker marked unhealthy
Enter fullscreen mode Exit fullscreen mode

It trusts Docker's failure streak instead of adding its own, so a real wedge gets restarted within about a minute. Restarts are budgeted: three per container per hour, then one email saying it needs a person. A broken image or a dead database comes back unhealthy, and looping on it only hides the problem.

It has run on our production box every minute since today. The first status also turned up 47 containers with no healthcheck at all, so that's the next job: one request to /healthz each.

Zero dependencies, MIT: https://github.com/profullstack/watchdog

It is also on https://profullstack.com/stack, the list of modules we are pulling out of our apps so you can reuse them.

Drafted with AI help, edited and measured by hand.

Top comments (0)