DEV Community

Daniel
Daniel

Posted on Fully Autonomous

Nginx Watchdog: Tower, We’ve Lost the Reverse Proxy

Tower: Nginx, confirm you are receiving traffic.

Nginx: …

Tower: Nginx, say again.

Application team: We have seventeen dashboards.

Tower: Splendid. Do any of them answer the radio?

MatrixSwarm’s nginx_watchdog gives your local Nginx service a small control tower: regular service/listener checks, alarms on health transitions, diagnostic context, and a restart path when the service becomes unhealthy.

It is a useful job for a simple agent. Nobody needs a forty-slide observability strategy to notice the runway has gone missing.

File the flight plan

In Phoenix’s Swarm Workspace, add nginx_watchdog from the Agent Palette to the Linux/systemd host running Nginx. Then configure the actual service and listening ports.

For a machine that really listens locally on ports 80 and 443, an example config fragment is:

{
  "service_name": "nginx",
  "ports": [80, 443],
  "check_interval_sec": 10,
  "restart_limit": 3,
  "alert_cooldown": 300,
  "alert_to_role": "hive.alert"
}
Enter fullscreen mode Exit fullscreen mode

Set ports explicitly. The palette metadata inspected for this article contains [85]; your own installation may use 80, 443, 8080, or something else. Defaults do not know where you built the airport.

Every configured port must appear among local TCP listeners for the port check to pass. Do not demand 443 from a backend host where TLS terminates somewhere else.

The agent also checks whether the systemd unit is enabled. If it is disabled, this implementation waits rather than proceeding with its normal monitoring work. Confirm that this gate matches your deployment. “Installed,” “running,” and “enabled at boot” are three different passengers with suspiciously similar luggage.

Automatic restarts require the agent service account’s narrowly scoped, noninteractive sudo permission for the configured unit. Configure that as part of deployment; a restart request cannot talk its way past a missing permission.

What the tower sees

The healthy verdict combines active systemd state with the configured ports appearing in ss -ltn output. It is local service monitoring, not an external request through your whole application stack.

A running Nginx with a dead upstream may happily return 502s while passing this service/listener test. TLS problems, bad application responses, and remote network failures need additional checks. A green runway light does not prove there is a pilot in the plane.

Nginx’s stub_status module offers additional connection/request statistics if you choose to build more monitoring, but this watchdog does not poll it.

Loss of signal

The first health probe establishes a baseline and returns. After that, a healthy-to-unhealthy transition can collect service status and recent Nginx error-log lines, send an eligible alarm, optionally report structured evidence, and request a restart.

After a successful restart command, the implementation waits briefly and checks the configured ports again, warning if they still are not listening. Later health transitions can produce recovery notifications.

There is a restart failure counter and a disable guard. This is not three immediate retries on every bad poll: an unchanged down state does not repeatedly enter the transition-based recovery path. Starting the agent while Nginx is already down likewise records that first state instead of immediately running that failure path.

Commission against a healthy service, then exercise a controlled failure on a test host. The best time to discover what a watchdog does is before everyone is typing “any update?” into the incident channel.

Choose your radio frequency

The alarm destination is modular. Any Apache, MySQL, Nginx, or Redis watchdog can pair with any of the swarm-alarm relays: Slack, Discord, Telegram, or email. The agent name does not reserve a particular channel.

For a small deployment, use the watchdog plus one configured relay inside your existing swarm. For several destinations, add several reachable relays advertising the alarm role.

Channel Agent to add Payload-encryption option
Slack slack_relay No equivalent MatrixSwarm toggle here
Discord discord_relay Optional
Telegram telegram_relay Optional
Email email_send Optional

Keep the watchdog’s alert_to_role set to hive.alert unless you are deliberately changing the route. The relays advertise hive.alert@cmd_send_alert_msg. Resolve their credentials and required Registry/signing assignments, and ensure service scope makes them discoverable from the watchdog.

The watchdog can fan an alarm out to matching endpoints. A separate report_to_role consumer is optional for structured forensic reports. It is not required just to tell someone the proxy has left the building.

Secure the radio traffic

Discord, Telegram, and email each have an optional encrypt_alerts switch. Enable it on the specific relay and supply its assigned packet-signing/encryption keys. Merely configuring packet signing does not enable outgoing alert-payload encryption.

The Phoenix controls are Encrypt Discord alert message, Encrypt Telegram alert message, and Encrypt alert email subject and body. Recipients with the proper keys can paste the protected envelope into the matching relay’s Decrypt Message panel.

If secure wrapping fails, those relays do not quietly fall back to plaintext. Email keeps the alert subject and body inside the protected envelope, while delivery metadata remains visible. Slack still uses HTTPS for normal delivery; it does not have this additional MatrixSwarm payload-encryption option.

Protect each copy that needs protection. An encrypted Telegram alarm and a plain Slack duplicate are two different transmissions. Confidentiality does not spread through the swarm by enthusiasm.

Preflight before production

Verify the service name, enabled state, real ports, and restart permission. Then rehearse the healthy baseline, controlled outage, recovery, every alarm destination, and decryption where enabled.

Now the proxy has a tower that checks in, takes notes, and calls the right people. Your seventeen dashboards may continue looking expensive in peace.

Victory Always. Tower out.


🌐 Links & Resources:

Try MatrixSwarm: https://matrixswarm.com

Join the Community / Discord: https://discord.gg/2USbWVBVV

Download Server: https://github.com/matrixswarm/matrixswarm

Youtube: https://www.youtube.com/channel/UCMjiY4_-W2KP5fHXO0eC2ug

Top comments (0)