The frontend is taking orders. The API is smiling professionally. Somewhere in the back, MySQL has locked the kitchen, turned off the lights, and left a note that says “connection refused.”
The load balancer would like table seven to know their request is very important to us.
MatrixSwarm’s mysql_watchdog is the shift manager who actually checks the kitchen. It watches the local MySQL or MariaDB service, collects useful clues when health changes, can request a restart, and gets an alarm to the humans who can do something about it.
No inspirational poster required. Just an agent with a job.
The two-item health menu
The current watchdog calls the database healthy when the configured systemd service is active and its configured TCP port appears among local listeners. The checks run on the Linux/systemd host where the database lives.
This is deliberately a small menu: service state and a listening port. It does not log in, execute a query, measure replication lag, or establish that your schema migration was a sound life choice. You do not need database credentials for those basic process/listener checks.
The configuration also exposes a socket path, but the current worker’s health decision uses the TCP listener check; the socket helper is not its fallback health path. A socket-only database needs a different or extended probe. Putting a socket path in a field does not make the worker use it. That would be configuration by wishful thinking, a surprisingly popular framework.
Open the kitchen in Phoenix
Add mysql_watchdog from the Agent Palette to your Swarm Workspace. Start with the facts on the database host:
| Setting | Example | What to verify |
|---|---|---|
| Service Name | mariadb |
Use the actual unit: it may instead be mysql or mysqld
|
| MySQL Port | 3306 |
Match the local TCP listener |
| Check Interval (sec) | 20 |
Choose a sensible polling interval |
| Alert Cooldown (sec) | 300 |
Controls eligible repeated failure notifications |
| Alert To Role | hive.alert |
Match the alarm relays’ advertised service |
These are example settings, not a complete deployment file. Configure the service account’s access to service status and listener inspection. For automatic recovery, its noninteractive sudo permission must cover restart of the exact database unit.
Give it a narrow permission for its job. Handing a dishwasher the keys to the entire shopping center is not operational elegance.
What happens when service goes sideways
The first probe records a baseline and returns. Once a healthy baseline exists, a transition to unhealthy triggers diagnostic collection, eligible alert delivery, optional forensic reporting, and a service restart request.
The agent can collect systemd status and recent lines from common MySQL/MariaDB error logs. This is useful evidence when “database unavailable” turns out to mean a full disk, a bad setting, or a permissions problem wearing a fake mustache.
A successful restart command is not proof that queries work. Subsequent health probes determine when the service/listener combination is healthy again, and the recovery path can send a notice. If the state stays down, the transition-based worker does not launch a new restart on every poll.
That also means starting this watchdog against an already-down database does not immediately exercise the failure-transition path. Commission it with the database healthy, then test a controlled outage in a test environment. Keep separate SQL/application checks for the meal actually reaching the table. MySQL’s mysqladmin documentation is a useful reference for additional database administration probes; those are not automatically performed by this watchdog.
The alarm menu is à la carte
Here is the part that saves unnecessary architecture meetings: Apache, MySQL, Nginx, and Redis watchdogs can each work with any of the swarm-alarm relays—Slack, Discord, Telegram, or email. There is no “MySQL only speaks email” rule hidden behind the refrigerator.
Pick the relay agent for your destination:
-
slack_relayfor Slack. -
discord_relayfor Discord. -
telegram_relayfor Telegram. -
email_sendfor email alerts through your SMTP configuration.
The watchdog sends to the role in alert_to_role, normally hive.alert. Alarm relays advertise hive.alert@cmd_send_alert_msg. Configure the selected relay’s credentials and required Registry/signing assignments, and check that service discovery scope and routing connect the agents.
One reachable relay is enough for one destination. Several matching relays can receive the same alarm, so the operator on Telegram and the team in Slack can both get the news. A separate report_to_role consumer is optional if you also want structured forensic data.
The watchdog owns the health check. The relay owns getting the message to people. This division of labor is considerably healthier than making the database server learn everyone’s vacation schedule.
Put the sensitive order slip in an envelope
Discord, Telegram, and email support optional alert-payload encryption. Set encrypt_alerts to true on each relay that should use it and provision its assigned packet-signing/encryption keys. Packet signing alone does not turn on encryption of the outgoing platform message.
Phoenix labels the switches Encrypt Discord alert message, Encrypt Telegram alert message, and Encrypt alert email subject and body. Authorized recipients can open the matching relay’s Decrypt Message panel and decrypt the encoded envelope with the proper keys.
If secure wrapping fails, the relay does not downgrade that alert to plaintext. Email’s protected content contains the alert subject and body; delivery metadata still exists. Slack has normal HTTPS delivery but no equivalent MatrixSwarm alert-payload encryption toggle in this implementation.
And encryption is per destination. A locked envelope on Telegram does not protect the duplicate you sent openly to Slack. That is not encryption; that is a takeaway order with a second receipt taped to the window.
Run a dinner rehearsal
Before relying on it, test a healthy baseline, a controlled failure, restart permissions, recovery detection, delivery to every chosen channel, and decryption where enabled. Do this on a test host, not by surprising the production dinner rush.
The goal is simple: when the database kitchen closes, the people holding the orders should hear about it before the customers start reviewing the restaurant.
Victory Always. Yes, chef.
🌐 Links & Resources:
Try MatrixSwarm: https://matrixswarm.com
Join the Community / Discord: https://discord.gg/2USbWVBVV
Download Server: https://github.com/matrixswarm/matrixswarm
Youtube: https://www.youtube.com/channel/UCMjiY4_-W2KP5fHXO0eC2ug
Top comments (0)