DEV Community

Daniel
Daniel

Posted on Fully Autonomous

RsyncBoy: “I Thought We Had Backups” Is Not a Recovery Plan

Incident report, 02:17.

Service status: unavailable.

Backup status: somebody remembers setting something up.

Recovery estimate: depends how quickly we can locate “somebody.”

Every ops team deserves better than a disaster-recovery plan built from folklore and a directory named backup_final_REAL.

Enter RsyncBoy, MatrixSwarm’s backup job dispatcher. In the palette, he answers to rsync_boy. His responsibilities include timestamped filesystem snapshots, MySQL dumps, scheduled execution, and keeping enough records that the next shift can tell what actually succeeded.

His hobbies include moving bytes and refusing to count “we started the process” as a successful backup. Management finds this attitude unnecessarily specific.

The job description is refreshingly concrete

RsyncBoy runs two supported kinds of job:

  • Filesystem timestamped snapshot: pull a directory from an SSH host onto the MatrixOS backup host, or push a local directory to an SSH destination. Each completed run gets its own dated directory.
  • MySQL timestamped dump: produce a database dump, optionally gzip it, calculate a SHA-256 checksum, write a manifest, and transfer the dump and manifest to the configured SSH destination.

The agent schedules the work; short-lived worker threads perform it. Phoenix gives you a live job table with the schedule, last success, running state, and backup-size information. You can edit jobs and execute a saved job immediately.

Which beats opening six terminals and trying to remember which green prompt belongs to the server that pays your mortgage.

First assignment: bring the website files home

For a concrete example, put MatrixOS on a backup machine and have it pull a small test directory from a separate SSH server. Here, “local” means the machine running the agent. Your Phoenix desktop is the control console, not the truck carrying the files.

Use current Phoenix and MatrixOS builds that include the RsyncBoy live panel and factories. The transfer hosts need the relevant command-line tools: rsync on both ends and SSH tooling on MatrixOS. Password or passphrase-based SSH modes also use sshpass; database jobs need the appropriate MySQL/MariaDB clients where the dump runs.

Then give the agent its assignment:

  1. Add rsync_boy to Swarm Workspace. Use your existing Matrix core, command ingress, and hive.rpc callback relay for Phoenix control. A typical transport pair is matrix_https and matrix_websocket; the actual backup data travels through SSH/rsync.
  2. Resolve the workspace’s Registry constraints. Assign packet signing, a dedicated persistent_state record, and the SSH connection details; configure the MySQL assignment for database work. Preserve the persistent-state record, identity, key, and universe across redeployments so the saved schedule remains accessible.
  3. Configure a trusted SSH host fingerprint. RsyncBoy checks the presented host key against the configured fingerprint. A hostname and a confident handshake are not identification documents.
  4. Open the agent’s configuration and choose Add Job. Select Filesystem timestamped snapshot, choose the SSH profile, and use settings like the ones below. These paths are examples—create a small test source and choose a destination owned by the MatrixOS user.
Field First test
Job ID website-files
Enabled Off, while checking the first run
Interval (sec) 86400
Pull Source Through SSH On
SSH Source Path /srv/backup-demo
Local Snapshot Root /srv/backups/website-files
Snapshot Prefix website
Hard-link Unchanged Files On
Keep Days 14

Save the configuration and deploy. Connect Phoenix, select RsyncBoy, and open its live jobs panel. Confirm the loaded job, then click Execute Now. Manual execution works for a disabled saved job, which makes it useful for this first check.

A newly enabled job with no matching successful history can become due on the next scheduler poll. A 24-hour interval does not promise a 24-hour grace period before its first run. Keep it disabled until its paths and credentials are ready; the scheduler does not read intentions from the ticket description.

What a completed filesystem run leaves behind

A successful sequence might produce this on the backup host:

/srv/backups/website-files/
├── website_20261007_120000/
│   ├── ...your files...
│   └── snapshot.manifest.json
├── website_20261008_120000/
│   ├── ...your files...
│   └── snapshot.manifest.json
└── latest -> website_20261008_120000
Enter fullscreen mode Exit fullscreen mode

The timestamps in snapshot names use UTC. During transfer, the new snapshot lives in a .partial directory. The worker promotes it to its final name and updates latest after transfer and manifest creation complete. That publication step keeps a half-finished directory from wearing the “latest completed snapshot” badge.

With hard-linking enabled, unchanged files can share storage with the previous snapshot through rsync’s --link-dest behavior. Matching preserved attributes and filesystem support matter; the upstream rsync manual explains the details. Treat completed snapshots as read-only: editing a shared hard-linked file in place can affect other snapshots using that same file.

Retention removes old matching snapshots according to Keep Days. Setting it to zero disables pruning. That setting should be a decision, not a ceremonial offering to the disk-full alert.

Use a separate snapshot root per job. A shared latest pointer is an excellent way to give two backup jobs an avoidable workplace dispute.

The scheduler keeps receipts

RsyncBoy records successful completion and the job definition in encrypted persistent state. With that state intact, restarting an unchanged job does not erase its successful-run history. A changed definition is evaluated as new work.

A failed run does not advance the last-success timestamp. The scheduler backs off before retrying, and it prevents another instance of the same job from launching while that job is running. Multiple due jobs can overlap, with staggered starts to avoid every transfer charging through the door at once.

Think of it as a shift supervisor who knows the difference between “assigned,” “in progress,” and “finished.” A rare and unsettling level of administrative competence.

The live panel’s Save Changes validates and persists the complete schedule, then waits for the agent’s acknowledgement. It also checks the configuration revision, so an operator with an old view cannot quietly overwrite someone else’s newer edits. Use Reload Server to get the current version before reconciling changes.

Once the first test succeeds, enable your job and save the schedule. Ordinary job edits can happen through the live panel. Adding an SSH profile that was never included in the deployment requires updating the workspace and redeploying; the agent cannot summon credentials from enthusiasm.

The database gets its own work order

For databases, choose MySQL timestamped dump and configure the MySQL Registry assignment, dump options, local staging directory, SSH destination, compression, and retention.

With Run MySQL Through SSH enabled, the database client runs on the selected SSH host and the dump is streamed back to MatrixOS. The worker then prepares the local file and uploads the result and manifest to that SSH host’s configured remote path. With the option disabled, the database client runs on MatrixOS instead. Check that topology deliberately when deciding where your recovery copies should live.

Use a dedicated remote directory for database dumps: the current remote cleanup matches old *.sql* files beneath that directory. It should not double as the office lost-and-found.

The manifest records the checksum and file details. It does not certify that the application can recover. Likewise, a filesystem snapshot is a file-copy operation, not an application-consistent freeze of a busy database. Choose a backup method and consistency settings suited to the workload.

Finally, keep the encryption claim precise: the schedule and success ledger are encrypted; backup files are not automatically encrypted at rest by these factories. SSH protects transfer traffic. Protect the backup storage itself according to what it holds.

Close the ticket with a restore

Retrieve a test file from a completed snapshot into a separate location and compare it with the original. For a database job, restore a test dump into an isolated database and check the result. Then make a harmless source change, run another backup, and confirm you can recover the expected versions.

The objective is a recovery you can demonstrate. A green status light is welcome, but it has never personally rebuilt a database at three in the morning.

RsyncBoy can handle the recurring shift. Your future self would appreciate a tested restore procedure and, if procurement approves it, coffee that does not taste like a UPS battery.


🌐 Links & Resources:

Try MatrixSwarm: https://matrixswarm.com

Join the Community / Discord: https://discord.gg/2USbWVBVV

Download Server: https://github.com/matrixswarm/matrixswarm

Youtube: https://www.youtube.com/channel/UCMjiY4_-W2KP5fHXO0eC2ug

Top comments (4)

Collapse
 
sgaggjhkjh profile image
sgaggjhkjh •

The incident report framing at the top is painfully relatable — "somebody remembers setting something up" is how most small teams actually experience backups until the day they need one. Starting there made the rest of the design land harder.

Two details stood out for me. First, treating completed snapshots as read-only when hard-linking is enabled — that is the kind of footnote that saves someone a genuinely awful afternoon. Second, the distinction between the scheduler knowing "assigned" vs "in progress" vs "finished," plus a failed run not advancing the last-success timestamp, is the real difference between monitoring and a checklist.

The closing line about the manifest is the thesis, really: checksums tell you the file is intact, only a restore tells you it works. Do you run scheduled restore drills against these snapshots, or is the restore still a manual "close the ticket" ritual?

Collapse
 
matrixswarm profile image
Daniel •

Thanks for the thoughtful read! Scheduled restore drills aren’t implemented in RsyncBoy yet; restore verification is still an operator step. The scheduler tracks backup completion, but that doesn’t establish recoverability.
A useful extension would be restoring into an isolated directory or disposable database, checking the recovered data, and recording a separate restore result and last-success timestamp. A successful backup shouldn’t conceal a failed restore drill.
You’ve identified the next gap to close. The manifest helps establish what was captured; the restore demonstrates what we can recover.

Collapse
 
sgaggjhkjh profile image
sgaggjhkjh •

Restore drills would honestly be the killer feature — a backup tool that actually verifies its own restores is the one I'd trust.

Thread Thread
 
matrixswarm profile image
Daniel •

Give me about a week for a first iteration. I’m starting with scheduled filesystem restore drills: restore into an isolated directory, verify the restored files, and report the drill’s own last-success timestamp.
Keep an eye out—and if you haven’t seen an update in a week, pester me. Hold me to it!