Operations memo, 09:00: The production swarm will notify us if anything goes wrong.
Operations memo, 09:07: The production swarm has stopped.
Operations memo, 09:08: We have received no notification. Morale remains inexplicably high.
A system that disappears cannot reliably announce its own disappearance. Somebody outside the room needs to take attendance.
MatrixSwarm’s harvester agent observes an assigned swarm through the target host’s matrixd process inventory, confirms repeated failures, and sends health notifications through the usual swarm relays. It can tell you that the expected universe is absent, that its agent count is below your minimum, or that the observation itself failed.
This article follows the current Phoenix workflow: one Harvester instance, one assigned universe, one SSH connection, observation only. We are issuing a clipboard, not authorizing a remote resurrection department.
What Harvester actually asks
The remote operation is fixed: matrixd list --json, using MatrixOS’s installed Python environment and script path. Harvester reads the structured inventory, finds the configured universe, and compares its agent count with your required minimum.
| Observation | Meaning |
|---|---|
| Assigned universe appears with enough agents | The configured inventory check passes |
| Assigned universe is absent | UNIVERSE_ABSENT |
| Universe appears with too few agents | A status such as AGENTS_3_BELOW_5
|
| SSH, command execution, or snapshot validation fails | An observation error, with a diagnostic status |
The minimum is a count, not a list of required agent identities. Five running processes are not proof that five specific jobs are working. A process can also remain alive while its application work is stuck.
Use Harvester for this inventory-level question. Pair it with the appropriate service, website, or application checks when you need to know whether useful work is happening.
Attendance and competence have always been separate columns on the spreadsheet.
Put the observer where it can survive
For a useful first deployment, put Harvester in a small observer swarm and point it at another already-deployed swarm.
Two universes on the same host can demonstrate a target-swarm failure. For host-outage coverage, use a separate observer host with a route to the target. If both share the failed machine, power supply, or network path, their opinions may become unavailable together.
Harvester runs on MatrixOS, so closing Phoenix does not stop its checks or the observer’s configured alert relays. Phoenix is your setup and inspection console. It does not need to spend the night staring at the process table.
Install MatrixOS, then build the observer
If the observer host needs MatrixOS, add and test its SSH record in Phoenix’s Registry, including the trusted host fingerprint. Open Railgun → Install MatrixOS… and choose Local Full Install with your source folder or Install from GitHub. For a fresh environment, choose Create new venv, install, and check the result. Installation requires root or passwordless sudo on that host.
The target must already have MatrixOS and the current matrixd list --json support. Keep its deployment record in your Phoenix vault; the assignment editor selects from those saved deployments.
Open Deploy → Swarm Workspaces and create a New workspace or Open an existing observer workspace. Beneath the matrix root, add:
| Agent | Assignment |
|---|---|
harvester |
Observe the selected remote universe |
matrix_https |
Accept Phoenix commands over HTTPS |
matrix_websocket |
Return replies and session traffic to Phoenix |
log_streamer |
Expose Agent Logs for diagnosis |
| Your chosen alert relay | Deliver health notifications |
Those ingress, egress, and log agents provide the connected Phoenix workflow. Harvester’s check itself reaches the target host over SSH. The target does not need an extra HTTPS/WebSocket pair merely to answer this inventory check.
Two SSH jobs are involved: Railgun deploys the observer to its host, and Harvester later connects from that host to the watched target. Keep those roles straight. “I could reach it from my laptop” is useful testimony, but your laptop has not been hired to run this check.
The assignment lives in the Registry
Harvester declares three constraints: packet_signing, symmetric_encryption, and harvester_assignment. Resolve them in the workspace using the appropriate Registry objects.
Create a Harvester Assignment Registry object. Its editor separates the watched deployment and connection policy from the agent’s ordinary runtime settings.
- Give the assignment a recognizable label.
- In Hive ↔ SSH, select the target’s saved deployment and its matching SSH Registry record.
- Leave Authority at Contact only (resurrection deferred).
- Set Minimum Agents to the smallest acceptable count for that target.
- Set failure/recovery thresholds and alert cooldown.
- For your first deployment, enable Send one confirmation when monitoring is established.
- Save the Registry object and assign it to Harvester’s harvester_assignment constraint.
Selecting the deployment attempts to select its recorded SSH object too. Validation checks the match when the deployment has that reference. Check the actual host, username, and universe before deploying; a carefully signed packet can still be addressed to the wrong building if your configuration is wrong.
The Registry assignment stores references and policy. When Phoenix compiles the directive, it resolves the deployment and SSH records from the unlocked vault and injects the selected connection profile and single target into Harvester’s encrypted directive.
Contact-only operation does not inject the watched swarm’s decryption key. It does require the assigned SSH authentication material needed to perform the observation.
Enable the actual checks
Open Harvester’s workspace Config editor and enable Enable matrixd checks. The shipped metadata starts disabled and has no assigned target.
A deployed but disabled Harvester is not monitoring quietly. It is following your instructions with devastating accuracy.
The runtime controls are:
| Setting | Default | Purpose |
|---|---|---|
| Enable matrixd checks | Off | Starts observation when the assignment is valid |
| Check Interval | 30 seconds | Delay between completed check cycles |
| Matrixd Timeout | 60 seconds | Total deadline for the supervised SSH checker |
| Alert Role | hive.alert |
Role used to find notification relays |
The assignment supplies these per-target defaults:
| Setting | Default |
|---|---|
| Minimum Agents | 1 |
| Failure Threshold | 3 consecutive failed observations |
| Recovery Threshold | 3 consecutive successful observations after a down state |
| Alert Cooldown | 300 seconds between eligible down reminders |
| Healthy Confirmation | Off until selected |
Choose a meaningful minimum. A six-agent deployment with a minimum of one can lose most of its intended workforce while still passing the count test. Conversely, a minimum above the legitimate running count guarantees complaints.
The current normal workflow provisions exactly one target. Multiple targets in one Harvester directive are rejected by its policy. Broader inventory fan-out is outside this release’s scope.
Pinned SSH and a bounded question
The target SSH record needs the expected SHA256 host-key fingerprint. Harvester verifies that identity, uses the provisioned authentication method, and runs its fixed inventory command. It does not accept an arbitrary command from a target row or roam through the network looking for swarms.
The transport uses Paramiko, whose SSH client documentation explains host-key handling and authentication options. Harvester supplies its own pinned-key policy instead of silently accepting an unfamiliar server.
Use credentials and routing that work from the observer host. A hostname resolved only on your workstation, an SSH agent available only in your desktop session, or a firewall rule allowing only your laptop will not become portable through optimism.
The fixed remote wrapper runs the MatrixOS inventory command as root, or uses noninteractive sudo -n when connecting as another user. Provision the target account to permit the required command; a successful interactive SSH login alone does not establish that permission. The check cannot answer a sudo password prompt.
The SSH work runs in a supervised checker process. The timeout covers startup, connection, authentication, and command execution. A stuck check is stopped and cleaned up before another one is allowed to proceed, while Harvester continues servicing its own heartbeat.
A remote machine may be unresponsive. That does not mean the observer should join it in a silent meditation retreat.
Launch the observer through Railgun
Save the workspace and click Deploy. Choose the observer host’s SSH record, a deployment label, observer universe, and runtime Linux user. Review the resolved directive and launch.
Railgun verifies the deployment connection, prepares the runtime account and applicable grants, and sends the sealed boot envelope over SSH stdin to MatrixD. MatrixD decrypts it in memory and starts the observer swarm. Check the launch result, then select the deployment and Connect.
The Harvester assignment inside that observer points to the watched target. Your observer universe and target universe are distinct responsibilities, even if a demonstration places them on the same server.
In Phoenix, use Agent Logs through log_streamer to inspect Harvester. The current agent has workspace configuration and Registry assignment editors; it does not provide a dedicated live Harvester inventory panel.
Look for the watching-target startup message and the initial observation. If you selected Healthy Confirmation and the first observation passes, verify the HEALTHY message at your chosen destination.
Read the roll call properly
Harvester tracks each observation through a small state machine:
| Event | When it happens |
|---|---|
| HEALTHY | Optional one-time confirmation on the first successful observation from the unknown state |
| DOWN | Consecutive failures reach the configured failure threshold |
| DOWN_REMINDER | The target remains down and the alert cooldown permits another reminder |
| RECOVERY | Consecutive successes reach the recovery threshold after a down state |
An isolated failed observation resets the success streak; an isolated successful observation resets the failure streak. Once down, the target needs the configured run of successful observations to be reported recovered.
The 30-second interval comes after check work. Three failures do not guarantee a precisely timed 90-second alert, especially if SSH is consuming its timeout. Treat the thresholds as consecutive observations rather than stopwatch promises.
Most importantly, DOWN can mean “the inventory check failed.” Read the accompanying status. UNIVERSE_ABSENT from a valid snapshot and SSH_AUTH_FAILED are very different findings. One says the selected universe was not in the inventory; the other says the observer could not establish the observation.
The uniform may say Harvester. It still has to show its evidence.
Wire the notice to a human
Harvester sends notifications directly to endpoints advertising its Alert Role, normally hive.alert. For this route you can use Slack, Discord, Telegram, or email through slack_relay, discord_relay, telegram_relay, or email_send.
Configure the relay’s destination and Registry requirements. Discord, Telegram, and email support optional signed alert-payload encryption through encrypt_alerts, with matching keys and compatible decryption support on the receiving side. This is separate from SSH or message-transport encryption.
For selective delivery, use a dedicated role shared by Harvester and the intended receiving relay. Otherwise, it considers all endpoints discovered for its configured role.
Harvester retains one current pending notification per target and retries failed swarm dispatches on later checks. It tracks successful recipient handoffs so an unavailable second relay does not cause needless repeats to the first. State changes can supersede an older pending notice; this is a current-health notification mechanism, not an unlimited historical outbox.
A successful packet handoff does not prove that a human’s inbox received or displayed the message. Verify the full route during setup.
Harvester’s condition counters and pending-delivery bookkeeping are in memory. Restarting the observer resets them. Do not assume a restart preserves outage history, retry history, or the once-per-run healthy confirmation state.
About the resurrection paperwork
Some older descriptions discuss a contact-and-resurrect capability. The current assignment editor exposes Contact only (resurrection deferred), disables the resurrection controls, and compiles automatic recovery as false. The running agent also keeps automatic recovery disabled.
Accordingly, this walkthrough promises observation and notification. A RECOVERY event means Harvester observed the target return to its required state. It does not mean Harvester restarted it.
Treat resurrection as deferred functionality. The existence of a form field in the filing cabinet does not mean the department has permission to operate a defibrillator.
A drill that proves the observer stays awake
Use a disposable target swarm with a known agent count and an observer that will remain running independently.
- Deploy the observer with Healthy Confirmation enabled. Verify its initial log and the actual HEALTHY notification.
- Stop only the disposable target swarm through your normal lifecycle procedure.
- Leave the observer running and wait for its configured failure confirmations.
- Verify DOWN and inspect the status that explains the failed observation.
- Relaunch the target through Railgun using the same expected universe and enough agents.
- Leave the observer untouched while successful observations accumulate, then verify RECOVERY.
If you restart the observer halfway through the drill, its state returns to unknown. That tests a new observer startup, not the down-to-recovered transition you were trying to demonstrate.
No need to sabotage production to establish that the clipboard works.
When attendance does not add up
| Symptom | First check |
|---|---|
| “Disabled by policy” | Enable matrixd checks in the workspace Config and redeploy |
| No valid target / fail closed | Assignment, referenced deployment, matching SSH object, and single-target limit |
SSH_HOST_KEY_MISMATCH |
Verify the real target host identity through a trusted channel; update the Registry only after checking the change |
SSH_AUTH_FAILED |
Provisioned credentials and authentication available to the observer |
SSH_DNS_ERROR / connection failure / timeout |
Name resolution and network access from the observer host |
MATRIXD_CHECK_FAILED |
Target installation, fixed command path, noninteractive permissions, and command output |
UNIVERSE_ABSENT |
Exact assigned universe and whether that target is actually running |
| Agent count is below minimum | Expected count versus legitimate target processes |
| Alerts remain pending | Receiving role, service discovery, relay health, and packet-dispatch logs |
Start with one assigned target, one meaningful minimum, and one verified relay. Keep the observer’s own availability visible too; every watchtower is still made of computers.
You now have something more useful than “I think that swarm is still up”: a repeated observation, an explicit failure policy, and a message from somewhere capable of sending it.
Victory Always. Roll call is now somebody’s actual job.
🌐 Links & Resources:
Try MatrixSwarm: https://matrixswarm.com
Join the Community / Discord: https://discord.gg/2USbWVBVV
Download Server: https://github.com/matrixswarm/matrixswarm
Youtube: https://www.youtube.com/channel/UCMjiY4_-W2KP5fHXO0eC2ug
Top comments (1)
tr.ee/dev-to
Some comments have been hidden by the post's author - find out more