The server is running. The port is open. The homepage returns 200.
Unfortunately, the homepage is now a maintenance apology wearing the company logo. The stylesheet has vanished. Checkout looks like a tax form transmitted by fax.
“Everything is green,” says the dashboard.
The customers have submitted a competing assessment.
MatrixSwarm’s site_sentinel checks the public page, required content, a sample of linked assets, an optional origin endpoint, and certificate lifetime. It also reads local access logs and measures the host’s resource pressure. Confirmed conditions become structured evidence for Forensic Detective.
Think of it as the building inspector who actually opens the door instead of admiring the illuminated EXIT sign from the car park.
What the inspection covers
| Evidence | What it helps you notice |
|---|---|
| Public HTTP response | The endpoint is unavailable or responding unusually slowly |
| Required text marker | A successful response contains the wrong or incomplete page |
| Linked CSS, JavaScript, and images | The page responds, but selected dependencies fail |
| Optional origin comparison | Public and origin paths disagree |
| TLS certificate lifetime | A currently valid public certificate is approaching expiration |
| Access-log patterns | High request rates, many distinct paths, and error-heavy or bot-like traffic |
| Local CPU, memory, and load | The monitoring host is under pressure |
Site Sentinel observes and reports. It does not restart Apache, ban addresses, purge caches, repair deployments, or quarantine files. Those are separate jobs. Its contribution is better evidence before somebody announces that rebooting everything is “the quickest option.”
Install the runtime and build the inspection team
If the host does not have MatrixOS, add and test its SSH connection in Phoenix’s Registry, including the trusted host fingerprint. Open Railgun → Install MatrixOS… and choose Local Full Install with your source folder or Install from GitHub. For a fresh environment, choose Create new venv, install, and verify the result. Installation requires root or passwordless sudo on the target.
Then open Deploy → Swarm Workspaces. Create a New workspace or Open an existing one. Keep the matrix root and add:
| Agent | Responsibility |
|---|---|
site_sentinel |
Collect and classify website and host evidence |
forensic_detective |
Receive status events, investigate critical incidents, and initiate alerts |
matrix_https |
Accept Phoenix commands through HTTPS ingress |
matrix_websocket |
Return swarm replies and session traffic to Phoenix |
log_streamer |
Make Agent Logs available in the cockpit |
| An alert relay | Deliver the resulting notifications to your chosen destination |
Resolve each agent’s declared Registry constraints and the HTTPS/WSS connection requirements. Forensic Detective needs its own Persistent State profile for the encrypted incident journal. Retain that profile for later deployments.
For the first exercise, open Forensic Detective’s Config and turn off Enable Oracle Analysis. Its normal investigation and immediate critical-incident alert do not require an LLM. You can later add oracle and configure the provider for an additional analysis step.
The inspection should work before we invite a consultant to narrate it.
Choose where the observer lives
URL checks run from the Site Sentinel host. Access logs and CPU/memory/load are also read from that host.
For a combined view of a web server’s pages, logs, and resource pressure, deploy the agent on that server and grant the runtime account the necessary read access to the configured log files. If you put it on a separate monitoring server, its resource metrics describe the monitoring server; pointing a URL at another machine does not move the CPU sensor there.
A separate observer is useful for external reachability, while an on-host observer provides local context. A monitor that loses its own host cannot report that host’s outage through sheer professional commitment.
Tell it what a healthy page looks like
Open Site Sentinel’s workspace Config editor. Under Public and Origin Targets, add a row:
| Field | Example or purpose |
|---|---|
url |
Your public page, such as https://shop.example.com/catalog
|
origin_url |
Optional direct origin URL for the equivalent page |
host_header |
Optional virtual-host name for the origin request |
expect |
A stable text marker that must appear in the returned page |
note |
A short description such as “Public catalog” |
Replace the example domain with a site you operate. Pick a marker tied to useful content, not something also present on every error page. “Catalog available” is a better diagnostic than the company name if the company name appears on the maintenance screen too.
The marker check is a case-insensitive substring search in the response text, using the first two million characters. It is not a DOM selector, screenshot comparison, checksum baseline, or JavaScript execution test. Leaving expect empty disables that particular content check.
A heading created only after browser JavaScript runs may not exist in the HTML fetched by this agent. Choose server-returned content for the first test.
The public door and the service entrance
The optional origin request helps compare what the public path serves with what the backend serves. When host_header is empty, the origin request uses the public URL’s hostname as its HTTP Host header.
For this comparison to mean something, use equivalent paths and make sure the origin URL actually reaches the origin rather than circling through the same CDN again.
HTTPS still requires valid certificate trust. An HTTP Host header does not change the TLS hostname used by the URL. Use an origin hostname whose certificate and routing are valid; certificate verification remains enabled, consistent with Requests’ documented behavior.
Here is the agent’s classification vocabulary:
| Observation | Classification |
|---|---|
| Public response succeeds, required marker is missing | CRITICAL: PUBLIC_CONTENT_MISMATCH |
| Public content passes, a sampled asset fails | WARNING: PUBLIC_ASSET_FAILURE |
| Public checks pass, configured origin fails | WARNING: ORIGIN_DEGRADED |
| Public request fails, origin checks pass | WARNING: EDGE_DEGRADED |
| Public request fails without a healthy origin result | CRITICAL: SITE_UNAVAILABLE |
| Otherwise healthy public response exceeds latency threshold | WARNING: PUBLIC_LATENCY |
| Otherwise healthy target has a certificate near expiration | WARNING: TLS_EXPIRING |
These labels describe observations, not a proven root cause. EDGE_DEGRADED is a reason to investigate the public path. It does not authorize a press release declaring the CDN guilty.
Also, if you did not configure an origin, there is no origin evidence. Read the underlying observations before concluding both paths failed.
“The HTML arrived” is a very low bar
Enable Verify linked CSS, JavaScript, and images to sample assets referenced directly in the HTML. The parser considers stylesheet links, script sources, and image sources. By default it checks up to 8 distinct HTTP/HTTPS assets; the configurable maximum is 25.
This catches useful failures such as a missing stylesheet or an unavailable script. It follows referenced URLs, so an asset may live on another host. It does not run scripts, inspect their behavior, follow every dependency inside CSS, or perform a complete browser journey.
A JavaScript file returning 200 can still contain a bug. Site Sentinel has an inspection clipboard, not supernatural access to the intentions of your frontend team.
The latency measurement covers the public page request, not browser rendering or the complete asset pass. Choose a threshold appropriate to the page and the observer’s location.
Traffic analysis: who has been rattling the doors?
Under Load and Scraper Detection, enable traffic analysis and list real access-log files on the deployment host, one per line. The palette suggests /var/log/apache2/access.log; replace or remove it when that is not your server’s actual path.
The runtime account must be able to read the files. At first open, the reader starts at the end; it does not replay yesterday’s traffic as an exciting new emergency. Subsequent cycles read appended entries, with handling for rotation and truncation.
Supported inputs include common/combined-style access logs and compatible JSON records. The analysis tracks request rate, distinct paths, error ratio, and bot-like user-agent strings for the busiest addresses.
The supplied settings include:
| Setting | Default |
|---|---|
| Traffic window | 60 seconds |
| Top addresses examined | 10 |
| Per-address warning / critical rate | 120 / 300 requests per minute |
| Total warning / critical rate | 600 / 1,500 requests per minute |
| Distinct-path threshold | 80 |
| Error-ratio threshold | 0.35, meaning 35% |
The per-address warning rate is not sufficient on its own for the aggressive-crawler warning. It also needs many distinct paths, enough errors, or a bot-like user agent on at least half the requests. The critical per-address rate can independently produce a critical classification. Total traffic has its own thresholds.
These are tuning defaults, not universal definitions of an attack. A busy API customer, a search crawler, and a broken client retry loop can all be noisy for different reasons. The agent identifies patterns; the human investigates the explanation.
Rates use when the agent reads entries into its observation window, rather than reconstructing a precise event-time timeline from the log timestamps. Batched log writes can therefore affect the apparent pattern.
Do not nominate the reverse proxy as Public Enemy Number One
The editor can prefer forwarded client addresses when compatible JSON fields or an appended forwarded-address field are available. That only helps if your proxy and logging configuration establish which upstream sources are trusted.
For example, Nginx’s real-IP module provides trusted-source configuration. Site Sentinel does not independently verify a logged forwarded address. Arbitrary client-supplied headers are not authenticated identity papers.
Set the ignored-address list thoughtfully. It defaults to loopback addresses. Ignored traffic is removed from these summaries, so adding half the internet to make the graph quieter is technically effective and operationally silly.
Pressure readings and fewer false alarms
Local resource warnings default to 80% CPU, 90% memory, or 1.5 load per CPU. Critical thresholds are 95% CPU, 97% memory, or 3.0 load per CPU. The load figure uses the one-minute system load divided by CPU count.
The agent combines these readings with traffic evidence for a separate pressure condition. CPU pressure alone can trigger it, even when no access log is configured. Missing logs do not magically mean zero real traffic; they mean that traffic evidence is unavailable.
The normal check interval is 30 seconds, request timeout 8 seconds, failure confirmation 3 cycles, and recovery confirmation 2 cycles. TLS warning defaults to 21 days and summary logging to 300 seconds.
Checks run through their work and then sleep for the interval. Multiple targets, origin probes, TLS checks, and slow assets add time. Three confirmations do not promise an exact 90-second incident-detection deadline.
Once a condition is active, the agent reports activation or a change in severity rather than repeating the same incident every cycle. Enough healthy observations produce an INFO RECOVERED report. Its counters and active-condition state live in memory and reset when the agent restarts.
Connect the evidence to the alarm
Keep Forensic Report Role at hive.forensics.data_feed for the standard arrangement. Forensic Detective advertises the corresponding status-ingestion handler.
The route is:
Site Sentinel → Forensic Detective → hive.alert relays → your destination.
Do not point Site Sentinel’s forensic role straight at a generic alert relay and expect interchangeable packet handling. It sends a structured status event. Detective turns qualifying incidents into notification packets.
In the current Detective implementation, confirmed CRITICAL events trigger incident investigation and alerts, subject to its cooldown. WARNING and INFO recovery events are ingested as evidence; they do not automatically generate their own external notifications. Site Sentinel events also share Detective’s service-level alert cooldown, so several targets failing close together may not each produce a separate immediate alert.
Use any of the usual relays: Slack, Discord, Telegram, or email. Discord, Telegram, and email support optional signed payload encryption with encrypt_alerts and the matching configured keys/decryption support. That is separate from transport encryption. Resolve the relay’s destination and Registry requirements, then verify actual receipt.
An alert route drawn beautifully on a diagram is still a hypothesis until the intended inbox makes a noise.
Deploy and run a harmless inspection drill
Save the workspace configuration and click Deploy. Select the target SSH record, deployment label, universe, and runtime account; review the resolved directive and applicable access grants. Railgun sends the sealed boot envelope over pinned SSH to MatrixD, which decrypts it in memory and launches the swarm. Check the result, then Connect to the deployment.
This agent’s current workflow uses its workspace Config editor and Agent Logs through log_streamer. It does not ship a dedicated live Site Sentinel control panel.
On a staging page you control:
- Serve a stable marker and a small stylesheet. Confirm normal summaries in Site Sentinel’s logs.
- Remove the marker while still returning HTTP 200.
- Wait for the configured failure confirmations and look for PUBLIC_CONTENT_MISMATCH.
- Check Detective’s ingestion/incident logs and verify the critical alert at your chosen destination.
- Restore the marker. After the recovery confirmations, check for RECOVERED in the evidence path.
- Separately test a missing sampled stylesheet and observe the warning classification, remembering the default external-alert policy.
You have now demonstrated the actual reason for using this agent: a server can answer politely while delivering the wrong thing.
If the inspection report looks odd
| Symptom | Check first |
|---|---|
| No targets checked | Add target rows; the default list is empty |
| Marker always fails | Raw returned HTML, redirects, maintenance templates, and browser-only content |
| Origin fails while public works | Origin route, equivalent path, Host header, and TLS hostname/trust |
| Traffic analysis is idle | Real log paths, read permissions, compatible format, and new appended entries |
| CPU figures seem unrelated | Which machine actually runs Site Sentinel |
| No forensic report arrives | Detective’s role, running state, and Persistent State assignment |
| Warning/recovery reaches logs but not chat | The current critical-only escalation policy |
Start with one meaningful page, one marker, and one verified notification route. Expand the inspection as you learn what normal looks like.
The dashboard may still say 200 OK. This time, somebody checked whether the building had a floor.
Victory Always. Inspect the premises before declaring the grand reopening.
🌐 Links & Resources:
Try MatrixSwarm: https://matrixswarm.com
Join the Community / Discord: https://discord.gg/2USbWVBVV
Download Server: https://github.com/matrixswarm/matrixswarm
Youtube: https://www.youtube.com/channel/UCMjiY4_-W2KP5fHXO0eC2ug
Top comments (0)