Hardening a Contributor Verification VPS Without Breaking Production
Running contributor verification workloads on the same VPS as production services creates an uncomfortable engineering constraint:
you want stronger isolation, but you cannot casually restart, reconfigure, or “clean up” the server.
That was the situation we faced while turning a MyZubster VPS into a more structured Contributor Verification Node.
The goal was not a full infrastructure migration.
The goal was simpler:
reduce unnecessary network exposure, map what is actually running, preserve live workloads, and produce reproducible technical evidence for every change.
This post documents that process.
Starting point
The VPS was already doing several jobs at once:
- production MyZubster gateway
- nginx reverse proxy
- PM2-managed applications
- Docker-based contributor verification environments
- MongoDB
- Monero services
- Bitcoin verification services
- IPFS
- Qdrant
- Ollama
- n8n
- contributor-specific bridges and pilots
That means a command such as:
pm2 restart all
would have been unacceptable.
The same applies to:
git pull
git reset --hard
docker system prune
ufw reset
These commands may be routine on disposable infrastructure.
They are not routine on a mixed production and verification node.
So the first rule became:
Observe first. Change one component at a time. Verify immediately.
Step 1: map the listening services
The first useful artifact was not a deployment script.
It was a socket inventory.
Using:
ss -lntp
we separated services into two broad categories.
Local-only services
Examples included:
127.0.0.1:5003 MyZubster Gateway
127.0.0.1:8787 BTC verifier
127.0.0.1:27017 MongoDB
127.0.0.1:6333 Qdrant
127.0.0.1:5678 n8n
127.0.0.1:8092 contributor bridge
These already followed a reasonable internal-service model.
Globally bound services
Other applications were listening on all interfaces:
*:5002
0.0.0.0:5005
0.0.0.0:5173
A globally bound socket does not automatically mean the application is publicly reachable.
The firewall still matters.
But it does increase the attack surface and makes the intended architecture less explicit.
So we inspected each port before changing anything.
Step 2: understand the reverse-proxy graph
The next question was:
Which services actually need to listen beyond localhost?
For the web frontend, nginx contained:
location / {
proxy_pass http://localhost:5173;
}
That established a clear dependency:
Internet
↓
nginx :443
↓
localhost:5173
There was no architectural reason for the Node process on 5173 to listen on every interface.
The application originally used:
server.listen(PORT, '0.0.0.0', () => {
We changed only that listener:
server.listen(PORT, '127.0.0.1', () => {
Then:
node --check server-static.js
pm2 restart myzubster-web
Only the affected PM2 process was restarted.
Step 3: verify before persisting
After the restart, we checked the actual socket:
ss -lntp | grep ':5173'
The expected result was:
127.0.0.1:5173
Then we tested the service directly:
curl -sS -o /dev/null \
-w '5173 local -> %{http_code}\n' \
http://127.0.0.1:5173/
Result:
5173 local -> 200
Next came nginx.
HTTP returned the expected redirect:
301
and HTTPS was tested locally with the real virtual host:
curl -k -sS -o /dev/null \
--resolve myzubster.com:443:127.0.0.1 \
-w 'HTTPS nginx -> %{http_code}\n' \
https://myzubster.com/
Result:
HTTPS nginx -> 200
Only after those checks did we persist PM2 state:
pm2 save
This sequence matters:
backup
→ edit
→ syntax check
→ restart one service
→ socket check
→ local application check
→ reverse-proxy check
→ persist
Not:
edit
→ restart everything
→ hope
Step 4: repeat the same method for the API service
Port 5002 belonged to an Urban Lab / escrow application.
Its server started with:
httpServer.listen(PORT, () => {
With Node.js, omitting the host can result in the service listening beyond localhost.
Before changing it, we inspected:
- nginx references
- frontend references
- active TCP connections
- WebSocket usage
- application routes
The frontend already referenced the service internally through:
const API_TARGET = 'http://127.0.0.1:5002';
No nginx route pointed directly at 5002.
No active direct connection required its public socket.
So the listener was changed to:
httpServer.listen(PORT, '127.0.0.1', () => {
Again, only that PM2 application was restarted.
The health endpoint remained available locally.
The frontend proxy returned an application-level 404 for one test route instead of a 502 Bad Gateway.
That distinction is useful.
A 404 meant:
the upstream service was reachable, but the requested route did not exist.
A 502 would have meant:
the frontend proxy could no longer reach the upstream service.
The difference prevented us from incorrectly rolling back a valid hardening change.
An unexpected finding: the TAZ frontend
While checking WebSocket dependencies before restricting 5002, we found this in the frontend:
return `${protocol}//${window.location.host}/api/taz/live`;
The browser was attempting to connect to:
wss://myzubster.com/api/taz/live
It also expected:
GET /api/taz/dashboard
GET /api/taz/xmr/summary
WS /api/taz/live
At first this looked like a potential consequence of the hardening work.
It was not.
Testing the production gateway directly showed:
404 Cannot GET /api/taz/dashboard
404 Cannot GET /api/taz/xmr/summary
404 Cannot GET /api/taz/live
That meant the mismatch already existed.
The frontend expected APIs that production did not currently expose.
This is an important operational lesson:
A broken feature discovered during hardening is not automatically a regression caused by hardening.
Reproduce the behavior against the actual runtime before drawing conclusions.
Nginx was not the root cause
The production nginx configuration was:
location / {
proxy_pass http://localhost:5173;
}
location /api/ {
proxy_pass http://localhost:5003;
}
So /api/taz/live correctly reached the MyZubster gateway on 5003.
The gateway itself returned the 404.
Adding this blindly:
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
would therefore not have fixed the issue.
WebSocket headers cannot create an application endpoint that does not exist.
We also discovered an existing realtime system
The repository already contained a realtime implementation.
Its Socket.IO server exposes:
/realtime
and the HTTP application mounts supporting routes under:
/api/realtime
The realtime module includes an explicit function:
attachRealtimeServer(httpServer)
This is valuable because it changes the architectural decision.
Instead of immediately creating another WebSocket subsystem for TAZ, the better question became:
Can the existing realtime gateway be reused?
That is usually preferable to maintaining:
Socket.IO realtime stack
+
raw WebSocket realtime stack
unless there is a strong reason to support both.
Another runtime mismatch
There was one more subtle discovery.
The production systemd service launches:
/root/myzubster/scripts/start-gateway-systemd.js
That script creates the HTTP server using:
const server = app.listen(port, host, ...)
The realtime implementation, however, is designed to be explicitly attached using:
attachRealtimeServer(server)
The regular backend startup path contains that attachment.
The production systemd startup path appears not to.
So the repository contains realtime code, but the active runtime path may not actually enable it.
This is exactly why runtime mapping matters.
Reading source code alone does not tell you what production is executing.
Repository architecture is not runtime architecture
One of the most useful lessons from this exercise was the difference between these statements:
"The repository supports realtime."
and:
"The production process currently exposes realtime."
They are not equivalent.
Likewise:
"The frontend contains a TAZ dashboard."
does not imply:
"The production backend exposes TAZ data."
For infrastructure work, the authoritative chain is closer to:
service manager
→ executable
→ startup script
→ imported application
→ bound socket
→ reverse proxy
→ externally observed response
Not simply:
repository search
Avoiding false evidence
Another important constraint was data integrity.
The frontend TAZ expects fields such as:
drinksServed
xmrReceived
robot.status
transactions
During repository searches, we found simulations and unrelated robot data.
It would have been easy to connect those values and make the dashboard “work.”
We deliberately did not.
A live dashboard should not turn synthetic or simulated data into apparent production telemetry.
The correct rule is:
no real source → no real metric.
A degraded state is better than fabricated evidence.
What changed
The useful hardening results were small but concrete:
MyZubsterWeb
0.0.0.0:5173
→
127.0.0.1:5173
and:
Urban Lab
*:5002
→
127.0.0.1:5002
Production HTTPS continued to respond correctly.
The gateway was already:
127.0.0.1:5003
The firewall remained active.
Docker verification workloads were not disturbed.
The production repository was not reset, pulled, cleaned, or replaced.
What we deliberately did not change
Some decisions are valuable precisely because they were postponed.
We did not:
- restart every PM2 process
- reset UFW
- expose new ports
- rewrite nginx prematurely
- connect TAZ to an unrelated WebSocket server
- turn simulation data into production telemetry
- run a destructive Git operation on the production checkout
- assume a repository feature was active in production
- call a technical checkpoint a scientific or operational validation
This is part of evidence-first infrastructure work too.
A documented non-change can be an engineering result.
Contributor Verification Node mindset
The larger goal is to evolve a general-purpose VPS toward a more structured Contributor Verification Node.
That does not mean giving contributors unrestricted shell access.
A safer direction is:
SSH identity
↓
dedicated Unix identity
↓
dedicated workspace
↓
container or sandbox
↓
explicit mounts
↓
explicit network access
↓
reproducible verifier
↓
immutable evidence
Contributor verification should produce evidence about a specific technical checkpoint.
It should not silently turn into production administration.
The main lesson
Hardening production infrastructure is often less about sophisticated security tooling and more about disciplined uncertainty reduction.
For each component:
What is running?
Why is it running?
Who reaches it?
Does it need a public socket?
What happens if I restrict it?
How do I prove that it still works?
What evidence supports the conclusion?
That process exposed two unnecessary public binds.
It also exposed a frontend/backend contract that production did not currently satisfy.
Both findings came from the same method:
map reality first, then change the smallest possible thing.
That is the approach we want to keep using as MyZubster moves toward independent contributor nodes and reproducible verification infrastructure.
Current checkpoint
At this stage:
5173 localhost-only ✅
5002 localhost-only ✅
5003 localhost-only ✅
nginx HTTPS ✅
PM2 persistence ✅
TAZ HTTP endpoints missing
TAZ live WebSocket missing
realtime subsystem present in code
production realtime wiring under verification
The next step is not “deploy more.”
It is to determine how the existing realtime subsystem should be connected to the active production server, then decide whether TAZ should use that system or receive a small read-only adapter of its own.
One verified dependency at a time.
Top comments (0)