DEV Community

Daniel Ioni
Daniel Ioni

Posted on

Hardening a Contributor Verification VPS Without Breaking Production

Hardening a Contributor Verification VPS Without Breaking Production

Running contributor verification workloads on the same VPS as production services creates an uncomfortable engineering constraint:

you want stronger isolation, but you cannot casually restart, reconfigure, or “clean up” the server.

That was the situation we faced while turning a MyZubster VPS into a more structured Contributor Verification Node.

The goal was not a full infrastructure migration.

The goal was simpler:

reduce unnecessary network exposure, map what is actually running, preserve live workloads, and produce reproducible technical evidence for every change.

This post documents that process.


Starting point

The VPS was already doing several jobs at once:

  • production MyZubster gateway
  • nginx reverse proxy
  • PM2-managed applications
  • Docker-based contributor verification environments
  • MongoDB
  • Monero services
  • Bitcoin verification services
  • IPFS
  • Qdrant
  • Ollama
  • n8n
  • contributor-specific bridges and pilots

That means a command such as:

pm2 restart all
Enter fullscreen mode Exit fullscreen mode

would have been unacceptable.

The same applies to:

git pull
git reset --hard
docker system prune
ufw reset
Enter fullscreen mode Exit fullscreen mode

These commands may be routine on disposable infrastructure.

They are not routine on a mixed production and verification node.

So the first rule became:

Observe first. Change one component at a time. Verify immediately.


Step 1: map the listening services

The first useful artifact was not a deployment script.

It was a socket inventory.

Using:

ss -lntp
Enter fullscreen mode Exit fullscreen mode

we separated services into two broad categories.

Local-only services

Examples included:

127.0.0.1:5003   MyZubster Gateway
127.0.0.1:8787   BTC verifier
127.0.0.1:27017  MongoDB
127.0.0.1:6333   Qdrant
127.0.0.1:5678   n8n
127.0.0.1:8092   contributor bridge
Enter fullscreen mode Exit fullscreen mode

These already followed a reasonable internal-service model.

Globally bound services

Other applications were listening on all interfaces:

*:5002
0.0.0.0:5005
0.0.0.0:5173
Enter fullscreen mode Exit fullscreen mode

A globally bound socket does not automatically mean the application is publicly reachable.

The firewall still matters.

But it does increase the attack surface and makes the intended architecture less explicit.

So we inspected each port before changing anything.


Step 2: understand the reverse-proxy graph

The next question was:

Which services actually need to listen beyond localhost?

For the web frontend, nginx contained:

location / {
    proxy_pass http://localhost:5173;
}
Enter fullscreen mode Exit fullscreen mode

That established a clear dependency:

Internet
   ↓
nginx :443
   ↓
localhost:5173
Enter fullscreen mode Exit fullscreen mode

There was no architectural reason for the Node process on 5173 to listen on every interface.

The application originally used:

server.listen(PORT, '0.0.0.0', () => {
Enter fullscreen mode Exit fullscreen mode

We changed only that listener:

server.listen(PORT, '127.0.0.1', () => {
Enter fullscreen mode Exit fullscreen mode

Then:

node --check server-static.js
pm2 restart myzubster-web
Enter fullscreen mode Exit fullscreen mode

Only the affected PM2 process was restarted.


Step 3: verify before persisting

After the restart, we checked the actual socket:

ss -lntp | grep ':5173'
Enter fullscreen mode Exit fullscreen mode

The expected result was:

127.0.0.1:5173
Enter fullscreen mode Exit fullscreen mode

Then we tested the service directly:

curl -sS -o /dev/null \
  -w '5173 local -> %{http_code}\n' \
  http://127.0.0.1:5173/
Enter fullscreen mode Exit fullscreen mode

Result:

5173 local -> 200
Enter fullscreen mode Exit fullscreen mode

Next came nginx.

HTTP returned the expected redirect:

301
Enter fullscreen mode Exit fullscreen mode

and HTTPS was tested locally with the real virtual host:

curl -k -sS -o /dev/null \
  --resolve myzubster.com:443:127.0.0.1 \
  -w 'HTTPS nginx -> %{http_code}\n' \
  https://myzubster.com/
Enter fullscreen mode Exit fullscreen mode

Result:

HTTPS nginx -> 200
Enter fullscreen mode Exit fullscreen mode

Only after those checks did we persist PM2 state:

pm2 save
Enter fullscreen mode Exit fullscreen mode

This sequence matters:

backup
→ edit
→ syntax check
→ restart one service
→ socket check
→ local application check
→ reverse-proxy check
→ persist
Enter fullscreen mode Exit fullscreen mode

Not:

edit
→ restart everything
→ hope
Enter fullscreen mode Exit fullscreen mode

Step 4: repeat the same method for the API service

Port 5002 belonged to an Urban Lab / escrow application.

Its server started with:

httpServer.listen(PORT, () => {
Enter fullscreen mode Exit fullscreen mode

With Node.js, omitting the host can result in the service listening beyond localhost.

Before changing it, we inspected:

  • nginx references
  • frontend references
  • active TCP connections
  • WebSocket usage
  • application routes

The frontend already referenced the service internally through:

const API_TARGET = 'http://127.0.0.1:5002';
Enter fullscreen mode Exit fullscreen mode

No nginx route pointed directly at 5002.

No active direct connection required its public socket.

So the listener was changed to:

httpServer.listen(PORT, '127.0.0.1', () => {
Enter fullscreen mode Exit fullscreen mode

Again, only that PM2 application was restarted.

The health endpoint remained available locally.

The frontend proxy returned an application-level 404 for one test route instead of a 502 Bad Gateway.

That distinction is useful.

A 404 meant:

the upstream service was reachable, but the requested route did not exist.

A 502 would have meant:

the frontend proxy could no longer reach the upstream service.

The difference prevented us from incorrectly rolling back a valid hardening change.


An unexpected finding: the TAZ frontend

While checking WebSocket dependencies before restricting 5002, we found this in the frontend:

return `${protocol}//${window.location.host}/api/taz/live`;
Enter fullscreen mode Exit fullscreen mode

The browser was attempting to connect to:

wss://myzubster.com/api/taz/live
Enter fullscreen mode Exit fullscreen mode

It also expected:

GET /api/taz/dashboard
GET /api/taz/xmr/summary
WS  /api/taz/live
Enter fullscreen mode Exit fullscreen mode

At first this looked like a potential consequence of the hardening work.

It was not.

Testing the production gateway directly showed:

404 Cannot GET /api/taz/dashboard
404 Cannot GET /api/taz/xmr/summary
404 Cannot GET /api/taz/live
Enter fullscreen mode Exit fullscreen mode

That meant the mismatch already existed.

The frontend expected APIs that production did not currently expose.

This is an important operational lesson:

A broken feature discovered during hardening is not automatically a regression caused by hardening.

Reproduce the behavior against the actual runtime before drawing conclusions.


Nginx was not the root cause

The production nginx configuration was:

location / {
    proxy_pass http://localhost:5173;
}

location /api/ {
    proxy_pass http://localhost:5003;
}
Enter fullscreen mode Exit fullscreen mode

So /api/taz/live correctly reached the MyZubster gateway on 5003.

The gateway itself returned the 404.

Adding this blindly:

proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
Enter fullscreen mode Exit fullscreen mode

would therefore not have fixed the issue.

WebSocket headers cannot create an application endpoint that does not exist.


We also discovered an existing realtime system

The repository already contained a realtime implementation.

Its Socket.IO server exposes:

/realtime
Enter fullscreen mode Exit fullscreen mode

and the HTTP application mounts supporting routes under:

/api/realtime
Enter fullscreen mode Exit fullscreen mode

The realtime module includes an explicit function:

attachRealtimeServer(httpServer)
Enter fullscreen mode Exit fullscreen mode

This is valuable because it changes the architectural decision.

Instead of immediately creating another WebSocket subsystem for TAZ, the better question became:

Can the existing realtime gateway be reused?

That is usually preferable to maintaining:

Socket.IO realtime stack
+
raw WebSocket realtime stack
Enter fullscreen mode Exit fullscreen mode

unless there is a strong reason to support both.


Another runtime mismatch

There was one more subtle discovery.

The production systemd service launches:

/root/myzubster/scripts/start-gateway-systemd.js
Enter fullscreen mode Exit fullscreen mode

That script creates the HTTP server using:

const server = app.listen(port, host, ...)
Enter fullscreen mode Exit fullscreen mode

The realtime implementation, however, is designed to be explicitly attached using:

attachRealtimeServer(server)
Enter fullscreen mode Exit fullscreen mode

The regular backend startup path contains that attachment.

The production systemd startup path appears not to.

So the repository contains realtime code, but the active runtime path may not actually enable it.

This is exactly why runtime mapping matters.

Reading source code alone does not tell you what production is executing.


Repository architecture is not runtime architecture

One of the most useful lessons from this exercise was the difference between these statements:

"The repository supports realtime."
Enter fullscreen mode Exit fullscreen mode

and:

"The production process currently exposes realtime."
Enter fullscreen mode Exit fullscreen mode

They are not equivalent.

Likewise:

"The frontend contains a TAZ dashboard."
Enter fullscreen mode Exit fullscreen mode

does not imply:

"The production backend exposes TAZ data."
Enter fullscreen mode Exit fullscreen mode

For infrastructure work, the authoritative chain is closer to:

service manager
→ executable
→ startup script
→ imported application
→ bound socket
→ reverse proxy
→ externally observed response
Enter fullscreen mode Exit fullscreen mode

Not simply:

repository search
Enter fullscreen mode Exit fullscreen mode

Avoiding false evidence

Another important constraint was data integrity.

The frontend TAZ expects fields such as:

drinksServed
xmrReceived
robot.status
transactions
Enter fullscreen mode Exit fullscreen mode

During repository searches, we found simulations and unrelated robot data.

It would have been easy to connect those values and make the dashboard “work.”

We deliberately did not.

A live dashboard should not turn synthetic or simulated data into apparent production telemetry.

The correct rule is:

no real source → no real metric.

A degraded state is better than fabricated evidence.


What changed

The useful hardening results were small but concrete:

MyZubsterWeb
0.0.0.0:5173
→
127.0.0.1:5173
Enter fullscreen mode Exit fullscreen mode

and:

Urban Lab
*:5002
→
127.0.0.1:5002
Enter fullscreen mode Exit fullscreen mode

Production HTTPS continued to respond correctly.

The gateway was already:

127.0.0.1:5003
Enter fullscreen mode Exit fullscreen mode

The firewall remained active.

Docker verification workloads were not disturbed.

The production repository was not reset, pulled, cleaned, or replaced.


What we deliberately did not change

Some decisions are valuable precisely because they were postponed.

We did not:

  • restart every PM2 process
  • reset UFW
  • expose new ports
  • rewrite nginx prematurely
  • connect TAZ to an unrelated WebSocket server
  • turn simulation data into production telemetry
  • run a destructive Git operation on the production checkout
  • assume a repository feature was active in production
  • call a technical checkpoint a scientific or operational validation

This is part of evidence-first infrastructure work too.

A documented non-change can be an engineering result.


Contributor Verification Node mindset

The larger goal is to evolve a general-purpose VPS toward a more structured Contributor Verification Node.

That does not mean giving contributors unrestricted shell access.

A safer direction is:

SSH identity
   ↓
dedicated Unix identity
   ↓
dedicated workspace
   ↓
container or sandbox
   ↓
explicit mounts
   ↓
explicit network access
   ↓
reproducible verifier
   ↓
immutable evidence
Enter fullscreen mode Exit fullscreen mode

Contributor verification should produce evidence about a specific technical checkpoint.

It should not silently turn into production administration.


The main lesson

Hardening production infrastructure is often less about sophisticated security tooling and more about disciplined uncertainty reduction.

For each component:

What is running?

Why is it running?

Who reaches it?

Does it need a public socket?

What happens if I restrict it?

How do I prove that it still works?

What evidence supports the conclusion?
Enter fullscreen mode Exit fullscreen mode

That process exposed two unnecessary public binds.

It also exposed a frontend/backend contract that production did not currently satisfy.

Both findings came from the same method:

map reality first, then change the smallest possible thing.

That is the approach we want to keep using as MyZubster moves toward independent contributor nodes and reproducible verification infrastructure.


Current checkpoint

At this stage:

5173  localhost-only  ✅
5002  localhost-only  ✅
5003  localhost-only  ✅
nginx HTTPS           ✅
PM2 persistence       ✅
TAZ HTTP endpoints    missing
TAZ live WebSocket    missing
realtime subsystem    present in code
production realtime   wiring under verification
Enter fullscreen mode Exit fullscreen mode

The next step is not “deploy more.”

It is to determine how the existing realtime subsystem should be connected to the active production server, then decide whether TAZ should use that system or receive a small read-only adapter of its own.

One verified dependency at a time.

Top comments (0)