ppmmx Standalone Stress-Test Report
Concurrent WHEP viewer capacity on 2 vCPU nodes
| Report | ppmmx standalone load-capacity and VPS-egress test |
| Version | v4 (final) |
| Date | 2026-10-07 |
| Software under test | ppmmx v1.19.1rc064, NODE_ROLE_STANDALONE
|
| Load generator |
whep-loadgen (pion/webrtc, RTP read-and-discard) |
| Providers compared | LightNode (Manila / Taipei / Singapore / Tokyo); AWS Lightsail (Singapore) |
Abstract
ppmmx in standalone mode was stress-tested for concurrent WHEP viewer
capacity on small cloud VPS instances, over two ingest paths (single-track SRT,
then multi-track WHIP with simultaneous HEVC + H.264), and across two providers.
The central findings are:
- The first provider tested (LightNode) never delivered its advertised 1 Gbps. Across three separately provisioned batches and four cities, measured egress was ~50??06 Mbps (per-instance, per-direction, both TCP and UDP, and to a third-party endpoint). A node on that network saturates the link at roughly 39 concurrent 2.5 Mbps viewers ??long before its CPU.
- On a provider that delivered real bandwidth (AWS Lightsail: ~2.4 Gbps public, ~5 Gbps private), a smaller 2 vCPU / 2 GB node completed 250 concurrent WHEP viewers with zero disconnects. The 300-viewer step ran the node out of memory, not bandwidth and not CPU.
- Multi-track WHIP ingest costs materially more than single-track SRT ??an ingest-only baseline of ~16% of a core and ~108 MB, and roughly double the per-viewer CPU at the top stable step.
- Memory, not CPU and not the network, is what a small standalone node hits first. Budget from measured egress and measured per-reader memory, not from the plan label.
Section 6 gives a side-by-side provider comparison.
1. Background
ppmmx is a low-latency live-video CDN node. The standalone role is the
self-contained deployment: a single instance both ingests a publisher
(WHIP/SRT/RTMP) and serves viewers directly over WebRTC (WHEP), with no
forwarding to other mmx nodes and no control-plane dependency.
The operational question is: how many concurrent viewers does one box hold,
and what limits it? This report answers that for the smallest sensible
instance sizes and two real ingest paths.
2. Method
2.1 Procedure
A stepped concurrency ladder is run against the node:
12 (10-min soak, Round 1) -> 25 -> 50 -> 75 -> 100 -> 150 -> 200 -> 250 -> 300
At each step, the target number of WHEP readers is opened, held for 2 minutes,
then closed together. A step is recorded PASS only if every session
establishes and closes with no disconnects; the ladder stops at the first step
that fails. Node CPU, RSS, available memory and the node's /metrics are
sampled throughout.
2.2 Tooling and measurement
-
Readers:
whep-loadgen, a purpose-built Go client that opens N realpion/webrtcPeerConnections against the WHEP endpoint. Each session completes full SDP / ICE / DTLS / SRTP and then reads RTP and discards it (no decoding), so the generator is not measuring its own video pipeline. Sessions are spread evenly across two load generators. -
Server sampling: process CPU and RSS every 2 s from
/proc/<pid>/{stat,status}(CPU as a delta, i.e. instantaneous, not a lifetime average); available memory and load average from/proc. -
Node
/metrics: session count, outbound bytes/packets, RTP loss, discarded frames. - Per-session: connect time, first-frame time, disconnects, bytes, packets.
- Ramp: 500 ms between session starts.
2.3 Test stream
-
Round 1: one H.264 1280?720@30, ~2.5 Mbps stream, generated with
ffmpeg(testsrc2+libx264) and published over SRT. Video only (no audio). - Round 2: the real product ingest path ??ppobs publishing over WHIP with simultaneous HEVC + H.264 multi-track (three H.264 simulcast layers + three HEVC layers per stream, each session also carrying Opus), pushed from a genuine consumer broadband uplink.
3. Round 1 ??SRT ingest, single H.264 track
Environment: 2 vCPU / 4 GB node and two 2 vCPU / 4 GB load generators, all on
LightNode (Manila), same datacenter.
| Step | Sessions | Verdict | Connected | Disconnects | mmx CPU avg / peak | mmx RSS avg / peak | Egress | Loss |
|---|---|---|---|---|---|---|---|---|
| baseline | 0 | ?? | ?? | ?? | 0.5% / 2.5% | 46 / 48 MB | ?? | 0 |
| 12 (10 min) | 12 | PASS | 12/12 | 0 | 26.4% / 30% | 114 / 116 MB | ~31 Mbps | 0 |
| 25 (2 min) | 25 | PASS | 25/25 | 0 | 39.5% / 45% | 167 / 175 MB | ~64 Mbps | 0 |
| 50 (2 min) | 50 | PASS | 50/50 | 0 | 63.1% / 79.5% | 274 / 294 MB | ~128 Mbps | 0 |
| 75 (2 min) | 75 | FAIL | 71/75 | 5 | 134% / 177% | 618 / 1068 MB | (overloaded) | 0* |
CPU is per-process where 100% = one full core. *RTP loss stays 0, but at
75 the server logs reader is too slow, discarding 27 frames ??the degradation
appears as dropped frames, not on-wire loss.
From 12 ??50 viewers, CPU rises almost linearly:
mmx CPU% ??15% + 0.97% ? concurrent_viewers (1 core = 100%)
about 1% of a core per 2.5 Mbps viewer plus ~15% fixed overhead. The last
stable step (50) used 0.63 of a core. At 75 the model breaks: measured CPU
reached 134% average / 177% peak (a two-core box at ~88%, run-queue load
~2.0); the B2 shard could not establish its last sessions (WHEP POST timeouts
and ICE timeouts), first-frame latency for that shard rose from ~0.3 s to
3.1 s average / 9.6 s worst, and the server began discarding frames for slow
readers.
Measured egress is ~2.55 Mbps per viewer. At the last stable step that is
~128 Mbps of a 1 Gbps port (~13%). (Caveat: this is the server's own send
count and may include bytes the provider later dropped; see ?5.)
RSS grows roughly linearly with readers (~14??6 MB/reader) and is reclaimed
after sessions close: the Go runtime returned memory over ~2?? minutes back
toward a ~170 MB baseline (from a 46 MB cold start), mem_alloc fell from
~486 MB to ~77 MB, and goroutines stayed flat at 18 ??no leak.
4. Round 2 ??ppobs WHIP ingest, HEVC + H.264 multi-track
Same node and ladder, but the real ingest path: ppobs over WHIP, three H.264 +
three HEVC simulcast layers and Opus per session, from a consumer uplink.
The uplink is a real internet path, not the LAN-like SRT link of Round 1. The
base layer arrived at 0% loss; the upper simulcast layers lost ~20% ??normal
behaviour for a home uplink once the stream exceeds the available upstream. Loss
on that link is a normal network condition, not a test artefact, but it means
the node's CPU below includes retransmission/discard work.
| Step | Verdict | Connected | Disconnects | mmx CPU avg / peak | mmx RSS avg / peak |
|---|---|---|---|---|---|
| ingest only (0 viewers) | ?? | ?? | ?? | 15.8% | 108 MB |
| 12 (3 min) | PASS | 12/12 | 0 | 32% / 39.5% | 118 / 127 MB |
| 25 (2 min) | PASS | 25/25 | 0 | 44.9% / 64% | 173 / 193 MB |
| 50 (2 min) | PASS | 50/50 | 0 | 130.7% / 157% | 315 / 351 MB |
| 75 (2 min) | FAIL | 73/75 | 0 | 109% / 164% | 443 / 653 MB |
Compared with Round 1:
- The ceiling on this provider is unchanged (~50 stable / 75 fails).
- Multi-track WHIP ingest is not free. With zero viewers, ingesting the two publishers (H.264 + HEVC, three layers each, plus two Opus tracks) already costs ~16% of a core and ~108 MB RSS; the idle Round 1 SRT node was ~0% and 46 MB.
- Per-viewer CPU roughly doubled: 0.63 core (Round 1, 50 viewers) ??1.31 core (Round 2).
Caveat. The 50-viewer step shows server-side frame discards and the uplink
reconnected mid-run, so these CPU figures bundle multi-track cost with
real-WAN loss handling. The direction is solid; treat the exact multiplier as
approximate.
5. VPS egress measurements
Before scaling to a larger instance, the actual egress of the provider's
instances was measured ??it did not match the plan label. Representative
measurements on a 4 vCPU / 8 GB LightNode instance (Manila):
| Test | Result |
|---|---|
| A??1 TCP, 16 streams | 103 Mbps |
| A??2 TCP, 16 streams | 106 Mbps |
| A??1 UDP (received) | 99.5 Mbps (97% loss) |
| A??1 and A??2 simultaneously | ~106 Mbps total |
| Upload to a third-party endpoint, 1 flow | ~104 Mbps |
| Upload to a third-party endpoint, 8 concurrent flows | ~103 Mbps aggregate |
Three separately provisioned batches, same-region and cross-region, four cities
(Manila, Taipei, Singapore, Tokyo), never exceeded ~106 Mbps; cross-region paths
fell to ~50 Mbps. There was no tc shaping inside the guest (only the default
mq/fq_codel) and ethtool reported the virtual NIC speed as unknown ??the
cap is on the provider's side, not in the VM.
Consequence: on that network a node tops out at ~39 concurrent 2.5 Mbps
viewers, i.e. network-bound well before its CPU ceiling. On budget VPS,
"1 Gbps" often denotes port speed, a shared uplink, or a burst credit, not
sustained per-instance egress.
6. Provider comparison ??LightNode vs AWS Lightsail
| Dimension | LightNode (as tested) | AWS Lightsail (as tested) |
|---|---|---|
| Region | Manila (also Taipei / Singapore / Tokyo) | Singapore |
| Node sizes used | 2 vCPU / 4 GB (Rounds 1??); 4 vCPU / 8 GB (egress probe) | 2 vCPU / 2 GB |
| Advertised egress | 1 Gbps | 1 Gbps (plan) |
| Measured private-network egress | not measured | ~4.6??.0 Gbps |
| Measured public egress (third-party endpoint) | ~50??06 Mbps | ~2.4 Gbps |
| Max completed concurrent WHEP viewers | 50 (2 vCPU / 4 GB; 75 failed ??network-induced) | 250 (2 vCPU / 2 GB; 300 = out-of-memory) |
| First binding constraint | Provider network (~100 Mbps cap) | Node memory (2 GB) |
| Fitness as a high-fan-out live edge | Poor | Suitable (scale RAM) |
7. AWS Lightsail capacity run
Same method, re-run on AWS Lightsail (Singapore) on delivering nodes. The node
was 2 vCPU / 2 GB and also runs the production ppmmx, so the benchmark
instance shared the box on a separate port set (webrtc 18888, ice 18188/udp,
srt 17890/udp, api 19996, metrics 19998) ??production was never stopped.
| Step | Verdict | Connected | Disconnects | CPU avg / peak | RSS avg / peak | MemAvailable min |
|---|---|---|---|---|---|---|
| 12, 25, 50, 75, 100 | PASS | all | 0 | ?? | ?? | ?? |
| 150 (2 min) | PASS | 150/150 | 0 | 101% / 138% | 640 / 713 MB | 723 MB |
| 200 (2 min) | PASS | 200/200 | 0 | 128% / 174% | 828 / 940 MB | 511 MB |
| 250 (2 min) | PASS | 250/250 | 0 | 154% / 202% | 1014 / 1163 MB | 289 MB |
| 300 (2 min) | host lost | (clients connected) | n/a | 163% / 204% | 1224 / 1514 MB | 33 MB |
- 12 ??250 completed cleanly; every client connected and closed with zero disconnects. Nothing failed at 75, unlike the LightNode run.
- At 250 the box is at its edge: CPU peak 202% (2-core saturated), RSS ~1.16 GB, 289 MB RAM remaining on the 2 GB box.
- At 300 the host ran out of memory. Clients did connect (150 per generator), but available memory fell to ~33 MB, CPU peaked at ~204%, and SSH stopped responding; the node had to be rebooted. 300 is not a passing capacity point ??it is an out-of-memory failure.
- The first binding resource is memory, then CPU ??not the network (~2.4 Gbps, barely used).
8. Findings and capacity guidance
A 2 vCPU / 2 GB
ppmmxstandalone node completed 250 concurrent 2.5 Mbps
WHEP viewers with zero disconnects on a network that delivered ~2.4 Gbps, and
failed at 300 on memory. The earlier "50 stable / 75 fails" result was caused
mainly by the first provider's ~100 Mbps egress cap, not by the node's CPU.
Guidance:
- On a network that genuinely delivers ?? Gbps, budget around 250 concurrent 2.5 Mbps viewers for a 2 vCPU / 2 GB node, and treat memory as the binding resource.
- Budget roughly 4?? MB of RSS per reader and keep headroom.
- Measure the VPS egress and per-reader memory before sizing from a plan label. "1 Gbps" is not a reliable predictor of sustained per-instance egress.
- Multi-track WHIP (HEVC + H.264) ingest costs materially more than single-track SRT at the same viewer count; plan ingest capacity accordingly.
(Rounds 1?? remain a valid relative comparison of SRT vs. multi-track WHIP
ingest cost; only their absolute viewer ceiling was distorted by the first
provider's egress cap.)
9. Limitations and caveats
- Ingest realism differs by round. Round 1 was a single, video-only H.264 track over SRT. Round 2 used the real multi-track WHIP path but from a consumer uplink with real loss, so its CPU figures mix multi-track cost with loss handling. Neither round used a loss-free multi-track source.
- Same-region clients. Load generators were co-located with the node, so this measures server capacity; real internet paths add loss/jitter that reduce effective capacity.
- Small instances only. 2 vCPU / 4 GB (Rounds 1??) and 2 vCPU / 2 GB (Lightsail). No dedicated 4 vCPU / 8 GB run was completed.
-
The Lightsail node was shared with production. The benchmark instance ran
alongside the live production
ppmmxon the same 2 GB box (separate ports, production untouched). This makes the memory ceiling worse than a dedicated node, so the 250-viewer result is a conservative floor. - Short holds. Only the 12-viewer Round 1 step was a 10-minute soak; all other steps were 2-minute holds. Long-term stability (thermals, multi-hour GC, memory creep) is not established.
- Isolated mode. Publish whitelist and playback auth were disabled (no control plane) so the media plane could be tested in isolation; production runs with them enabled.
- Load-generator headroom was not instrumented. Each Lightsail generator backed up to 150 sessions with no observed bottleneck, but its own CPU was not sampled.
10. Reproducibility
All tooling is scripted:
# server (A)
sudo ./mmx-host-setup.sh --binary ./mmx-linux-amd64 --config ./standalone.test.yml
# each load generator (B1, B2)
./loadgen-host-setup.sh --binary ./whep-loadgen-linux-amd64
# drive the ladder from B1
MMX_HOST=<A_IP> MMX_SSH=root@<A_IP> ./loadgen-fleet.sh \
--hosts "local,root@<B2_IP>" \
--loadgen ~/whep-loadgen-linux-amd64 \
--bitrate 2500k --out ./results
The runner splits each step across hosts, samples the server's CPU/RSS and
/metrics throughout, writes a CSV per step, and stops at the first step that
cannot establish every session. standalone.test.yml disables the admin
listener and opens /metrics + the control API to the load generators only; the
Lightsail run used an equivalent config on a separate port set so it could
coexist with the production node.
Per-step CSVs and logs for every run in this report are archived alongside it.






Top comments (0)