DEV Community

greg Tham
greg Tham

Posted on

ppmmx Standalone Stress-Test Report: Concurrent WHEP Viewer Capacity on 2 vCPU Nodes

ppmmx Standalone Stress-Test Report

Concurrent WHEP viewer capacity on 2 vCPU nodes

Report ppmmx standalone load-capacity and VPS-egress test
Version v4 (final)
Date 2026-10-07
Software under test ppmmx v1.19.1rc064, NODE_ROLE_STANDALONE
Load generator whep-loadgen (pion/webrtc, RTP read-and-discard)
Providers compared LightNode (Manila / Taipei / Singapore / Tokyo); AWS Lightsail (Singapore)

Abstract

ppmmx in standalone mode was stress-tested for concurrent WHEP viewer
capacity on small cloud VPS instances, over two ingest paths (single-track SRT,
then multi-track WHIP with simultaneous HEVC + H.264), and across two providers.
The central findings are:

  1. The first provider tested (LightNode) never delivered its advertised 1 Gbps. Across three separately provisioned batches and four cities, measured egress was ~50??06 Mbps (per-instance, per-direction, both TCP and UDP, and to a third-party endpoint). A node on that network saturates the link at roughly 39 concurrent 2.5 Mbps viewers ??long before its CPU.
  2. On a provider that delivered real bandwidth (AWS Lightsail: ~2.4 Gbps public, ~5 Gbps private), a smaller 2 vCPU / 2 GB node completed 250 concurrent WHEP viewers with zero disconnects. The 300-viewer step ran the node out of memory, not bandwidth and not CPU.
  3. Multi-track WHIP ingest costs materially more than single-track SRT ??an ingest-only baseline of ~16% of a core and ~108 MB, and roughly double the per-viewer CPU at the top stable step.
  4. Memory, not CPU and not the network, is what a small standalone node hits first. Budget from measured egress and measured per-reader memory, not from the plan label.

Section 6 gives a side-by-side provider comparison.

1. Background

ppmmx is a low-latency live-video CDN node. The standalone role is the
self-contained deployment: a single instance both ingests a publisher
(WHIP/SRT/RTMP) and serves viewers directly over WebRTC (WHEP), with no
forwarding to other mmx nodes and no control-plane dependency.

The operational question is: how many concurrent viewers does one box hold,
and what limits it?
This report answers that for the smallest sensible
instance sizes and two real ingest paths.

2. Method

2.1 Procedure

A stepped concurrency ladder is run against the node:

12 (10-min soak, Round 1)  ->  25  ->  50  ->  75  ->  100  ->  150  ->  200  ->  250  ->  300
Enter fullscreen mode Exit fullscreen mode

At each step, the target number of WHEP readers is opened, held for 2 minutes,
then closed together. A step is recorded PASS only if every session
establishes and closes with no disconnects; the ladder stops at the first step
that fails. Node CPU, RSS, available memory and the node's /metrics are
sampled throughout.

2.2 Tooling and measurement

  • Readers: whep-loadgen, a purpose-built Go client that opens N real pion/webrtc PeerConnections against the WHEP endpoint. Each session completes full SDP / ICE / DTLS / SRTP and then reads RTP and discards it (no decoding), so the generator is not measuring its own video pipeline. Sessions are spread evenly across two load generators.
  • Server sampling: process CPU and RSS every 2 s from /proc/<pid>/{stat,status} (CPU as a delta, i.e. instantaneous, not a lifetime average); available memory and load average from /proc.
  • Node /metrics: session count, outbound bytes/packets, RTP loss, discarded frames.
  • Per-session: connect time, first-frame time, disconnects, bytes, packets.
  • Ramp: 500 ms between session starts.

2.3 Test stream

  • Round 1: one H.264 1280?720@30, ~2.5 Mbps stream, generated with ffmpeg (testsrc2 + libx264) and published over SRT. Video only (no audio).
  • Round 2: the real product ingest path ??ppobs publishing over WHIP with simultaneous HEVC + H.264 multi-track (three H.264 simulcast layers + three HEVC layers per stream, each session also carrying Opus), pushed from a genuine consumer broadband uplink.

3. Round 1 ??SRT ingest, single H.264 track

Environment: 2 vCPU / 4 GB node and two 2 vCPU / 4 GB load generators, all on
LightNode (Manila), same datacenter.

Step Sessions Verdict Connected Disconnects mmx CPU avg / peak mmx RSS avg / peak Egress Loss
baseline 0 ?? ?? ?? 0.5% / 2.5% 46 / 48 MB ?? 0
12 (10 min) 12 PASS 12/12 0 26.4% / 30% 114 / 116 MB ~31 Mbps 0
25 (2 min) 25 PASS 25/25 0 39.5% / 45% 167 / 175 MB ~64 Mbps 0
50 (2 min) 50 PASS 50/50 0 63.1% / 79.5% 274 / 294 MB ~128 Mbps 0
75 (2 min) 75 FAIL 71/75 5 134% / 177% 618 / 1068 MB (overloaded) 0*

CPU is per-process where 100% = one full core. *RTP loss stays 0, but at
75 the server logs reader is too slow, discarding 27 frames ??the degradation
appears as dropped frames, not on-wire loss.

CPU vs concurrency

From 12 ??50 viewers, CPU rises almost linearly:

mmx CPU% ??15%  +  0.97%  ?  concurrent_viewers     (1 core = 100%)
Enter fullscreen mode Exit fullscreen mode

about 1% of a core per 2.5 Mbps viewer plus ~15% fixed overhead. The last
stable step (50) used 0.63 of a core. At 75 the model breaks: measured CPU
reached 134% average / 177% peak (a two-core box at ~88%, run-queue load
~2.0); the B2 shard could not establish its last sessions (WHEP POST timeouts
and ICE timeouts), first-frame latency for that shard rose from ~0.3 s to
3.1 s average / 9.6 s worst, and the server began discarding frames for slow
readers.

Egress vs concurrency

Measured egress is ~2.55 Mbps per viewer. At the last stable step that is
~128 Mbps of a 1 Gbps port (~13%). (Caveat: this is the server's own send
count and may include bytes the provider later dropped; see ?5.)

Memory vs concurrency

CPU and RSS over time

RSS grows roughly linearly with readers (~14??6 MB/reader) and is reclaimed
after sessions close: the Go runtime returned memory over ~2?? minutes back
toward a ~170 MB baseline (from a 46 MB cold start), mem_alloc fell from
~486 MB to ~77 MB, and goroutines stayed flat at 18 ??no leak.

4. Round 2 ??ppobs WHIP ingest, HEVC + H.264 multi-track

Same node and ladder, but the real ingest path: ppobs over WHIP, three H.264 +
three HEVC simulcast layers and Opus per session, from a consumer uplink.

The uplink is a real internet path, not the LAN-like SRT link of Round 1. The
base layer arrived at 0% loss; the upper simulcast layers lost ~20% ??normal
behaviour for a home uplink once the stream exceeds the available upstream. Loss
on that link is a normal network condition, not a test artefact, but it means
the node's CPU below includes retransmission/discard work.

Step Verdict Connected Disconnects mmx CPU avg / peak mmx RSS avg / peak
ingest only (0 viewers) ?? ?? ?? 15.8% 108 MB
12 (3 min) PASS 12/12 0 32% / 39.5% 118 / 127 MB
25 (2 min) PASS 25/25 0 44.9% / 64% 173 / 193 MB
50 (2 min) PASS 50/50 0 130.7% / 157% 315 / 351 MB
75 (2 min) FAIL 73/75 0 109% / 164% 443 / 653 MB

CPU: round 1 vs round 2

Compared with Round 1:

  • The ceiling on this provider is unchanged (~50 stable / 75 fails).
  • Multi-track WHIP ingest is not free. With zero viewers, ingesting the two publishers (H.264 + HEVC, three layers each, plus two Opus tracks) already costs ~16% of a core and ~108 MB RSS; the idle Round 1 SRT node was ~0% and 46 MB.
  • Per-viewer CPU roughly doubled: 0.63 core (Round 1, 50 viewers) ??1.31 core (Round 2).

Caveat. The 50-viewer step shows server-side frame discards and the uplink
reconnected mid-run, so these CPU figures bundle multi-track cost with
real-WAN loss handling. The direction is solid; treat the exact multiplier as
approximate.

5. VPS egress measurements

Before scaling to a larger instance, the actual egress of the provider's
instances was measured ??it did not match the plan label. Representative
measurements on a 4 vCPU / 8 GB LightNode instance (Manila):

Test Result
A??1 TCP, 16 streams 103 Mbps
A??2 TCP, 16 streams 106 Mbps
A??1 UDP (received) 99.5 Mbps (97% loss)
A??1 and A??2 simultaneously ~106 Mbps total
Upload to a third-party endpoint, 1 flow ~104 Mbps
Upload to a third-party endpoint, 8 concurrent flows ~103 Mbps aggregate

Three separately provisioned batches, same-region and cross-region, four cities
(Manila, Taipei, Singapore, Tokyo), never exceeded ~106 Mbps; cross-region paths
fell to ~50 Mbps. There was no tc shaping inside the guest (only the default
mq/fq_codel) and ethtool reported the virtual NIC speed as unknown ??the
cap is on the provider's side, not in the VM.

Consequence: on that network a node tops out at ~39 concurrent 2.5 Mbps
viewers
, i.e. network-bound well before its CPU ceiling. On budget VPS,
"1 Gbps" often denotes port speed, a shared uplink, or a burst credit, not
sustained per-instance egress.

6. Provider comparison ??LightNode vs AWS Lightsail

Dimension LightNode (as tested) AWS Lightsail (as tested)
Region Manila (also Taipei / Singapore / Tokyo) Singapore
Node sizes used 2 vCPU / 4 GB (Rounds 1??); 4 vCPU / 8 GB (egress probe) 2 vCPU / 2 GB
Advertised egress 1 Gbps 1 Gbps (plan)
Measured private-network egress not measured ~4.6??.0 Gbps
Measured public egress (third-party endpoint) ~50??06 Mbps ~2.4 Gbps
Max completed concurrent WHEP viewers 50 (2 vCPU / 4 GB; 75 failed ??network-induced) 250 (2 vCPU / 2 GB; 300 = out-of-memory)
First binding constraint Provider network (~100 Mbps cap) Node memory (2 GB)
Fitness as a high-fan-out live edge Poor Suitable (scale RAM)

7. AWS Lightsail capacity run

Same method, re-run on AWS Lightsail (Singapore) on delivering nodes. The node
was 2 vCPU / 2 GB and also runs the production ppmmx, so the benchmark
instance shared the box on a separate port set (webrtc 18888, ice 18188/udp,
srt 17890/udp, api 19996, metrics 19998) ??production was never stopped.

Step Verdict Connected Disconnects CPU avg / peak RSS avg / peak MemAvailable min
12, 25, 50, 75, 100 PASS all 0 ?? ?? ??
150 (2 min) PASS 150/150 0 101% / 138% 640 / 713 MB 723 MB
200 (2 min) PASS 200/200 0 128% / 174% 828 / 940 MB 511 MB
250 (2 min) PASS 250/250 0 154% / 202% 1014 / 1163 MB 289 MB
300 (2 min) host lost (clients connected) n/a 163% / 204% 1224 / 1514 MB 33 MB

AWS Lightsail capacity

  • 12 ??250 completed cleanly; every client connected and closed with zero disconnects. Nothing failed at 75, unlike the LightNode run.
  • At 250 the box is at its edge: CPU peak 202% (2-core saturated), RSS ~1.16 GB, 289 MB RAM remaining on the 2 GB box.
  • At 300 the host ran out of memory. Clients did connect (150 per generator), but available memory fell to ~33 MB, CPU peaked at ~204%, and SSH stopped responding; the node had to be rebooted. 300 is not a passing capacity point ??it is an out-of-memory failure.
  • The first binding resource is memory, then CPU ??not the network (~2.4 Gbps, barely used).

8. Findings and capacity guidance

A 2 vCPU / 2 GB ppmmx standalone node completed 250 concurrent 2.5 Mbps
WHEP viewers with zero disconnects on a network that delivered ~2.4 Gbps, and
failed at 300 on memory. The earlier "50 stable / 75 fails" result was caused
mainly by the first provider's ~100 Mbps egress cap, not by the node's CPU.

Guidance:

  • On a network that genuinely delivers ?? Gbps, budget around 250 concurrent 2.5 Mbps viewers for a 2 vCPU / 2 GB node, and treat memory as the binding resource.
  • Budget roughly 4?? MB of RSS per reader and keep headroom.
  • Measure the VPS egress and per-reader memory before sizing from a plan label. "1 Gbps" is not a reliable predictor of sustained per-instance egress.
  • Multi-track WHIP (HEVC + H.264) ingest costs materially more than single-track SRT at the same viewer count; plan ingest capacity accordingly.

(Rounds 1?? remain a valid relative comparison of SRT vs. multi-track WHIP
ingest cost; only their absolute viewer ceiling was distorted by the first
provider's egress cap.)

9. Limitations and caveats

  • Ingest realism differs by round. Round 1 was a single, video-only H.264 track over SRT. Round 2 used the real multi-track WHIP path but from a consumer uplink with real loss, so its CPU figures mix multi-track cost with loss handling. Neither round used a loss-free multi-track source.
  • Same-region clients. Load generators were co-located with the node, so this measures server capacity; real internet paths add loss/jitter that reduce effective capacity.
  • Small instances only. 2 vCPU / 4 GB (Rounds 1??) and 2 vCPU / 2 GB (Lightsail). No dedicated 4 vCPU / 8 GB run was completed.
  • The Lightsail node was shared with production. The benchmark instance ran alongside the live production ppmmx on the same 2 GB box (separate ports, production untouched). This makes the memory ceiling worse than a dedicated node, so the 250-viewer result is a conservative floor.
  • Short holds. Only the 12-viewer Round 1 step was a 10-minute soak; all other steps were 2-minute holds. Long-term stability (thermals, multi-hour GC, memory creep) is not established.
  • Isolated mode. Publish whitelist and playback auth were disabled (no control plane) so the media plane could be tested in isolation; production runs with them enabled.
  • Load-generator headroom was not instrumented. Each Lightsail generator backed up to 150 sessions with no observed bottleneck, but its own CPU was not sampled.

10. Reproducibility

All tooling is scripted:

# server (A)
sudo ./mmx-host-setup.sh --binary ./mmx-linux-amd64 --config ./standalone.test.yml

# each load generator (B1, B2)
./loadgen-host-setup.sh --binary ./whep-loadgen-linux-amd64

# drive the ladder from B1
MMX_HOST=<A_IP> MMX_SSH=root@<A_IP> ./loadgen-fleet.sh \
  --hosts "local,root@<B2_IP>" \
  --loadgen ~/whep-loadgen-linux-amd64 \
  --bitrate 2500k --out ./results
Enter fullscreen mode Exit fullscreen mode

The runner splits each step across hosts, samples the server's CPU/RSS and
/metrics throughout, writes a CSV per step, and stops at the first step that
cannot establish every session. standalone.test.yml disables the admin
listener and opens /metrics + the control API to the load generators only; the
Lightsail run used an equivalent config on a separate port set so it could
coexist with the production node.

Per-step CSVs and logs for every run in this report are archived alongside it.

Top comments (0)