DEV Community

greg Tham
greg Tham

Posted on

ppmmx Standalone Stress-Test Report — The Concurrency Where CPU Hits 80% on a 1 vCPU Node

ppmmx Standalone Stress-Test Report

1 vCPU / 2GB self-hosted node: the concurrency at which CPU usage hits 80%

Node under test Self-hosted ppmmx standalone node, LightNode, 1 vCPU / 2GB, Ubuntu 24.04
Deployment Official public deploy.sh (ppmmxDocker), docker compose up -d --build, default bridge network + port mapping
Software under test ppmmx v1.19.1rc064
Load generator whep-loadgen (pion/webrtc, RTP read-and-discard)
Ingest Local ffmpeg (testsrc2 + libx264, 1280×720@30, ~2.5 Mbps, video-only) pushed over the public internet via SRT
Date 2026-10-08

Summary

This 1 vCPU / 2GB node stayed stable through 18 concurrent viewers (CPU ≤ 53%). The moment the 19th viewer joined, CPU climbed from ~40% to 88-103% within about 15 seconds and stayed there, and the server started dropping frames for multiple sessions (reader is too slow). The concurrency at which CPU crosses 80% is 19.

This isn't a smooth curve — it's a cliff. There's almost no transition zone between 18 and 19 viewers. The reason is straightforward: this is a single-core machine with no second core to absorb GC pauses, network polling, and bursty processing. Once the one core gets close to saturated, SRT/WebRTC processing can't keep up in real time, buffers start backing up (the SRT ingest's RTT jumped from ~4ms to 100-500ms during the 19-viewer run), which drives CPU even higher — a small congestion-collapse feedback loop rather than linear degradation.

Along the way, testing also surfaced and fixed two real bugs affecting every customer who self-hosts via licenseCode-only deployment (see "Findings along the way" below); both fixes are already live in production.

1. Method

Structurally identical to the first round of an earlier 2 vCPU stress test, just on a 1 vCPU box with a finer-grained ladder:

10 -> 15 -> 17 -> 18 -> 19 -> 20
Enter fullscreen mode Exit fullscreen mode

Each step ramps up to the target number of WHEP readers and holds for 45-50 seconds (covering the full ramp plus a steady-state window), sampling the container's CPU% via docker stats --no-stream every 3 seconds (single-core box, so 100% = one full core saturated) along with memory, while watching server logs for reader is too slow warnings or RTT anomalies. Test stream: ffmpeg testsrc2 + libx264, 1280×720@30, ~2.5 Mbps over SRT — matching the first round of the earlier report (no audio, no multi-track).

Measured RTT between the load generator and the node under test was ~4ms, close to same-datacenter quality, so this measures the server's own capacity rather than real-world internet-path behavior.

2. Results

Concurrency Verdict Connected CPU (steady-state mean / peak) Server-side anomalies
10 Stable 10/10 26% / 30% None
15 Mostly stable 15/15 ~45% (with two ~95-98% transient spikes, likely a periodic background task rather than sustained load) None
17 Stable 17/17 47% / 64% None
18 Stable (last clean step) 18/18 45% / 53% None
19 Crosses 80%, overload begins 19/19 Climbs from ~40% to 88-103% within ~15s of full ramp, then stays high reader is too slow, discarding N frames (multiple sessions)
20 Saturated 20/20 94-104% (high from the start of steady state) Same as above, plus SRT ingest RTT rising from ~4ms to 100-500ms

CPU is the container-level percentage from docker stats; 100% = one full core saturated.

3. The Key Finding: A Cliff, Not a Gradient

From 10 to 18 viewers, CPU climbs roughly linearly with concurrency (~26% → ~45-53%, about 2-3% per viewer) — consistent in direction with the linear model seen on multi-core machines. But the step from 18 to 19 blows right past that trend:

  • Extrapolating the 10-18 slope would predict ~48-55% at 19 viewers.
  • The actual measurement was 88-103%, and it didn't happen instantly — with 19 viewers already connected and CPU still sitting at ~40%, it took roughly 15 seconds before it started climbing, eventually settling at a high plateau, with frame-drop warnings logged for multiple readers along the way.

This echoes the "CPU rises roughly linearly with concurrency until one step suddenly runs away" pattern seen on a 2 vCPU machine in the earlier report, but the cliff on a single-core machine arrives much earlier and is much narrower — with no second core to fall back on, once the main processing path can't keep up in real time, SRT/WebRTC buffer backlog and retransmit/frame-drop handling consume even more CPU, creating a self-reinforcing congestion spiral rather than a graceful degradation.

4. Conclusions and Recommendations

  • This 1 vCPU / 2GB node's practical concurrency ceiling is 18 viewers at 2.5 Mbps WHEP (peak CPU 53%, leaving headroom); 19 is the point where CPU crosses 80% and frame drops begin.
  • Don't estimate a 1 vCPU node's capacity by simply halving a 2 vCPU model like "CPU% ≈ 15 + 0.97% × N" (that would optimistically suggest ~35 viewers) — a single core's behavior near saturation is a non-linear cliff, not half of a linear curve.
  • Memory stayed healthy throughout (peak 161.8 MiB, well under the 2GB limit) — in this test, CPU hit its ceiling first, not memory, unlike the earlier report's conclusion that small standalone nodes hit memory first. The difference is simply that the CPU budget dropped from 2 cores to 1, so CPU became the binding constraint instead.
  • If this node needs to serve more viewers, upgrade to 2 vCPU rather than adding RAM.

5. Findings Along the Way (fixed / logged)

Testing itself got blocked twice by real bugs, both of which got handled:

  1. A licenseCode-only self-hosted node's publish whitelist never syncs. In the node-registration protocol, after ppcenter resolves a node's identity from its license code, it never echoed the resolved node secret back to the node — yet the whitelist-sync endpoint requires that same secret as a credential. The node could never obtain it, so its whitelist stayed permanently empty and no appId could ever publish. Fixed: the registration acknowledgment now includes that resolved value (only for licenseCode-only registrations), and the node caches it lazily after a successful registration. Already deployed to production and verified on the test node (publishing was accepted normally).
  2. WHEP playback can't connect under the default bridge-network docker-compose.yml deployment. The container defaults to gathering ICE candidates from its network interfaces, but under bridge networking the container only sees its internal Docker bridge IP, not the public one — so WebRTC's ICE candidates point at an address nothing outside the host can ever reach. SRT ingest is unaffected (it's a plain port forward), but playback just times out waiting to connect. This test worked around it by manually adding a public-IP override to the container config, and confirmed that fixes it; the underlying code hasn't been patched yet and is tracked as a separate follow-up.

6. Limitations and Caveats

  • Single-track SRT H.264 only, no audio, no multi-track. Matches the earlier report's first round for comparability, but doesn't represent the cost of a real multi-track WHIP ingest path.
  • Short hold times. Each step ran 45-50 seconds; long-term stability (thermal throttling, multi-hour GC behavior) wasn't tested.
  • Load generator and node under test were near the same datacenter (~4ms RTT). This measures server capacity; real internet-path loss/jitter would further reduce effective capacity.
  • Default limits were raised for testing, then restored. The node's default per-path reader cap was temporarily set to unlimited to see the real CPU ceiling; the ICE workaround from section 5 was likewise temporary. Both were restored to factory defaults and verified by restart before this report was finalized.
  • Single test run, not repeated. The 18/19 boundary could shift by ±1 viewer on a repeat run (the 15-viewer step already showed one hard-to-explain transient spike). "18 is safe, 19 crosses the line" is a solid conclusion, but treat the exact boundary as "somewhere between 18 and 19" rather than an absolute number.

Top comments (0)