<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: greg Tham</title>
    <description>The latest articles on DEV Community by greg Tham (@greg_tham_9527).</description>
    <link>https://dev.to/greg_tham_9527</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4167978%2F35714fcb-d1df-46bd-accb-fd9808772b74.jpg</url>
      <title>DEV Community: greg Tham</title>
      <link>https://dev.to/greg_tham_9527</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/greg_tham_9527"/>
    <language>en</language>
    <item>
      <title>ppmmx Standalone Stress-Test Report — The Concurrency Where CPU Hits 80% on a 1 vCPU Node</title>
      <dc:creator>greg Tham</dc:creator>
      <pubDate>Thu, 08 Oct 2026 09:40:00 +0000</pubDate>
      <link>https://dev.to/greg_tham_9527/ppmmx-standalone-stress-test-report-the-concurrency-where-cpu-hits-80-on-a-1-vcpu-node-4l2h</link>
      <guid>https://dev.to/greg_tham_9527/ppmmx-standalone-stress-test-report-the-concurrency-where-cpu-hits-80-on-a-1-vcpu-node-4l2h</guid>
      <description>&lt;h1&gt;
  
  
  ppmmx Standalone Stress-Test Report
&lt;/h1&gt;

&lt;p&gt;&lt;strong&gt;1 vCPU / 2GB self-hosted node: the concurrency at which CPU usage hits 80%&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Node under test&lt;/td&gt;
&lt;td&gt;Self-hosted ppmmx standalone node, LightNode, &lt;strong&gt;1 vCPU / 2GB&lt;/strong&gt;, Ubuntu 24.04&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Deployment&lt;/td&gt;
&lt;td&gt;Official public &lt;code&gt;deploy.sh&lt;/code&gt; (&lt;code&gt;ppmmxDocker&lt;/code&gt;), &lt;code&gt;docker compose up -d --build&lt;/code&gt;, default bridge network + port mapping&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Software under test&lt;/td&gt;
&lt;td&gt;ppmmx &lt;code&gt;v1.19.1rc064&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Load generator&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;whep-loadgen&lt;/code&gt; (pion/webrtc, RTP read-and-discard)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ingest&lt;/td&gt;
&lt;td&gt;Local &lt;code&gt;ffmpeg&lt;/code&gt; (&lt;code&gt;testsrc2&lt;/code&gt; + &lt;code&gt;libx264&lt;/code&gt;, 1280×720@30, ~2.5 Mbps, video-only) pushed over the public internet via &lt;strong&gt;SRT&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Date&lt;/td&gt;
&lt;td&gt;2026-10-08&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Summary
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;This 1 vCPU / 2GB node stayed stable through 18 concurrent viewers (CPU ≤ 53%). The moment the 19th viewer joined, CPU climbed from ~40% to 88-103% within about 15 seconds and stayed there, and the server started dropping frames for multiple sessions (&lt;code&gt;reader is too slow&lt;/code&gt;). The concurrency at which CPU crosses 80% is 19.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This isn't a smooth curve — it's a cliff. There's almost no transition zone between 18 and 19 viewers. The reason is straightforward: this is a single-core machine with no second core to absorb GC pauses, network polling, and bursty processing. Once the one core gets close to saturated, SRT/WebRTC processing can't keep up in real time, buffers start backing up (the SRT ingest's RTT jumped from ~4ms to 100-500ms during the 19-viewer run), which drives CPU even higher — a small congestion-collapse feedback loop rather than linear degradation.&lt;/p&gt;

&lt;p&gt;Along the way, testing also surfaced and fixed two real bugs affecting every customer who self-hosts via &lt;code&gt;licenseCode&lt;/code&gt;-only deployment (see "Findings along the way" below); both fixes are already live in production.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Method
&lt;/h2&gt;

&lt;p&gt;Structurally identical to the first round of an earlier 2 vCPU stress test, just on a 1 vCPU box with a finer-grained ladder:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;10 -&amp;gt; 15 -&amp;gt; 17 -&amp;gt; 18 -&amp;gt; 19 -&amp;gt; 20
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each step ramps up to the target number of WHEP readers and holds for 45-50 seconds (covering the full ramp plus a steady-state window), sampling the container's CPU% via &lt;code&gt;docker stats --no-stream&lt;/code&gt; every 3 seconds (single-core box, so 100% = one full core saturated) along with memory, while watching server logs for &lt;code&gt;reader is too slow&lt;/code&gt; warnings or RTT anomalies. Test stream: &lt;code&gt;ffmpeg testsrc2 + libx264&lt;/code&gt;, 1280×720@30, ~2.5 Mbps over SRT — matching the first round of the earlier report (no audio, no multi-track).&lt;/p&gt;

&lt;p&gt;Measured RTT between the load generator and the node under test was ~4ms, close to same-datacenter quality, so this measures the &lt;strong&gt;server's own&lt;/strong&gt; capacity rather than real-world internet-path behavior.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Results
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Concurrency&lt;/th&gt;
&lt;th&gt;Verdict&lt;/th&gt;
&lt;th&gt;Connected&lt;/th&gt;
&lt;th&gt;CPU (steady-state mean / peak)&lt;/th&gt;
&lt;th&gt;Server-side anomalies&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Stable&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;10/10&lt;/td&gt;
&lt;td&gt;26% / 30%&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;15&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Mostly stable&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;15/15&lt;/td&gt;
&lt;td&gt;~45% (with two ~95-98% transient spikes, likely a periodic background task rather than sustained load)&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;17&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Stable&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;17/17&lt;/td&gt;
&lt;td&gt;47% / 64%&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;18&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Stable (last clean step)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;18/18&lt;/td&gt;
&lt;td&gt;45% / 53%&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;19&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Crosses 80%, overload begins&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;19/19&lt;/td&gt;
&lt;td&gt;Climbs from ~40% to 88-103% within ~15s of full ramp, then stays high&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;reader is too slow, discarding N frames&lt;/code&gt; (multiple sessions)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Saturated&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;20/20&lt;/td&gt;
&lt;td&gt;94-104% (high from the start of steady state)&lt;/td&gt;
&lt;td&gt;Same as above, plus SRT ingest RTT rising from ~4ms to 100-500ms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;small&gt;CPU is the container-level percentage from &lt;code&gt;docker stats&lt;/code&gt;; 100% = one full core saturated.&lt;/small&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  3. The Key Finding: A Cliff, Not a Gradient
&lt;/h2&gt;

&lt;p&gt;From 10 to 18 viewers, CPU climbs roughly linearly with concurrency (~26% → ~45-53%, about 2-3% per viewer) — consistent in direction with the linear model seen on multi-core machines. But the step from 18 to 19 blows right past that trend:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Extrapolating the 10-18 slope would predict ~48-55% at 19 viewers.&lt;/li&gt;
&lt;li&gt;The actual measurement was 88-103%, and it didn't happen instantly — with 19 viewers already connected and CPU still sitting at ~40%, it took roughly 15 seconds before it started climbing, eventually settling at a high plateau, with frame-drop warnings logged for multiple readers along the way.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This echoes the "CPU rises roughly linearly with concurrency until one step suddenly runs away" pattern seen on a 2 vCPU machine in the earlier report, but &lt;strong&gt;the cliff on a single-core machine arrives much earlier and is much narrower&lt;/strong&gt; — with no second core to fall back on, once the main processing path can't keep up in real time, SRT/WebRTC buffer backlog and retransmit/frame-drop handling consume even more CPU, creating a self-reinforcing congestion spiral rather than a graceful degradation.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Conclusions and Recommendations
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;This 1 vCPU / 2GB node's practical concurrency ceiling is 18 viewers at 2.5 Mbps WHEP&lt;/strong&gt; (peak CPU 53%, leaving headroom); &lt;strong&gt;19 is the point where CPU crosses 80% and frame drops begin.&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Don't estimate a 1 vCPU node's capacity by simply halving a 2 vCPU model like "CPU% ≈ 15 + 0.97% × N" (that would optimistically suggest ~35 viewers) — a single core's behavior near saturation is a non-linear cliff, not half of a linear curve.&lt;/li&gt;
&lt;li&gt;Memory stayed healthy throughout (peak 161.8 MiB, well under the 2GB limit) — in this test, &lt;strong&gt;CPU hit its ceiling first, not memory&lt;/strong&gt;, unlike the earlier report's conclusion that small standalone nodes hit memory first. The difference is simply that the CPU budget dropped from 2 cores to 1, so CPU became the binding constraint instead.&lt;/li&gt;
&lt;li&gt;If this node needs to serve more viewers, upgrade to 2 vCPU rather than adding RAM.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  5. Findings Along the Way (fixed / logged)
&lt;/h2&gt;

&lt;p&gt;Testing itself got blocked twice by real bugs, both of which got handled:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;A licenseCode-only self-hosted node's publish whitelist never syncs.&lt;/strong&gt; In the node-registration protocol, after ppcenter resolves a node's identity from its license code, it never echoed the resolved node secret back to the node — yet the whitelist-sync endpoint requires that same secret as a credential. The node could never obtain it, so its whitelist stayed permanently empty and &lt;strong&gt;no appId could ever publish&lt;/strong&gt;. Fixed: the registration acknowledgment now includes that resolved value (only for licenseCode-only registrations), and the node caches it lazily after a successful registration. &lt;strong&gt;Already deployed to production&lt;/strong&gt; and verified on the test node (publishing was accepted normally).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;WHEP playback can't connect under the default bridge-network &lt;code&gt;docker-compose.yml&lt;/code&gt; deployment.&lt;/strong&gt; The container defaults to gathering ICE candidates from its network interfaces, but under bridge networking the container only sees its internal Docker bridge IP, not the public one — so WebRTC's ICE candidates point at an address nothing outside the host can ever reach. SRT ingest is unaffected (it's a plain port forward), but playback just times out waiting to connect. This test worked around it by manually adding a public-IP override to the container config, and confirmed that fixes it; &lt;strong&gt;the underlying code hasn't been patched yet&lt;/strong&gt; and is tracked as a separate follow-up.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  6. Limitations and Caveats
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Single-track SRT H.264 only, no audio, no multi-track.&lt;/strong&gt; Matches the earlier report's first round for comparability, but doesn't represent the cost of a real multi-track WHIP ingest path.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Short hold times.&lt;/strong&gt; Each step ran 45-50 seconds; long-term stability (thermal throttling, multi-hour GC behavior) wasn't tested.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Load generator and node under test were near the same datacenter (~4ms RTT).&lt;/strong&gt; This measures server capacity; real internet-path loss/jitter would further reduce effective capacity.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Default limits were raised for testing, then restored.&lt;/strong&gt; The node's default per-path reader cap was temporarily set to unlimited to see the real CPU ceiling; the ICE workaround from section 5 was likewise temporary. Both were restored to factory defaults and verified by restart before this report was finalized.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Single test run, not repeated.&lt;/strong&gt; The 18/19 boundary could shift by ±1 viewer on a repeat run (the 15-viewer step already showed one hard-to-explain transient spike). "18 is safe, 19 crosses the line" is a solid conclusion, but treat the exact boundary as "somewhere between 18 and 19" rather than an absolute number.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>webrtc</category>
      <category>cdn</category>
      <category>performance</category>
      <category>go</category>
    </item>
    <item>
      <title>Live End-to-End Latency: P2P 70ms vs WHIP 108ms vs SRT 386ms</title>
      <dc:creator>greg Tham</dc:creator>
      <pubDate>Wed, 07 Oct 2026 15:42:55 +0000</pubDate>
      <link>https://dev.to/greg_tham_9527/live-end-to-end-latency-p2p-70ms-vs-whip-108ms-vs-srt-386ms-2alh</link>
      <guid>https://dev.to/greg_tham_9527/live-end-to-end-latency-p2p-70ms-vs-whip-108ms-vs-srt-386ms-2alh</guid>
      <description>&lt;h1&gt;
  
  
  Live Latency Performance Test Report
&lt;/h1&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Attribute&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Document type&lt;/td&gt;
&lt;td&gt;Test report (method + result analysis)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Subject&lt;/td&gt;
&lt;td&gt;PPCDN end-to-end live latency: ppobs publish ??ingest ??delivery ??pplayer playback&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Key metric&lt;/td&gt;
&lt;td&gt;End-to-end latency (the player panel's "P2P Delay")&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Result summary&lt;/td&gt;
&lt;td&gt;P2P direct &lt;strong&gt;70ms&lt;/strong&gt;, WHIP ingest &lt;strong&gt;108ms&lt;/strong&gt;, SRT ingest &lt;strong&gt;386ms&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Updated&lt;/td&gt;
&lt;td&gt;2026-09-29&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  0. Summary
&lt;/h2&gt;

&lt;p&gt;After tuning, the measured end-to-end latency of three access/playback modes:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Scenario&lt;/th&gt;
&lt;th&gt;Path&lt;/th&gt;
&lt;th&gt;Measured&lt;/th&gt;
&lt;th&gt;Screenshot&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;P2P direct&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Player ??publisher direct (no Edge)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;70ms&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;?4.1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;WHIP ingest&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;WHIP ingest ??Edge delivery ??playback&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;108ms&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;?4.2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;SRT ingest&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;SRT ingest ??Edge delivery ??playback&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;386ms&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;?4.3&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Conclusion&lt;/strong&gt;: latency &lt;strong&gt;P2P &amp;lt; WHIP &amp;lt; SRT&lt;/strong&gt;. SRT is about &lt;strong&gt;278ms&lt;/strong&gt; higher than WHIP,&lt;br&gt;
mainly from the ingest-side SRT receive window (TSBPD), a fixed delivery delay; WHIP&lt;br&gt;
ingest has no such window and is clearly lower; P2P direct has the shortest path and&lt;br&gt;
the lowest latency.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Goal &amp;amp; scope
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Quantify and compare P2P direct, WHIP ingest and SRT ingest end-to-end latency.&lt;/li&gt;
&lt;li&gt;Verify the latency gains after tuning and document a reproducible, publishable method.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Not tested&lt;/strong&gt;: throughput ceiling, concurrency scale, billing correctness, recording path.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Metric definition and measurement principle
&lt;/h2&gt;

&lt;h3&gt;
  
  
  2.1 End-to-end latency
&lt;/h3&gt;

&lt;p&gt;The one-way delay from the &lt;strong&gt;ppobs capture instant&lt;/strong&gt; (in-bitstream mark) to &lt;strong&gt;pplayer&lt;br&gt;
rendering&lt;/strong&gt;, covering capture, encode, ingest, delivery (Origin??dge or P2P direct),&lt;br&gt;
jitter buffer and decode. The player panel's "P2P Delay" is this value (the label is&lt;br&gt;
shared by the Edge / P2P paths; the actual path is shown as Connection in the panel).&lt;/p&gt;

&lt;h3&gt;
  
  
  2.2 Measurement path
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;ppobs writes an &lt;strong&gt;absolute UTC timestamp&lt;/strong&gt; into the bitstream (H.264/H.265 SEI);
Origin/Edge pass it through hop by hop without decoding or transcoding.&lt;/li&gt;
&lt;li&gt;ppplayer subtracts it from a clock calibrated against ppcenter:
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;  corrected_now   = local_clock + offset
  one_way_delay   = corrected_now - embedded_timestamp
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;Readings add a fixed &lt;strong&gt;+50ms&lt;/strong&gt; compensation for the encode/decode/render overhead
between the mark and rendering that cannot be measured directly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cross-browser measurable&lt;/strong&gt;: to avoid non-Chromium browsers (e.g. iOS Safari, which
lacks WebCodecs Insertable Streams and cannot read the in-bitstream SEI) seeing only
an estimate, Edge nodes parse the ingest stream's SEI and push &lt;code&gt;OBS_TIMESTAMP&lt;/code&gt; to the
player over the ABR control WebSocket, so the measured value is readable on iOS Safari too.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  2.3 Boundaries
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Readings come from the ppplayer panel; full percentiles are aggregated per minute/day
by ppcenter and visible in the console.&lt;/li&gt;
&lt;li&gt;Single machine, single stream, limited window ??it does &lt;em&gt;not&lt;/em&gt; represent a network-wide SLA.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  3. Test configuration
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Item&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Publisher&lt;/td&gt;
&lt;td&gt;ppobs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ingest / connection&lt;/td&gt;
&lt;td&gt;P2P direct / WHIP / SRT (three-way comparison)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Codec&lt;/td&gt;
&lt;td&gt;H.264 (HEVC included in multitrack scenarios)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Player&lt;/td&gt;
&lt;td&gt;pplayer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Playback buffer&lt;/td&gt;
&lt;td&gt;ppplayer default&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  4. Measured results
&lt;/h2&gt;

&lt;h3&gt;
  
  
  4.1 P2P direct ??70ms
&lt;/h3&gt;

&lt;p&gt;The player is &lt;strong&gt;P2P-direct&lt;/strong&gt; to the publisher (no Edge); end-to-end latency &lt;strong&gt;70ms&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4b8zt4fsh98h9ql8n27e.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4b8zt4fsh98h9ql8n27e.jpg" alt="P2P direct measured end-to-end latency 70ms" width="799" height="437"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  4.2 WHIP ingest ??108ms
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;WHIP ingest&lt;/strong&gt;, Edge delivery; end-to-end latency &lt;strong&gt;108ms&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbhj2ksvps076lmrl1b0z.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbhj2ksvps076lmrl1b0z.jpg" alt="WHIP ingest measured end-to-end latency 108ms" width="800" height="556"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  4.3 SRT ingest ??386ms
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;SRT ingest&lt;/strong&gt;, Edge delivery; end-to-end latency &lt;strong&gt;386ms&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpjnsalof5rapibz9n2iw.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpjnsalof5rapibz9n2iw.jpg" alt="SRT ingest measured end-to-end latency 386ms" width="800" height="556"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Analysis
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Scenario&lt;/th&gt;
&lt;th&gt;Measured&lt;/th&gt;
&lt;th&gt;vs WHIP&lt;/th&gt;
&lt;th&gt;Main difference&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;P2P direct&lt;/td&gt;
&lt;td&gt;70ms&lt;/td&gt;
&lt;td&gt;??8ms&lt;/td&gt;
&lt;td&gt;Shortest path: player directly to publisher, no ingest/cascade&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;WHIP ingest&lt;/td&gt;
&lt;td&gt;108ms&lt;/td&gt;
&lt;td&gt;baseline&lt;/td&gt;
&lt;td&gt;No TSBPD receive window, small ingest overhead&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SRT ingest&lt;/td&gt;
&lt;td&gt;386ms&lt;/td&gt;
&lt;td&gt;+278ms&lt;/td&gt;
&lt;td&gt;Ingest SRT receive window (TSBPD) adds a fixed delivery delay, plus weak-network retransmit buffering&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;P2P direct is lowest&lt;/strong&gt;: it removes the Origin??dge cascade and ingest queue, leaving
only capture/encode + end-to-end network + decode/render.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;WHIP beats SRT&lt;/strong&gt;: WHIP is WebRTC-based and has no fixed TSBPD receive window on
ingest ??consistent with the design intent (SRT's receive window is deliberately
enlarged for weak-network retransmission: a loss-robustness ??latency trade-off).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;SRT is highest&lt;/strong&gt;: the ~278ms overhead closely matches the SRT receive window
(hundreds of ms at its floor); on a lossy uplink, retransmission/receive buffering
pushes it higher. On a clean uplink SRT falls back near its receive-window floor.&lt;/li&gt;
&lt;li&gt;The gaps are stable and reproducible, indicating the difference comes from the
&lt;strong&gt;protocol/link structure&lt;/strong&gt;, not occasional noise.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  6. SLA note
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Target&lt;/th&gt;
&lt;th&gt;Ceiling&lt;/th&gt;
&lt;th&gt;Path&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;P80&lt;/td&gt;
&lt;td&gt;??300ms&lt;/td&gt;
&lt;td&gt;??600ms&lt;/td&gt;
&lt;td&gt;P2P direct&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;P99&lt;/td&gt;
&lt;td&gt;??2s&lt;/td&gt;
&lt;td&gt;??&lt;/td&gt;
&lt;td&gt;P2P direct&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This test's &lt;strong&gt;P2P direct 70ms&lt;/strong&gt; is well within the P80 target (??00ms). Note the SRT&lt;br&gt;
ingest path's latency is governed by the ingest receive window and is not the same&lt;br&gt;
metric as the P2P direct SLA.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. Conclusion &amp;amp; limitations
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Conclusion&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;End-to-end: P2P direct (70ms) &amp;lt; WHIP ingest (108ms) &amp;lt; SRT ingest (386ms).&lt;/li&gt;
&lt;li&gt;For the lowest latency, prefer &lt;strong&gt;P2P direct&lt;/strong&gt;, then &lt;strong&gt;WHIP ingest&lt;/strong&gt;; &lt;strong&gt;SRT ingest&lt;/strong&gt;
trades several times WHIP's latency for weak-uplink retransmission robustness ??the
choice depends on your tolerance for weak networks.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Known limitations&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Single machine, single stream, limited window ??not a network-wide SLA.&lt;/li&gt;
&lt;li&gt;SRT readings depend on uplink loss/retransmission and receive-window adaptation, and vary over time.&lt;/li&gt;
&lt;li&gt;Mobile devices are strongly affected by carrier network and signal strength.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  8. Revision history
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Date&lt;/th&gt;
&lt;th&gt;Revision&lt;/th&gt;
&lt;th&gt;Notes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;2026-09-26&lt;/td&gt;
&lt;td&gt;Draft&lt;/td&gt;
&lt;td&gt;Test plan finalized&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2026-09-29&lt;/td&gt;
&lt;td&gt;Results&lt;/td&gt;
&lt;td&gt;Merged end-to-end latency analysis; added 5-scenario measurements&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2026-09-29&lt;/td&gt;
&lt;td&gt;Tuned results&lt;/td&gt;
&lt;td&gt;Updated to a P2P/WHIP/SRT comparison (70/108/386ms) with screenshots&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

</description>
      <category>webrtc</category>
      <category>streaming</category>
      <category>latency</category>
      <category>video</category>
    </item>
    <item>
      <title>ppmmx Standalone Stress-Test Report: Concurrent WHEP Viewer Capacity on 2 vCPU Nodes</title>
      <dc:creator>greg Tham</dc:creator>
      <pubDate>Wed, 07 Oct 2026 14:45:36 +0000</pubDate>
      <link>https://dev.to/greg_tham_9527/ppmmx-standalone-stress-test-report-concurrent-whep-viewer-capacity-on-2-vcpu-nodes-4df5</link>
      <guid>https://dev.to/greg_tham_9527/ppmmx-standalone-stress-test-report-concurrent-whep-viewer-capacity-on-2-vcpu-nodes-4df5</guid>
      <description>&lt;h1&gt;
  
  
  ppmmx Standalone Stress-Test Report
&lt;/h1&gt;

&lt;p&gt;&lt;strong&gt;Concurrent WHEP viewer capacity on 2 vCPU nodes&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Report&lt;/td&gt;
&lt;td&gt;ppmmx standalone load-capacity and VPS-egress test&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Version&lt;/td&gt;
&lt;td&gt;v4 (final)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Date&lt;/td&gt;
&lt;td&gt;2026-10-07&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Software under test&lt;/td&gt;
&lt;td&gt;ppmmx &lt;code&gt;v1.19.1rc064&lt;/code&gt;, &lt;code&gt;NODE_ROLE_STANDALONE&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Load generator&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;whep-loadgen&lt;/code&gt; (pion/webrtc, RTP read-and-discard)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Providers compared&lt;/td&gt;
&lt;td&gt;LightNode (Manila / Taipei / Singapore / Tokyo); AWS Lightsail (Singapore)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Abstract
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;ppmmx&lt;/code&gt; in &lt;code&gt;standalone&lt;/code&gt; mode was stress-tested for concurrent WHEP viewer&lt;br&gt;
capacity on small cloud VPS instances, over two ingest paths (single-track SRT,&lt;br&gt;
then multi-track WHIP with simultaneous HEVC + H.264), and across two providers.&lt;br&gt;
The central findings are:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The first provider tested (LightNode) never delivered its advertised
1 Gbps.&lt;/strong&gt; Across three separately provisioned batches and four cities,
measured egress was &lt;strong&gt;~50??06 Mbps&lt;/strong&gt; (per-instance, per-direction, both TCP
and UDP, and to a third-party endpoint). A node on that network saturates the
&lt;em&gt;link&lt;/em&gt; at roughly 39 concurrent 2.5 Mbps viewers ??long before its CPU.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;On a provider that delivered real bandwidth (AWS Lightsail: ~2.4 Gbps
public, ~5 Gbps private), a &lt;em&gt;smaller&lt;/em&gt; 2 vCPU / 2 GB node completed 250
concurrent WHEP viewers with zero disconnects.&lt;/strong&gt; The 300-viewer step ran the
node out of &lt;strong&gt;memory&lt;/strong&gt;, not bandwidth and not CPU.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Multi-track WHIP ingest costs materially more than single-track SRT&lt;/strong&gt; ??an
ingest-only baseline of ~16% of a core and ~108 MB, and roughly double the
per-viewer CPU at the top stable step.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Memory, not CPU and not the network, is what a small standalone node hits
first.&lt;/strong&gt; Budget from measured egress and measured per-reader memory, not from
the plan label.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Section 6 gives a side-by-side provider comparison.&lt;/p&gt;
&lt;h2&gt;
  
  
  1. Background
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;ppmmx&lt;/code&gt; is a low-latency live-video CDN node. The &lt;code&gt;standalone&lt;/code&gt; role is the&lt;br&gt;
self-contained deployment: a single instance both &lt;strong&gt;ingests&lt;/strong&gt; a publisher&lt;br&gt;
(WHIP/SRT/RTMP) and &lt;strong&gt;serves&lt;/strong&gt; viewers directly over WebRTC (WHEP), with no&lt;br&gt;
forwarding to other mmx nodes and no control-plane dependency.&lt;/p&gt;

&lt;p&gt;The operational question is: &lt;strong&gt;how many concurrent viewers does one box hold,&lt;br&gt;
and what limits it?&lt;/strong&gt; This report answers that for the smallest sensible&lt;br&gt;
instance sizes and two real ingest paths.&lt;/p&gt;
&lt;h2&gt;
  
  
  2. Method
&lt;/h2&gt;
&lt;h3&gt;
  
  
  2.1 Procedure
&lt;/h3&gt;

&lt;p&gt;A stepped concurrency ladder is run against the node:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;12 (10-min soak, Round 1)  -&amp;gt;  25  -&amp;gt;  50  -&amp;gt;  75  -&amp;gt;  100  -&amp;gt;  150  -&amp;gt;  200  -&amp;gt;  250  -&amp;gt;  300
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;At each step, the target number of WHEP readers is opened, held for 2 minutes,&lt;br&gt;
then closed together. A step is recorded &lt;strong&gt;PASS&lt;/strong&gt; only if every session&lt;br&gt;
establishes and closes with no disconnects; the ladder stops at the first step&lt;br&gt;
that fails. Node CPU, RSS, available memory and the node's &lt;code&gt;/metrics&lt;/code&gt; are&lt;br&gt;
sampled throughout.&lt;/p&gt;
&lt;h3&gt;
  
  
  2.2 Tooling and measurement
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Readers:&lt;/strong&gt; &lt;code&gt;whep-loadgen&lt;/code&gt;, a purpose-built Go client that opens N real
&lt;code&gt;pion/webrtc&lt;/code&gt; &lt;code&gt;PeerConnection&lt;/code&gt;s against the WHEP endpoint. Each session
completes full SDP / ICE / DTLS / SRTP and then reads RTP and discards it (no
decoding), so the generator is not measuring its own video pipeline. Sessions
are spread evenly across two load generators.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Server sampling:&lt;/strong&gt; process CPU and RSS every 2 s from
&lt;code&gt;/proc/&amp;lt;pid&amp;gt;/{stat,status}&lt;/code&gt; (CPU as a delta, i.e. instantaneous, not a
lifetime average); available memory and load average from &lt;code&gt;/proc&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Node &lt;code&gt;/metrics&lt;/code&gt;:&lt;/strong&gt; session count, outbound bytes/packets, RTP loss,
discarded frames.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Per-session:&lt;/strong&gt; connect time, first-frame time, disconnects, bytes, packets.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ramp:&lt;/strong&gt; 500 ms between session starts.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;
  
  
  2.3 Test stream
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Round 1:&lt;/strong&gt; one H.264 1280?720@30, ~2.5 Mbps stream, generated with &lt;code&gt;ffmpeg&lt;/code&gt;
(&lt;code&gt;testsrc2&lt;/code&gt; + &lt;code&gt;libx264&lt;/code&gt;) and published over &lt;strong&gt;SRT&lt;/strong&gt;. Video only (no audio).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Round 2:&lt;/strong&gt; the real product ingest path ??&lt;strong&gt;ppobs&lt;/strong&gt; publishing over
&lt;strong&gt;WHIP&lt;/strong&gt; with &lt;strong&gt;simultaneous HEVC + H.264 multi-track&lt;/strong&gt; (three H.264 simulcast
layers + three HEVC layers per stream, each session also carrying Opus),
pushed from a genuine consumer broadband uplink.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;
  
  
  3. Round 1 ??SRT ingest, single H.264 track
&lt;/h2&gt;

&lt;p&gt;Environment: 2 vCPU / 4 GB node and two 2 vCPU / 4 GB load generators, all on&lt;br&gt;
LightNode (Manila), same datacenter.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Step&lt;/th&gt;
&lt;th&gt;Sessions&lt;/th&gt;
&lt;th&gt;Verdict&lt;/th&gt;
&lt;th&gt;Connected&lt;/th&gt;
&lt;th&gt;Disconnects&lt;/th&gt;
&lt;th&gt;mmx CPU avg / peak&lt;/th&gt;
&lt;th&gt;mmx RSS avg / peak&lt;/th&gt;
&lt;th&gt;Egress&lt;/th&gt;
&lt;th&gt;Loss&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;baseline&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;??&lt;/td&gt;
&lt;td&gt;??&lt;/td&gt;
&lt;td&gt;??&lt;/td&gt;
&lt;td&gt;0.5% / 2.5%&lt;/td&gt;
&lt;td&gt;46 / 48 MB&lt;/td&gt;
&lt;td&gt;??&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;12 (10 min)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;PASS&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;12/12&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;26.4% / 30%&lt;/td&gt;
&lt;td&gt;114 / 116 MB&lt;/td&gt;
&lt;td&gt;~31 Mbps&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;25 (2 min)&lt;/td&gt;
&lt;td&gt;25&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;PASS&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;25/25&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;39.5% / 45%&lt;/td&gt;
&lt;td&gt;167 / 175 MB&lt;/td&gt;
&lt;td&gt;~64 Mbps&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;50 (2 min)&lt;/td&gt;
&lt;td&gt;50&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;PASS&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;50/50&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;63.1% / 79.5%&lt;/td&gt;
&lt;td&gt;274 / 294 MB&lt;/td&gt;
&lt;td&gt;~128 Mbps&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;75 (2 min)&lt;/td&gt;
&lt;td&gt;75&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;FAIL&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;71/75&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;134% / 177%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;618 / &lt;strong&gt;1068 MB&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;(overloaded)&lt;/td&gt;
&lt;td&gt;0*&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;small&gt;CPU is per-process where 100% = one full core. *RTP loss stays 0, but at&lt;br&gt;
75 the server logs &lt;code&gt;reader is too slow, discarding 27 frames&lt;/code&gt; ??the degradation&lt;br&gt;
appears as dropped frames, not on-wire loss.&lt;/small&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7xmzmizvw0ui7lfcatsd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7xmzmizvw0ui7lfcatsd.png" alt="CPU vs concurrency" width="800" height="480"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;From 12 ??50 viewers, CPU rises almost linearly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;mmx CPU% ??15%  +  0.97%  ?  concurrent_viewers     (1 core = 100%)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;about &lt;strong&gt;1% of a core per 2.5 Mbps viewer&lt;/strong&gt; plus ~15% fixed overhead. The last&lt;br&gt;
stable step (50) used 0.63 of a core. At 75 the model breaks: measured CPU&lt;br&gt;
reached &lt;strong&gt;134% average / 177% peak&lt;/strong&gt; (a two-core box at ~88%, run-queue load&lt;br&gt;
~2.0); the B2 shard could not establish its last sessions (WHEP POST timeouts&lt;br&gt;
and ICE timeouts), first-frame latency for that shard rose from ~0.3 s to&lt;br&gt;
&lt;strong&gt;3.1 s average / 9.6 s worst&lt;/strong&gt;, and the server began discarding frames for slow&lt;br&gt;
readers.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4s4pmxsq7s2846jjcj12.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4s4pmxsq7s2846jjcj12.png" alt="Egress vs concurrency" width="800" height="480"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Measured egress is ~&lt;strong&gt;2.55 Mbps per viewer&lt;/strong&gt;. At the last stable step that is&lt;br&gt;
~128 Mbps of a 1 Gbps port (~13%). &lt;em&gt;(Caveat: this is the server's own send&lt;br&gt;
count and may include bytes the provider later dropped; see ?5.)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frg36u5ntl5odqolo33k0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frg36u5ntl5odqolo33k0.png" alt="Memory vs concurrency" width="800" height="480"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fti6hwwij7l2hux1b7oju.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fti6hwwij7l2hux1b7oju.png" alt="CPU and RSS over time" width="800" height="391"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;RSS grows roughly linearly with readers (~14??6 MB/reader) and is &lt;strong&gt;reclaimed&lt;/strong&gt;&lt;br&gt;
after sessions close: the Go runtime returned memory over ~2?? minutes back&lt;br&gt;
toward a ~170 MB baseline (from a 46 MB cold start), &lt;code&gt;mem_alloc&lt;/code&gt; fell from&lt;br&gt;
~486 MB to ~77 MB, and &lt;code&gt;goroutines&lt;/code&gt; stayed flat at 18 ??no leak.&lt;/p&gt;
&lt;h2&gt;
  
  
  4. Round 2 ??ppobs WHIP ingest, HEVC + H.264 multi-track
&lt;/h2&gt;

&lt;p&gt;Same node and ladder, but the real ingest path: ppobs over WHIP, three H.264 +&lt;br&gt;
three HEVC simulcast layers and Opus per session, from a consumer uplink.&lt;/p&gt;

&lt;p&gt;The uplink is a real internet path, not the LAN-like SRT link of Round 1. The&lt;br&gt;
&lt;strong&gt;base layer arrived at 0% loss&lt;/strong&gt;; the upper simulcast layers lost ~20% ??normal&lt;br&gt;
behaviour for a home uplink once the stream exceeds the available upstream. Loss&lt;br&gt;
on that link is a normal network condition, not a test artefact, but it means&lt;br&gt;
the node's CPU below includes retransmission/discard work.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Step&lt;/th&gt;
&lt;th&gt;Verdict&lt;/th&gt;
&lt;th&gt;Connected&lt;/th&gt;
&lt;th&gt;Disconnects&lt;/th&gt;
&lt;th&gt;mmx CPU avg / peak&lt;/th&gt;
&lt;th&gt;mmx RSS avg / peak&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;ingest only (0 viewers)&lt;/td&gt;
&lt;td&gt;??&lt;/td&gt;
&lt;td&gt;??&lt;/td&gt;
&lt;td&gt;??&lt;/td&gt;
&lt;td&gt;15.8%&lt;/td&gt;
&lt;td&gt;108 MB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;12 (3 min)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;PASS&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;12/12&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;32% / 39.5%&lt;/td&gt;
&lt;td&gt;118 / 127 MB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;25 (2 min)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;PASS&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;25/25&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;44.9% / 64%&lt;/td&gt;
&lt;td&gt;173 / 193 MB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;50 (2 min)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;PASS&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;50/50&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;130.7% / 157%&lt;/td&gt;
&lt;td&gt;315 / 351 MB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;75 (2 min)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;FAIL&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;73/75&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;109% / 164%&lt;/td&gt;
&lt;td&gt;443 / 653 MB&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8xnnwc8joynr3ujbmmwb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8xnnwc8joynr3ujbmmwb.png" alt="CPU: round 1 vs round 2" width="800" height="480"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Compared with Round 1:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;The ceiling on this provider is unchanged (~50 stable / 75 fails).&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Multi-track WHIP ingest is not free.&lt;/strong&gt; With zero viewers, ingesting the two
publishers (H.264 + HEVC, three layers each, plus two Opus tracks) already
costs &lt;strong&gt;~16% of a core and ~108 MB RSS&lt;/strong&gt;; the idle Round 1 SRT node was ~0% and
46 MB.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Per-viewer CPU roughly doubled&lt;/strong&gt;: 0.63 core (Round 1, 50 viewers) ??&lt;strong&gt;1.31
core&lt;/strong&gt; (Round 2).&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Caveat.&lt;/strong&gt; The 50-viewer step shows server-side frame discards and the uplink&lt;br&gt;
reconnected mid-run, so these CPU figures bundle multi-track cost with&lt;br&gt;
real-WAN loss handling. The &lt;em&gt;direction&lt;/em&gt; is solid; treat the exact multiplier as&lt;br&gt;
approximate.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2&gt;
  
  
  5. VPS egress measurements
&lt;/h2&gt;

&lt;p&gt;Before scaling to a larger instance, the actual egress of the provider's&lt;br&gt;
instances was measured ??it did not match the plan label. Representative&lt;br&gt;
measurements on a 4 vCPU / 8 GB LightNode instance (Manila):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Test&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A??1 TCP, 16 streams&lt;/td&gt;
&lt;td&gt;103 Mbps&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A??2 TCP, 16 streams&lt;/td&gt;
&lt;td&gt;106 Mbps&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A??1 UDP (received)&lt;/td&gt;
&lt;td&gt;99.5 Mbps (97% loss)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A??1 and A??2 &lt;strong&gt;simultaneously&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;~106 Mbps total&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Upload to a third-party endpoint, 1 flow&lt;/td&gt;
&lt;td&gt;~104 Mbps&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Upload to a third-party endpoint, &lt;strong&gt;8 concurrent flows&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~103 Mbps aggregate&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Three separately provisioned batches, same-region and cross-region, four cities&lt;br&gt;
(Manila, Taipei, Singapore, Tokyo), never exceeded ~106 Mbps; cross-region paths&lt;br&gt;
fell to ~50 Mbps. There was no &lt;code&gt;tc&lt;/code&gt; shaping inside the guest (only the default&lt;br&gt;
&lt;code&gt;mq&lt;/code&gt;/&lt;code&gt;fq_codel&lt;/code&gt;) and &lt;code&gt;ethtool&lt;/code&gt; reported the virtual NIC speed as unknown ??the&lt;br&gt;
cap is on the provider's side, not in the VM.&lt;/p&gt;

&lt;p&gt;Consequence: on that network a node tops out at ~&lt;strong&gt;39 concurrent 2.5 Mbps&lt;br&gt;
viewers&lt;/strong&gt;, i.e. &lt;strong&gt;network-bound well before its CPU ceiling&lt;/strong&gt;. On budget VPS,&lt;br&gt;
"1 Gbps" often denotes port speed, a shared uplink, or a burst credit, not&lt;br&gt;
sustained per-instance egress.&lt;/p&gt;
&lt;h2&gt;
  
  
  6. Provider comparison ??LightNode vs AWS Lightsail
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;LightNode (as tested)&lt;/th&gt;
&lt;th&gt;AWS Lightsail (as tested)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Region&lt;/td&gt;
&lt;td&gt;Manila (also Taipei / Singapore / Tokyo)&lt;/td&gt;
&lt;td&gt;Singapore&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Node sizes used&lt;/td&gt;
&lt;td&gt;2 vCPU / 4 GB (Rounds 1??); 4 vCPU / 8 GB (egress probe)&lt;/td&gt;
&lt;td&gt;2 vCPU / 2 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Advertised egress&lt;/td&gt;
&lt;td&gt;1 Gbps&lt;/td&gt;
&lt;td&gt;1 Gbps (plan)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Measured &lt;strong&gt;private-network&lt;/strong&gt; egress&lt;/td&gt;
&lt;td&gt;not measured&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~4.6??.0 Gbps&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Measured &lt;strong&gt;public&lt;/strong&gt; egress (third-party endpoint)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~50??06 Mbps&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~2.4 Gbps&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Max &lt;strong&gt;completed&lt;/strong&gt; concurrent WHEP viewers&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;50&lt;/strong&gt; (2 vCPU / 4 GB; 75 failed ??network-induced)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;250&lt;/strong&gt; (2 vCPU / 2 GB; 300 = out-of-memory)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;First binding constraint&lt;/td&gt;
&lt;td&gt;Provider network (~100 Mbps cap)&lt;/td&gt;
&lt;td&gt;Node memory (2 GB)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fitness as a high-fan-out live edge&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Poor&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Suitable&lt;/strong&gt; (scale RAM)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;h2&gt;
  
  
  7. AWS Lightsail capacity run
&lt;/h2&gt;

&lt;p&gt;Same method, re-run on AWS Lightsail (Singapore) on delivering nodes. The node&lt;br&gt;
was 2 vCPU / 2 GB and also runs the production &lt;code&gt;ppmmx&lt;/code&gt;, so the benchmark&lt;br&gt;
instance shared the box on a &lt;strong&gt;separate port set&lt;/strong&gt; (webrtc 18888, ice 18188/udp,&lt;br&gt;
srt 17890/udp, api 19996, metrics 19998) ??production was never stopped.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Step&lt;/th&gt;
&lt;th&gt;Verdict&lt;/th&gt;
&lt;th&gt;Connected&lt;/th&gt;
&lt;th&gt;Disconnects&lt;/th&gt;
&lt;th&gt;CPU avg / peak&lt;/th&gt;
&lt;th&gt;RSS avg / peak&lt;/th&gt;
&lt;th&gt;MemAvailable min&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;12, 25, 50, 75, 100&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;PASS&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;all&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;??&lt;/td&gt;
&lt;td&gt;??&lt;/td&gt;
&lt;td&gt;??&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;150 (2 min)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;PASS&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;150/150&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;101% / 138%&lt;/td&gt;
&lt;td&gt;640 / 713 MB&lt;/td&gt;
&lt;td&gt;723 MB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;200 (2 min)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;PASS&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;200/200&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;128% / 174%&lt;/td&gt;
&lt;td&gt;828 / 940 MB&lt;/td&gt;
&lt;td&gt;511 MB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;250 (2 min)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;PASS&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;250/250&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;154% / 202%&lt;/td&gt;
&lt;td&gt;1014 / 1163 MB&lt;/td&gt;
&lt;td&gt;289 MB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;300 (2 min)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;host lost&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;(clients connected)&lt;/td&gt;
&lt;td&gt;n/a&lt;/td&gt;
&lt;td&gt;163% / &lt;strong&gt;204%&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;1224 / &lt;strong&gt;1514 MB&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;33 MB&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flyo0997ed78jwrg3rwge.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flyo0997ed78jwrg3rwge.png" alt="AWS Lightsail capacity" width="799" height="305"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;12 ??250 completed cleanly&lt;/strong&gt;; every client connected and closed with zero
disconnects. Nothing failed at 75, unlike the LightNode run.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;At 250 the box is at its edge&lt;/strong&gt;: CPU peak 202% (2-core saturated), RSS
~1.16 GB, &lt;strong&gt;289 MB&lt;/strong&gt; RAM remaining on the 2 GB box.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;At 300 the host ran out of memory.&lt;/strong&gt; Clients did connect (150 per
generator), but available memory fell to &lt;strong&gt;~33 MB&lt;/strong&gt;, CPU peaked at ~204%, and
SSH stopped responding; the node had to be rebooted. &lt;strong&gt;300 is not a passing
capacity point ??it is an out-of-memory failure.&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;The first binding resource is &lt;strong&gt;memory&lt;/strong&gt;, then CPU ??not the network (~2.4 Gbps,
barely used).&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;
  
  
  8. Findings and capacity guidance
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A 2 vCPU / 2 GB &lt;code&gt;ppmmx&lt;/code&gt; standalone node completed 250 concurrent 2.5 Mbps&lt;br&gt;
WHEP viewers with zero disconnects on a network that delivered ~2.4 Gbps, and&lt;br&gt;
failed at 300 on memory. The earlier "50 stable / 75 fails" result was caused&lt;br&gt;
mainly by the first provider's ~100 Mbps egress cap, not by the node's CPU.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Guidance:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;On a network that genuinely delivers ?? Gbps, budget around &lt;strong&gt;250 concurrent
2.5 Mbps viewers&lt;/strong&gt; for a 2 vCPU / 2 GB node, and treat &lt;strong&gt;memory&lt;/strong&gt; as the
binding resource.&lt;/li&gt;
&lt;li&gt;Budget roughly &lt;strong&gt;4?? MB of RSS per reader&lt;/strong&gt; and keep headroom.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Measure the VPS egress and per-reader memory before sizing from a plan
label.&lt;/strong&gt; "1 Gbps" is not a reliable predictor of sustained per-instance
egress.&lt;/li&gt;
&lt;li&gt;Multi-track WHIP (HEVC + H.264) ingest costs materially more than single-track
SRT at the same viewer count; plan ingest capacity accordingly.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;(Rounds 1?? remain a valid relative comparison of SRT vs. multi-track WHIP&lt;br&gt;
ingest cost; only their absolute viewer ceiling was distorted by the first&lt;br&gt;
provider's egress cap.)&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  9. Limitations and caveats
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Ingest realism differs by round.&lt;/strong&gt; Round 1 was a single, video-only H.264
track over SRT. Round 2 used the real multi-track WHIP path but from a
consumer uplink with real loss, so its CPU figures mix multi-track cost with
loss handling. Neither round used a loss-free multi-track source.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Same-region clients.&lt;/strong&gt; Load generators were co-located with the node, so
this measures &lt;em&gt;server&lt;/em&gt; capacity; real internet paths add loss/jitter that
reduce effective capacity.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Small instances only.&lt;/strong&gt; 2 vCPU / 4 GB (Rounds 1??) and 2 vCPU / 2 GB
(Lightsail). No dedicated 4 vCPU / 8 GB run was completed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The Lightsail node was shared with production.&lt;/strong&gt; The benchmark instance ran
alongside the live production &lt;code&gt;ppmmx&lt;/code&gt; on the same 2 GB box (separate ports,
production untouched). This makes the memory ceiling &lt;em&gt;worse&lt;/em&gt; than a dedicated
node, so the 250-viewer result is a conservative floor.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Short holds.&lt;/strong&gt; Only the 12-viewer Round 1 step was a 10-minute soak; all
other steps were 2-minute holds. Long-term stability (thermals, multi-hour GC,
memory creep) is not established.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Isolated mode.&lt;/strong&gt; Publish whitelist and playback auth were disabled (no
control plane) so the media plane could be tested in isolation; production
runs with them enabled.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Load-generator headroom was not instrumented.&lt;/strong&gt; Each Lightsail generator
backed up to 150 sessions with no observed bottleneck, but its own CPU was not
sampled.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;
  
  
  10. Reproducibility
&lt;/h2&gt;

&lt;p&gt;All tooling is scripted:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# server (A)&lt;/span&gt;
&lt;span class="nb"&gt;sudo&lt;/span&gt; ./mmx-host-setup.sh &lt;span class="nt"&gt;--binary&lt;/span&gt; ./mmx-linux-amd64 &lt;span class="nt"&gt;--config&lt;/span&gt; ./standalone.test.yml

&lt;span class="c"&gt;# each load generator (B1, B2)&lt;/span&gt;
./loadgen-host-setup.sh &lt;span class="nt"&gt;--binary&lt;/span&gt; ./whep-loadgen-linux-amd64

&lt;span class="c"&gt;# drive the ladder from B1&lt;/span&gt;
&lt;span class="nv"&gt;MMX_HOST&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&amp;lt;A_IP&amp;gt; &lt;span class="nv"&gt;MMX_SSH&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;root@&amp;lt;A_IP&amp;gt; ./loadgen-fleet.sh &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--hosts&lt;/span&gt; &lt;span class="s2"&gt;"local,root@&amp;lt;B2_IP&amp;gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--loadgen&lt;/span&gt; ~/whep-loadgen-linux-amd64 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--bitrate&lt;/span&gt; 2500k &lt;span class="nt"&gt;--out&lt;/span&gt; ./results
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The runner splits each step across hosts, samples the server's CPU/RSS and&lt;br&gt;
&lt;code&gt;/metrics&lt;/code&gt; throughout, writes a CSV per step, and stops at the first step that&lt;br&gt;
cannot establish every session. &lt;code&gt;standalone.test.yml&lt;/code&gt; disables the admin&lt;br&gt;
listener and opens &lt;code&gt;/metrics&lt;/code&gt; + the control API to the load generators only; the&lt;br&gt;
Lightsail run used an equivalent config on a separate port set so it could&lt;br&gt;
coexist with the production node.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Per-step CSVs and logs for every run in this report are archived alongside it.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>webrtc</category>
      <category>cdn</category>
      <category>performance</category>
      <category>go</category>
    </item>
  </channel>
</rss>
