DEV Community

greg Tham
greg Tham

Posted on Originally published at docs.pp-cdn.org

EP01: The interactive-video dilemma

Episode 1 of 20 in the **PPCDN Low-Latency Live Streaming Course* — building a self-hosted, low-latency live-streaming CDN from protocols to global delivery. Originally published on the PPCDN docs site; see the full course index for all episodes.*

0. Goals of this episode

By the end, the viewer should be able to answer two questions:

  1. Is my interactive live-streaming / interactive-game project also stuck in the latency, legacy-protocol and single-point/sovereignty traps?
  2. If my pipeline still uses RTMP or HTTP-FLV, what happens on a weak network, and why is that not an exaggeration?

This episode only lays out problems and verifiable facts — no cure; the cure is
EP02–EP20.


1. Opening hook (script notes)

If you've built interactive live streaming — real-person video, online auctions,
co-streamed games, or games with barrage feedback — you may have hit this: the host
has already moved to the next step, but the viewer's picture is still on the previous one.

That's usually not just "the network happened to hiccup". It's the combined result of
capture, encoding, transport, jitter resistance and decode/render. Today we break these
problems into several real dilemmas and separate protocol mechanics, configuration and
measured data.


2. Three real dilemmas

2.1 Dilemma 1: the latency paradox — mature protocols are the opposite of real-time business needs

HLS/DASH is the most mature way to deliver internet video. Based on "segment then
download", it inherently carries seconds of latency. That barely matters for one-way
VOD, but it is a structural conflict for highly interactive scenarios:

Scenario Why latency-sensitive Consequence of traditional segmentation
Live video / real-time trading The interaction window is measured in seconds; the picture must match the business state Users see a stale picture and can't participate
Online auctions / flash sales Bid order decides fairness; the picture can't lag the business state Ordering confusion, disputes
Interactive classrooms Teacher-student Q&A and exercises need instant feedback Interaction degrades into one-way playback
Social co-streaming Conversation depends on natural turn-taking Talking over each other, latency stacking — unusable

Do real-time protocols, usually UDP-based (WebRTC/SRT), fix it automatically? No.
End-to-end latency is the combined result of capture, encoding, network queueing and
propagation, loss recovery, receiver jitter buffer, decode and render. Some waits are
bounded by parameters, some vary dynamically with network and implementation; you can't
just add a few defaults together and call it a measurement, and you certainly can't treat
one buffer parameter as a strict mathematical upper bound for an entire protocol.

The project's confirmed, same-caliber measurements are below. They are PPCDN case data,
not a guarantee of WebRTC, SRT, or any codec in all environments:

Scenario Measured end-to-end latency Notes
P2P direct ~70ms Real link, media bypasses Origin/Edge forwarding
WHIP ingest (via Edge) ~108ms Measured on the current actual pull path
SRT ingest (via Edge) ~386ms Measured on the current actual pull path

The WHIP-vs-SRT gap mainly comes from the ingest protocol itself — SRT's receive window
(TSBPD) inherently queues longer than WebRTC ingest does. That's a protocol difference,
not a codec difference
: HEVC currently doesn't even participate in P2P direct connect, it
only goes through Edge delivery, so the encoding format isn't the variable driving these
numbers. P2P skipping a server hop usually helps latency, but the result still depends on
the network path, NAT traversal outcome, encoder and player. A protocol name only tells
you which mechanisms are available; whether it meets a real-time goal must be measured with
unified instrumentation, endpoints and clock caliber.

2.2 Dilemma 2: a protocol being retired — RTMP/HTTP-FLV on weak networks

2.2.1 An industry migration that already happened

  • Adobe formally ended Flash Player support on 2020-12-31, and mainstream browsers then removed the Flash runtime. RTMP (Real-Time Messaging Protocol) was originally designed for Flash Player playback in the browser — that "native browser playback" path no longer exists.
  • What people call "RTMP live" in a browser today doesn't actually use RTMP for playback: it's re-wrapped as HTTP-FLV (same TCP transport, but the browser still needs a JavaScript library like flv.js to de-mux client-side), or transcoded to HLS/DASH (segment download buys playback compatibility at the cost of real-time — see Dilemma 1).
  • WHIP (WebRTC-HTTP ingestion) and WHEP (WebRTC-HTTP egress) matured during IETF standardization and have been adopted by mainstream publishers including OBS Studio and several real-time services as the next-generation ingest/playback protocols — precisely to replace RTMP at the "ingest" position.
  • This isn't one vendor's preference but the joint direction of browser capability, codec standards and protocol standardization over recent years: the playback-side RTMP/Flash combo is gone, and ingest-side RTMP is being replaced by WHIP/SRT.

2.2.2 Root cause: why it's a disaster on weak networks, not hyperbole

RTMP and HTTP-FLV are both built on TCP. TCP delivers an "ordered, complete" byte stream:
if any packet is lost, later data that already arrived must wait for its retransmission
before reaching the application — this is Head-of-Line Blocking, a design property of
TCP itself, not a defect of some implementation.

On a weak network (with some loss rate), this cascades:

  1. A packet is lost → retransmit, waiting at least one more round-trip (RTT);
  2. During retransmission, all later-arrived audio/video is blocked in the buffer and can't play;
  3. Sustained loss or repeated retransmits → blocking can accumulate from tens of ms to seconds or more, and has no theoretical upper bound;
  4. The player buffer fills or drains → stutter, catch-up, and in the worst case a reconnect.

This differs from common WebRTC/SRT trade-offs: an application can bound how long it waits
for retransmission based on a packet's timeliness and drop data once it's stale, avoiding
endless accumulation for the sake of completeness. SRT does this with ARQ, TSBPD and
stale-packet drop; WebRTC does it with RTP/RTCP feedback, congestion control, jitter buffer
and decoder frame-drop policy. Neither offers a strict mathematical latency bound
independent of network, configuration and implementation: congestion, queueing, outages,
reconnects or implementation policy can still push latency well up.

In one sentence: TCP byte streams prioritize reliable, ordered delivery, so loss causes
head-of-line blocking; a real-time media stack can drop stale data by timeliness, which
usually makes latency easier to control — but it is not an unconditional upper bound.

2.2.3 Protocol comparison

Dimension RTMP / HTTP-FLV SRT WebRTC (WHIP / WHEP)
Transport TCP UDP + ARQ Usually UDP + RTP/RTCP; can fall back to TURN/TCP/TLS
Loss handling TCP must deliver reliably in order; possible head-of-line blocking ARQ requests retransmit; TSBPD delivers by time; configurable stale-drop can discard late data Whether NACK/FEC is used depends on negotiation/implementation; the receiver can drop stale frames
Weak-network latency Queueing/retransmit can accumulate to seconds or kill the stream Engineering bounds via latency etc., but no unconditional math bound Jitter buffer and congestion control adjust dynamically; no unconditional math bound
Native browser playback No longer supported (relies on the discontinued Flash Player) Not supported (usually needs protocol conversion server-side) Native, standard Web APIs
Current role Legacy protocol exiting the playback path The modern choice for weak uplinks The mainstream real-time playback protocol

2.2.4 A concrete architecture choice

This isn't abstract — the project's public technical notes say it plainly: ingest supports
both WHIP and SRT (the SRT ingest measurements in §2.1 are exactly that path), playback is
uniformly WHEP, and HTTP-FLV playback endpoints are deliberately not provided. In other
words, "no HTTP-FLV" isn't an omission but a constraint locked in at protocol-selection
time; SRT is not contradictory — it addresses weak-uplink robustness, while playback still
goes only through WHEP. Such constraints are increasingly common in mature real-time
systems and are the concrete form of the "industry migration" discussed here.

2.3 Dilemma 3: the single-point and sovereignty paradox — the easier it is, the more you hand over your lifeline

Beyond protocols, real projects carry two equally important structural risks that don't show
up in latency numbers:

  • Single point of failure: an Origin is usually a single copy — the sole entry for all publishing. When it fails, the result is not "degradation" but a whole-path outage: every viewer loses the picture at once. This isn't theory — the architecture design doc lists "Origin publish disconnect" as a known failure mode, and explicitly not an automatic, seamless failover: the publisher must re-publish to a new Origin before any viewer can recover. This is an objective risk point in many real interactive-live architectures today, and something to weigh carefully later: whether a single point is worth eliminating is an engineering calculation, not the dogma "a single point must be eliminated".
  • Content sovereignty: to cut engineering effort, some teams outsource distribution to a single third party, even treating it as a "backup fallback". It looks like double insurance, but actually hands business continuity to a third party you can't control — their rate limiting, service changes or cross-border compliance shifts can all directly affect your availability, and you have no say in those decisions.

These two are still blank spots in most discussions about interactive-live projects —
few systematically assess whether a single point is worth eliminating, and few discuss how to
weigh content sovereignty (EP03 expands on it). This episode just points them out.


3. Diagrams

3.1 End-to-end latency composition: defaults can't replace measurement

Capture → Encode → Send queue → Network propagation/queueing/loss recovery
        → Receiver reorder & jitter buffer → Decode → Render

These stages overlap or vary dynamically; configuration values can't be mechanically added.

PPCDN same-caliber measurements:
P2P direct ~70ms
WHIP ingest (via Edge) ~108ms, SRT ingest (via Edge) ~386ms
(the gap comes from the ingest protocol's receive window, not the codec —
HEVC currently doesn't participate in P2P direct connect)
Enter fullscreen mode Exit fullscreen mode

Diagram points (narration cues): unify the measurement endpoints and clock first, then
look at per-stage instrumentation. Network, encoder and buffer policy can each become the
bottleneck; you can't infer end-to-end latency from a single RTP packet count, an SRT
latency parameter, or a player's target buffer value.

3.2 The protocol triangle: latency, weak-network robustness, ecosystem compatibility

                    Lowest latency
                    ╱        ╲
                   ╱ WebRTC    ╲
                  ╱ (WHIP/WHEP) ╲
                 ╱________________╲
   Best weak-net robustness ──── Best ecosystem compatibility
     (SRT, trading retransmit     (RTMP once unified the world
      window for weak-net           via the Flash Player;
      availability)                 after Flash ended, this
                                    corner collapsed)
Enter fullscreen mode Exit fullscreen mode

Diagram points (narration cues): RTMP's historical edge was exactly the "ecosystem
compatibility" corner — almost every browser could play it via Flash. After Flash ended,
that corner no longer holds for RTMP, and it never had an upper bound on the
"weak-network robustness" corner either (see the head-of-line analysis in §2.2.2). That's
the graphical explanation of "RTMP/HTTP-FLV is being retired": it wasn't beaten by a
competitor; its one advantage collapsed on its own.

3.3 Weak-network loss handling: head-of-line blocking vs dropping by timeliness

[RTMP / HTTP-FLV (TCP)]
frame1  frame2  frame3(lost)  frame4  frame5 ...
                   │
                   ▼
     TCP demands "ordered + complete" delivery
                   │
                   ▼
     frame4, frame5 must queue until frame3 is retransmitted
                   │
                   ▼
     Retransmit waits ≥ 1 RTT; on a weak network, repeats stack to seconds, no theoretical bound
                   │
                   ▼
     Player buffer fills → stutter / catch-up / reconnect

[SRT / WebRTC (real-time media policy, typically UDP)]
frame1  frame2  frame3(lost)  frame4  frame5 ...
                   │
                   ▼
     Try retransmit, FEC, or wait for reorder per protocol/implementation policy
                   │
                   ▼
     Once stale, recovery can be abandoned to avoid further backlog
                   │
                   ▼
     Viewer experience: usually trades localized quality loss for more controllable latency
Enter fullscreen mode Exit fullscreen mode

3.4 A real project's architecture today: single point and third-party dependency

                 [Viewers / Internet]
                          │  egress
        ┌─────────────────┼─────────────────┐
        │                 │                 │
   [Edge 1]          [Edge 2]   ...    [Edge N]
        └─────────────────┼─────────────────┘
                internal forwarding
                          │
                   [Origin]   ← sole entry, single point, failure = whole-path outage
                          │
                    ingest
                          │
              [publish site · long online]

     [Origin] ──backup route──> [third-party live service]  ← continuity depends on a third party's decisions
Enter fullscreen mode Exit fullscreen mode

Diagram points (narration cues): this is the distribution architecture of one real
interactive-live project today, not a hypothetical. Both risks of Dilemma 3 map onto it —
the Origin is the sole entry (single-point risk), and the backup route hangs off a
third-party cloud (sovereignty risk). This is why the course starts from "dilemmas" rather
than jumping straight to solutions.


4. Wrap-up and next episode (script notes)

The dilemmas seen here reduce to one line: a traditional live path optimized for one-way
viewing cannot be assumed to meet real-time interaction goals
. Dynamic or configured
buffering, TCP head-of-line blocking, architecture single points and third-party
dependencies all need scenario-specific measurement and trade-offs — not conclusions from a
protocol name alone.

Next episode, we start taking apart how PPCDN addresses these architecturally — trying P2P
first and falling back to Edge in sequence on failure, controlling weak-network backlog by
media timeliness, and keeping the critical path in your own hands.

Top comments (0)