DEV Community

greg Tham
greg Tham

Posted on Originally published at docs.pp-cdn.org

EP02: The self-hosted CDN answer — PPCDN architecture overview

Episode 2 of 20 in the **PPCDN Low-Latency Live Streaming Course* — building a self-hosted, low-latency live-streaming CDN from protocols to global delivery. Originally published on the PPCDN docs site; see the full course index for all episodes.*

Recap EP01 "The interactive-video dilemma": the latency paradox, the retiring RTMP/HTTP-FLV, the single-point and sovereignty paradox
Next EP03 "Content sovereignty: control it yourself, the ultimate moat"

0. Goals of this episode

By the end, the viewer should be able to:

  1. Draw PPCDN's topology on a whiteboard: publisher, Origin, Edge, viewers, and where the control plane handling scheduling/signaling sits.
  2. Explain the engineering response PPCDN gives to each of EP01's three dilemmas — the low-latency path, the real-time protocol choice, and a controllable distributed topology.
  3. Explain the three concrete engineering reasons "self-hosting can beat big-cloud services on cost-effectiveness", rather than the empty claim "self-hosting is cheaper".

1. Opening hook (script notes)

Last episode we broke down three dilemmas: dynamically-accumulating latency across the
pipeline, RTMP/HTTP-FLV exiting the browser playback ecosystem, and single points and
third-party dependencies. This episode stops posing problems — we draw PPCDN's overall
architecture and see how it responds to each.


2. Architecture overview: the whole picture in one diagram

2.1 Three tiers + one control plane

PPCDN's media path has just three tiers, plus a control plane that never touches media data:

  • Publisher: the in-house ppobs (deeply customized from OBS Studio), or any standard WHIP/SRT publisher.
  • Origin: the source that receives the publish stream — the sole entry for media into the system.
  • Edge (multiple): pulls from Origin and serves viewers; multiple viewers of the same stream on one Edge share a single origin-pull link.
  • ppcenter (control plane): scheduling, auth, node management, P2P signaling, billing metadata and health reports. It carries no media and proxies no media packets.

Keeping the control plane and media plane fully separate is the most fundamental design
principle in the whole architecture; nearly every later "high availability" conclusion rests on it.

2.2 Why separate the control plane from the media plane

If scheduling/auth "decision logic" shared a path with audio/video, a control-plane wobble
would wobble viewers too. PPCDN keeps media packets out of ppcenter: the control plane
mainly participates in connection setup, auth and scheduling; after setup, media takes an
independent path, and a control-plane failure does not interrupt in-progress viewing.
This
fundamentally shrinks the failure domain and lets the media plane run independently and
steadily.

Another easily missed premise is clock synchronization. Short-lived credentials'
iat/exp/txTime, cross-device latency instrumentation and fault detection all depend on
time: nodes sync their clocks via NTP and monitor skew, and auth allows only a clear, finite
clock-skew tolerance. The whole network's clocks stay self-consistent with ppcenter as the
common reference — subtracting timestamps across nodes is only meaningful within that frame.


3. How each of the three dilemmas is answered

3.1 Answering Dilemma 1 (latency paradox): codec-aware routing to the low-latency path

EP01 explained that the Edge path adds a server hop, and latency is further shaped by
encoding, network, recovery and buffer policy. PPCDN's architectural answer is to route by
the player's decode capability, opening an additional, usually-shorter P2P direct path for
H.264 clients
:

  • Routing by decode capability at playback time: HEVC-capable clients go straight to the Edge HEVC stream — HEVC is usually more bandwidth-efficient than H.264 at equal quality, and "~30% less" is a commonly-cited interval midpoint in the industry and inside this project, not a number this project has itself measured and validated yet (EP20 digs into the measurement method and its limits), though this path does avoid the first-frame uncertainty of P2P setup. Clients without HEVC first try a P2P direct connect for H.264.
  • On the P2P route, ppcenter first runs the NAT eligibility check and atomically allocates a signaling slot (at most 3 direct connections per publisher by platform default — defaultMaxP2PSessions=3, adjustable by a superadmin up to 50); the two sides then relay SDP/ICE through ppcenter to establish the direct connection — ppcenter only relays signaling and never parses media.
  • Once P2P connects, media flows straight from publisher to viewer, bypassing Origin/Edge and consuming zero edge egress — the lowest-latency path.
  • On any failed check, timeout or setup failure, it cleanly falls back to the Edge H.264 stream, with no error or black screen.
  • P2P only uses the publisher's measured spare uplink; on detected uplink congestion it is sacrificed immediately, so the main publish's quality and bandwidth budget are never affected.

EP13 covers NAT traversal and eligibility details; for now, remember one line: route by
decode capability first; H.264 connects directly when it can, and cleanly falls back to Edge
when it can't
.

3.2 Answering Dilemma 2 (retiring protocols): WHIP/SRT ingest, WHEP playback throughout

EP01 explained that TCP's reliable-ordered byte stream can hit head-of-line blocking on loss.
PPCDN's answer is to prefer a real-time transport stack that can bound recovery by media timeliness:

  • Ingest supports both WHIP and SRT: typical paths use UDP and bound recovery by media timeliness, so a weak network is not dragged down by TCP head-of-line blocking.
  • Delivery and playback uniformly use WHEP; Edge offers only low-latency WebRTC playback and no HTTP-FLV endpoint — a constraint locked at the requirements stage.
  • The weak-network answer isn't "tough out the loss" but proactive degradation: the server steps up/down between simulcast layers based on loss and RTT — first lowering bitrate (without interrupting the publish), and only lowering resolution if bitrate is already at the floor and still short (that step has a brief interruption). EP04 and EP07 cover the transport comparison and weak-network mechanisms in detail.

3.3 Answering Dilemma 3 (single point and sovereignty): shrink the failure domain, keep control in your hands

  • The Edge tier scales horizontally: Edge deploys across multiple nodes, scheduled by node pool and preferring the node with the most free capacity; multiple viewers of the same stream on one Edge share a single origin-pull link, so one overloaded Edge doesn't drag down everything.
  • Edge is autonomous: before each origin pull, Edge dynamically resolves the Origin address and short-lived media credentials from ppcenter (POST /internal/mmx/v1/origin/resolve) and caches the result in memory — when the control center is unavailable, it reuses the unexpired cache to keep reconnecting, so the media plane doesn't wobble with the control plane (docs/ppcdn-development-progress.zh-CN.md §3.5, status "completed").
  • Control plane and media plane are separated: a control-plane failure does not interrupt in-progress viewing; the media path keeps running steadily.
  • Content sovereignty: the whole system is self-hostable with nodes under your control, so the entire publish-to-delivery path is yours; the sovereignty trade-offs and compliance strategy are covered in EP03.

4. Diagrams

4.1 Overall architecture

                         ppcenter (control plane)
              scheduling · auth · P2P signaling · billing metadata · health reports
                    ▲                              ▲
                    │ control / signaling          │ control / signaling
                    │                              │
  ppobs publisher ──WHIP/SRT──▶ Origin ──forward──▶ Edge 1..N ──WHEP──▶ viewer browser
        │                                                                    ▲
        └──── P2P direct: H.264 first, fall back to Edge (HEVC goes straight to Edge) ────┘
Enter fullscreen mode Exit fullscreen mode

Diagram points (narration cues): note the control plane (ppcenter) is drawn above the
path with dashed arrows — it only takes part in the "establish connection" decision, not on
the media path. Media travels the solid path below, and P2P direct is a shortcut that bypasses
Origin/Edge straight from publisher to viewer.

4.2 Playback routing decision

Viewer requests playback
        │
        ▼
pplayer detects whether the browser can hardware-decode HEVC
        │
        ├─ HEVC supported ──▶ go straight to Edge, append /hevc/whep
        │
        └─ No HEVC
                │
                ▼
        Try a P2P direct connect for H.264 first
                │
                ├─ ppcenter eligibility passes + a free signaling slot ──▶ relay SDP/ICE ──▶ use P2P
                │
                └─ Not eligible / timeout / setup fails / slots full ──▶ cleanly fall back to Edge H.264
Enter fullscreen mode Exit fullscreen mode

4.3 Control/media separation: a control-plane failure does not affect the media plane

                    ppcenter (control plane) fails
                            │
                            ▼
              Established media sessions keep transferring steadily
                            │
                            ▼
        After recovery, all endpoints re-register and rebuild leases
Enter fullscreen mode Exit fullscreen mode

Diagram points (narration cues): control plane and media plane are fully separated, so a
control-center failure does not interrupt in-progress viewing; after recovery, all endpoints
re-register and rebuild leases. This is the key design that lets PPCDN shrink the failure
domain.


5. Cost-effectiveness: why self-hosting can beat big-cloud services

The sections above address "can self-hosting achieve low latency and control risk". This one
answers a more practical question: at equal investment, why can self-hosting beat big-cloud
live/RTC services like Tencent Cloud LEB or Huawei Cloud on cost-effectiveness?
Not via
scale discounts, but via three engineering levers that big clouds structurally struggle to copy.

5.1 Custom ppobs: the whole path is optimizable end to end, a lever big clouds don't have

Big-cloud live/RTC services can usually optimize only "their own side" — origin and edge
delivery; the publisher is either a customer's generic SDK or standard OBS, so capture,
encoding, instrumentation and QoS policy are outside the provider's control — the provider and
the client are two teams, two codebases.

PPCDN's ppobs is a deeply customized publisher, co-designed with the server under one
architecture, enabling coordinated optimization a "buy an SDK" model can't:

  • Absolute timestamps embedded in the bitstream (SEI for H.264/HEVC, metadata OBU for AV1) let the player compute real end-to-end latency — this requires changing encoder/packetizer logic, not a config knob, and a generic SDK can't.
  • P2P and the main publish share one encoding pass, avoiding per-viewer re-encoding and saving publisher CPU; independent send queues and bandwidth budgets for P2P and the main publish keep "turn on P2P to save money" from hurting the host's quality.
  • Weak-network degradation (lower bitrate first, without interrupting the publish) is triggered jointly by encoder and server, not by the server's one-sided rate-limit guess.

These coordinated optimizations are essentially the added value of "end-to-end programmability":
only by owning both ends can you tune the whole path as one system instead of bolting two black
boxes together.

5.2 WHIP/SRT simulcast: multiple layers from one client encode pass, no server transcode

Multi-layer quality (e.g. 1080p/720p/360p simulcast) is usually achieved by server-side
transcoding
in big clouds — Tencent Cloud LEB bills transcoding per output layer, so more
layers and longer duration cost more. That's industry practice, not an exception.

ppobs uses WHIP/SRT simulcast to output multiple layers from the encoder in one pass; the
server only forwards, no decode, no transcode. For customers who need multiple layers anyway,
eliminating transcoding often cuts the total bill substantially — in many setups, multi-layer
transcoding is the second-largest cost after bandwidth, and removing it materially changes the
bill structure.

5.3 ppobs + pplayer coordinated direct connect: ~70ms, really measured

Whether P2P direct connect works and how low latency can go depends not only on "having a NAT
traversal library" but on whether publisher and player can coordinate eligibility decisions and
timestamp alignment (see §3.1, diagram 4.2). This project's P2P direct connect is already
proven, with end-to-end latency of about 70ms; the ingest path delivered via Edge measured
about 108ms for WHIP and about 386ms for SRT (the gap comes from the ingest protocol's
receive window, not the encoding format).


6. Wrap-up and next episode (script notes)

This episode's architecture diagrams answered EP01's three dilemmas: codec-aware routing adds
a shorter P2P direct path for H.264 clients, the real-time protocol choice escapes TCP
head-of-line backlog risk, and separating the control and media planes while scaling Edge
horizontally shrinks the failure domain and keeps control in your hands. With a custom client,
no server transcoding and end-to-end coordination, PPCDN delivers a self-hosted solution that
is clearly stronger on latency, cost and controllability.

Next episode, we tackle "sovereignty" first — why "under your control" is a harder moat than
any performance number, and how to weigh self-hosting against relying on a third party.

Top comments (0)