DEV Community

Conexa Remote
Conexa Remote

Posted on

Building a Sub-50ms Remote Desktop: Why Direct WebRTC P2P Beats Relay-Heavy Cloud Architectures

When building or selecting remote desktop tools, developers and IT administrators frequently hit a common wall: input lag and latency jitter.

In this article, we’ll look under the hood at why traditional cloud-relay architectures introduce latency bottlenecks and how modern direct Peer-to-Peer (P2P) pipelines powered by WebRTC and native desktop duplication APIs can achieve sub-50ms glass-to-glass latency.


1. The Bottleneck: Cloud Relays vs. Direct Peer-to-Peer

In traditional client-server remote control architectures, the remote machine captures its screen, encodes the frame, and sends the video packet to a centralized cloud relay cluster. The relay then forwards the packet down to the technician's workstation.

While cloud relays make NAT traversal trivial, they introduce three fundamental issues:

  1. Compounded Round-Trip Time (RTT): The video stream travels from Machine A → Data Center → Machine B. Even on fast fiber, routing packets through an intermediary server thousands of kilometers away adds 60–150ms of unavoidable latency.
  2. Bandwidth Saturation & Transcoding Overhead: When thousands of concurrent sessions hit relay servers, providers often throttle bitrate, drop framerates, or compress color depth (4:2:0 subsampling), causing blurry fonts.
  3. Data Sovereignty Concerns: Having a third-party server process active desktop video feeds introduces privacy and compliance overhead.

The P2P WebRTC Alternative

With direct WebRTC peer-to-peer streaming:

  • The cloud server acts only as a lightweight signaling / rendezvous coordinator to exchange SDP (Session Description Protocol) offers and ICE candidates.
  • Once the direct socket is established via STUN, 100% of media and input data flows directly between the two endpoints over encrypted DTLS-SRTP.
  • The video stream never touches a middleman server during the session.
[Tech Workstation] <==========================> [Target Machine]
         \         (Direct WebRTC DTLS-SRTP)         /
          \                                         /
           \===> [Signaling / STUN Server] <======/
                     (Only for 0.5s Handshake)
Enter fullscreen mode Exit fullscreen mode

2. Low-Overhead Capture: DXGI & ScreenCaptureKit

Networking is only half the equation. If your screen capture takes 30ms before encoding even starts, sub-50ms glass-to-glass latency is impossible.

On Windows: Desktop Duplication API (DXGI)

Instead of legacy GDI captures (BitBlt), modern remote desktop clients tap directly into the GPU pipeline via the DirectX Graphics Infrastructure (DXGI) Desktop Duplication API:

  • Frames are acquired directly from GPU video memory (VRAM).
  • Frame buffers are processed without unnecessary host memory (RAM) round-trips (zero-copy GPU rendering).
  • Dirty rect tracking ensures only changed screen regions are pushed to the video encoder.

On macOS: ScreenCaptureKit

On macOS 12.3+, legacy CGDisplayStream APIs have been replaced with Apple's modern ScreenCaptureKit:

  • Native hardware-accelerated screen capture running directly on Apple Silicon / Metal engines.
  • Significantly lower CPU utilization and consistent 60fps frame delivery.

3. Real-World Trade-Offs in P2P Architectures

Direct P2P architectures offer massive speed and privacy advantages, but they require robust engineering around real-world edge cases:

  1. Symmetric NATs & Enterprise Firewalls: Around 10–15% of enterprise networks use strict symmetric NATs where direct UDP hole punching fails. In these rare cases, seamless fallback to TURN relaying is essential.
  2. Congestion Control: Network jitter requires dynamic bitrate adaptation (e.g., Google Congestion Control / GCC) to adapt resolution on degraded Wi-Fi without crashing the session.
  3. Security Handshake: Encrypting all data channels with DTLS 1.3 and media tracks with SRTP (AES-256) ensures privacy equivalent to a direct VPN tunnel.

4. Our Implementation at Conexa Remote

When we designed Conexa Remote, we applied these exact engineering principles:

  • Sub-50ms Glass-to-Glass Latency: Leveraging native DXGI, ScreenCaptureKit, and WebRTC data channels for fluid, responsive mouse and keyboard input.
  • Clientless Ad-Hoc Connections: For attended helpdesk support, end users don't need to create accounts, configure passwords, or install bloated runtimes. They open the standalone binary and read an instant 9-digit PIN.
  • Floating Concurrent Capacity Model: Instead of rigid named-user licensing taxes, businesses share a floating pool: the Professional tier (29 €/mo for 2 concurrent active connections) includes unlimited technician logins and unlimited endpoints in the address book (+5 €/mo per extra concurrent slot).
  • 100% EU Infrastructure: Rendezvous and signaling infrastructure is hosted exclusively within the EU (Romania), with zero third-party telemetry.
  • Permanent Free Tier: 1 concurrent session with zero time limits or artificial session timeouts.

Conclusion

Real-time remote control shouldn't feel like driving a rover on Mars. By bypassing relay proxies in favor of direct WebRTC peer-to-peer pipelines and modern GPU capture APIs, remote support can be instantaneous, lightweight, and respectful of user privacy.

Feel free to inspect the architecture and test the Windows/macOS builds at conexaremote.com. We'd love to hear feedback from fellow network and systems engineers!

Top comments (0)