DEV Community

Cover image for Scalable Video Streaming Architecture: Micro‑services vs Monolith
Amitesh0512
Amitesh0512

Posted on Originally published at amiteshsurwar.com

Scalable Video Streaming Architecture: Micro‑services vs Monolith

Quick Answer

Learn how to build a production‑grade, scalable video streaming architecture using microservices, CDNs, and analytics—optimized for latency, cost, and reliability.

Single CDN Push Fails During Viewer Surge

When a 200k‑viewer spike hits a live concert stream, the naive “push to a single CDN and call it a day” approach collapses under a cocktail of bandwidth bursts, cache misses, and a single monolith that can’t autoscale. In the wild, this manifests as 3‑second latency spikes, 30‑second re‑buffer bursts, and a 15 % spike in egress costs that the finance team didn’t budget for.

Real‑world Example

Last year a music‑festival operator ran a live event from a single Azure region. They had a monolith that ingested RTMP, ran FFmpeg, and served HLS from a shared Blob container. During the peak 8‑minute set, the ingestion queue filled to 200 k messages, the FFmpeg process stalled, and the CDN reported a 60 % cache‑miss rate. The result: 25 % of viewers saw >5 s of re‑buffering and 18 % abandoned the stream within the first minute.

Trade‑offs

  • Monolith vs. Micro‑services: A single process is easier to ship but forces you to scale everything together. Splitting the pipeline into stateless services lets you autoscale the transcoder independently, but you pay for the operational overhead of a message bus and service discovery.
  • Live vs. Batch Ingest: A low‑latency RTMP‑to‑HLS path requires in‑memory packet buffers and a tight end‑to‑end pipeline, but it exposes you to jitter in the upstream source. A batch ingest pipeline can tolerate upstream hiccups but introduces a 30‑second delay that kills live‑event UX.
  • CDN Caching Strategy: Caching every rendition at every edge node maximizes hit rates but inflates storage costs and can lead to stale DRM‑protected content. A tiered cache that only stores the most popular bitrates reduces egress but requires a more sophisticated cache‑invalidation logic.
  • Spot vs. On‑Demand Transcoding Workers: Spot VMs cut transcoding costs by 60 % but can be reclaimed at any time, forcing you to keep a fallback on‑demand pool or risk dropping segments. On‑demand VMs offer stability at 2× cost.
  • Stateless vs. Stateful Queues: Kafka gives you ordering guarantees and replayability but adds a 12 h retention cost. Azure Service Bus offers simpler APIs but lacks the same throughput for high‑volume live events.

Latency KPIs, Queue Choice & Transcoder Scaling

  1. Define the Latency KPI – If sub‑second latency is a product requirement, go for a dedicated RTMP ingestion service with a single‑partition Kafka topic per channel and keep the transcoder warm.
  2. Choose the Queue – For >100k concurrent viewers, Kafka’s high throughput and partitioning is a must. Use idempotent producers and sequence numbers to guard against re‑balance induced ordering loss.
  3. Transcoder Scaling – Deploy a warm pool of 3 FFmpeg pods per channel. Use an HPA that scales on queue_length and set a maxReplicas that respects your spot‑VM quota.
  4. CDN Tiering – Cache 720p/1080p at Tier‑1 POPs, 4K at Tier‑2 regional caches, and pull the rest from origin. Use short‑lived signed URLs (30 s) and push a CDN purge on DRM revocation.
  5. Compliance – Run ingestion in the country of origin, store raw chunks locally, and replicate only encrypted HLS fragments to global edges. Use Azure Private Link to keep traffic inside the VNet.
  6. Observability – Emit OpenTelemetry metrics for buffer_health and rebuffer_seconds. Trigger autoscaling or alerts when rebuffer_seconds > 2 s for >10 % of sessions.
  7. Cost Optimisation – Run a cost simulation: 60 % spot + 40 % on‑demand transcoding yields a 35 % reduction in egress while keeping 95 % uptime.

When This Fails in Production

  • Kafka Partition Rebalance During Live – A rebalance can drop the ordering of chunks, causing playback gaps. Fix: use a single partition per channel or enable idempotent producers with sequence numbers.
  • DRM Token Cache Stampede – 10k simultaneous token refreshes can overwhelm the license server. Fix: local in‑memory cache with a 5 s TTL and a leaky bucket limiter on the license endpoint.
  • Cold‑Start Transcoder Latency – FFmpeg containers take 8 s to start, dropping the first few seconds of a live stream. Fix: maintain a pool of 2–3 warm transcoder pods and a readiness probe that waits for the first FFmpeg binary load.
  • Edge Cache Invalidation Lag – A revoked token may still be served from a stale cache for up to 5 min. Fix: push a CDN purge immediately on token revocation and use a short signed URL TTL.
  • Unbounded Queue Growth – If the ingestion rate spikes faster than transcoding can keep up, the raw‑chunks queue grows unbounded. Fix: back‑pressure the ingestion service by exposing a queue depth metric and throttling incoming RTMP packets.

Common Mistakes Engineers Make

  • Assuming the CDN will automatically handle DRM‑protected content; many CDNs cache the license URL as a static asset.
  • Using a shared Blob container for both raw and transcoded assets; this causes contention and unpredictable IOPS.
  • Relying on CPU metrics for HPA; transcoding is I/O bound, so queue depth is a better metric.
  • Neglecting to purge the CDN on DRM revocation; stale keys can expose content for hours.
  • Over‑optimising for cost by running all transcoding on spot VMs without a fallback; this leads to dropped segments during spot evictions.

Better Approach Based on Experience

In a production environment, I would:

  • Use a dedicated RTMP ingestion service per region with a single‑partition Kafka topic to guarantee ordering and minimal latency.
  • Implement a warm pool of FFmpeg containers that are kept alive in a ready state and only spun up when the queue depth exceeds 200.
  • Adopt a multi‑tier CDN strategy, caching only the most popular bitrates at the edge and using origin shields to reduce duplicate pulls.
  • Leverage Azure Private Link and Azure Front Door for secure, low‑latency routing to the CDN while keeping all traffic within the VNet.
  • Instrument everything with OpenTelemetry and set up automated alerts for >5 % rebuffer spikes, which trigger an autoscaling rule that adds transcoder pods in the affected region.
  • Run a cost‑simulation before launch that mixes spot and on‑demand instances, validates the CDN purge latency, and ensures the DRM license server can handle a 5 k requests/second burst.
Architecture Option Latency Impact Cost Impact Reliability Impact
Microservices + CDN + Real‑time Analytics Low—edge caching + immediate data processing reduces buffering Higher—multiple services and real‑time data pipelines High—service isolation, auto‑scaling, real‑time failover
Microservices + CDN + Batch Analytics Low—edge caching same as above Moderate—batch jobs less frequent than real‑time High—service isolation, auto‑scaling, but delayed anomaly detection
Monolithic + CDN Medium—single deployment, may not scale to edge nodes quickly Low—fewer services, simpler ops Medium—single point of failure, harder to isolate faults

Performance Considerations

  • Use ffmpeg -threads 8 and -preset veryfast to balance CPU usage and encoding speed.
  • Store raw chunks in Azure Blob Storage with the Hot tier and set the Cache-Control: immutable header for subtitles to push them into edge TTLs forever.
  • Set the CDN minTTL to 30 s for HLS segments; this prevents the CDN from fetching the same segment from origin for each request.
  • Enable origin shield on the regional node to avoid duplicate pulls from storage during a flash surge.
  • Use a maxReplicas of 30 for transcoder HPA to handle a 200k viewer spike, but cap the total CPU usage at 80 % to avoid thrashing.

Scaling Notes

  • Scale the ingestion service horizontally by adding more RTMP listeners; each listener writes to a dedicated Kafka partition.
  • Use a queue_length metric that aggregates across all partitions to decide when to spin up transcoder pods.
  • For global events, deploy the CDN edge nodes in all major regions and keep the transcoder pool region‑specific; this avoids cross‑region data transfer costs.
  • Implement a back‑pressure signal from the transcoder to the ingestion service when the queue depth > 5000, causing the ingestion service to drop non‑essential metadata packets.

Migration Checklist

  1. Catalog all API endpoints and map them to the four target services.
  2. Containerize each service and externalise all configuration via Kubernetes secrets.
  3. Deploy a staging cluster with HPA enabled; run a 5 % traffic shadow test.
  4. Introduce Kafka topics incrementally – start with raw‑chunks only, keep the legacy pull‑based transcoder as a fallback.
  5. Enable CDN edge caching for a single region; validate cache‑hit ratios before expanding globally.
  6. Instrument every service with OpenTelemetry; set up alerts on >5 % increase in rebuffer rate.
  7. Run a cost‑simulation (Azure Pricing Calculator) for spot‑vs‑on‑demand transcoder mix.
  8. Cutover live traffic during a low‑viewership window; keep the monolith on standby for 48 hours.

Conclusion

Scalable video streaming is not a set of optional knobs but a disciplined architecture where latency, cost, and compliance are baked into every layer. Treat sub‑second latency as a product KPI, not a after‑thought, and you’ll halve churn for live‑event platforms and unlock premium pricing tiers.

Related Articles

Top comments (0)