Choosing an NVR keeps running into the same wall: there are no public performance numbers with stated measurement conditions. Third-party cross-product benchmarks don't exist; what you can find is either vendor-quoted partial metrics (detector inference latency, say) or editorial estimate tables that omit the hardware, bitrate, and framerate conditions — the numbers can't be compared against each other.
So I measured MiBeeNvr's own capacity and resource curves: three hardware tiers, recording ramps from 16 to 68 cameras, concurrent playback on top of full load, pure forwarding, an fsync-policy comparison, plus a read-only observation of a 7×24 production machine. The methodology, pass/fail criteria, and test scripts are all published with this post. The goal is not to rank products — with no comparable public methodology, ranking means nothing — but to answer two questions that actually come up during capacity planning: how many cameras a machine of a given tier can carry, and whether the bottleneck lands on CPU, memory, or storage.
Conclusions first, data and tools below:
- High-spec x86: recording 64×1080p + 4×4K at 24% total CPU (16-thread basis) and 171 MB of NVR process memory, API p50 in the low tens of milliseconds throughout; stacking 32 concurrent playback sessions (FLV / LL-HLS / RTSP) on top produced zero failures. The tier stops at 68 cameras because the traffic generator ran out of headroom, not because the NVR struggled.
- Entry ARM NAS (RK3399 + eMMC-class storage): CPU sits at two tenths while the bottleneck is entirely in storage — once the volume saturates, the health endpoint degrades to 2 s, but recording, per-camera accounting, and recording queries are unaffected. Switch to 30-camera pure forwarding and CPU drops to three tenths with a perfectly clean API distribution: plenty of headroom as a live gateway.
- ARM SBC production box (Banana Pi M5): 18 real cameras running 7×24, API p50 still 13.5 ms during a window with 80% disk iowait — two orders of magnitude away from the saturated eMMC tier's 2 s. Storage medium sets the experience ceiling of the low tier, and that conclusion reproduced in production.
These are vendor self-test numbers, and I'm not going to pretend otherwise: every figure's measurement conditions are in this post, the tools and criteria are public, and anyone can rerun the thing end to end.
Test Environment
| Device | Role | CPU | RAM | Storage | Deployment |
|---|---|---|---|---|---|
| Banana Pi M5 | Production (read-only observation) | Amlogic S905X3, ARMv8 ×4 | 4GB | 2.7TB SATA | bare-metal systemd |
| RK3399 NAS | Controlled benchmark (low tier) | RK3399 ×6 | 4GB | eMMC-class volume | fnOS + Docker |
| x86 workstation | Controlled benchmark (high tier) | i5-12600KF, 10C16T | 64GB | NVMe | bare metal |
| x86 laptop | Traffic generation / idle measurement | 8-core | 16GB | SSD | — |
Two environment details are worth stating up front. The RK3399 NAS's storage volume is a single-member RAID1 with thin provisioning on an eMMC partition — 2× write amplification. That detail keeps resurfacing in the data below; it is the key to reading this tier. And the production machine happened to be running a historical-archive migration during the observation window, so the same device yielded both a "heavy IO stress" and a "business as usual" window.
The build under test was a single unified v0.13.0 preview build; the production observer ran a preview build from the same series.
Criteria: How It Was Measured and Counted
The four-piece load and measurement setup:
-
Synthetic cameras: pre-rendered MP4s pushed in real time (
-repacing) over RTMP, at two specs — 1080p25 H.264 2.5 Mbps (2 s GOP, no B-frames) and 4K25 H.264 16 Mbps. The pushers reconnect on drop, matching real camera behavior. Per-camera accounting comes from periodic snapshots of the NVR's/api/streamscounters to verify each camera's ingest bitrate. -
Resource sampling: on Linux, 5-second reads of raw
/proc/statcounters (so iowait can be broken out),/proc/meminfo, and NVR process RSS; the Windows host sampled performance counters at the same cadence. Disk figures come from iostat. -
API probe: every 5 seconds, request
/api/health(unauthenticated, lightweight) and/api/recordings?limit=5(auth + SQLite query path), recording the full latency distribution and failure count. - Playback: real pull clients on FLV / LL-HLS / RTSP watching concurrently, measuring throughput, time-to-first-byte, and live-edge latency (LL-HLS computed from the playlist's PROGRAM-DATE-TIME).
Judgment criteria: CPU includes soft interrupts, iowait is broken out separately; the "steady-state window" is the 4 minutes starting 60–75 s after the last pusher starts; API latencies in tables are percentiles over that window. Every curve is backed by 5-second raw data.
Configuration criteria: except for the fsync comparison section, everything ran on factory defaults — storage.durability at its default strict level (fsync on every finalized segment), io.budget_bytes_per_sec at 0 (background IO budget off), preallocation on. The headline numbers are therefore measured in the default posture, with no opt-in tuning switches polishing them.
Load conversion: using the community's usual MP/s dimension, 1080p25 = 51.8 MP/s per camera, 4K25 = 207.4 MP/s. This run's x86 peak of 68 cameras ≈ 4,145 MP/s; RK3399 pure forwarding of 30 cameras ≈ 1,555 MP/s. Note that MiBeeNvr records in the compressed domain without decoding, so MP/s here is a throughput dimension — comparable against generic planning methods, not equivalent to decode-style workloads.
The scripts behind this setup are collected in the appendix at the end of the post.
High-Spec x86: Recording Ramp and Concurrent Playback
From 16 cameras up to 68 (64×1080p + 4×4K), 3 minutes of steady state per tier, with per-camera ingest held precisely at 2.5/16 Mbps throughout:
| Load | Total ingest | Total CPU (16T) | NVR process memory | health p50/p95 | recordings p50/p95 |
|---|---|---|---|---|---|
| Baseline (1 real camera) | 2.0 Mbps | 11.4% | 57 MB | — | — |
| 16×1080p | 42.2 Mbps | 12.8% | 82 MB | 8.5 / 15.2 ms | 9.4 / 38.5 ms |
| 32×1080p | 82.3 Mbps | 14.3% | 105 MB | 8.9 / 14.5 ms | 9.0 / 14.0 ms |
| 48×1080p | 122.4 Mbps | 15.3% | 131 MB | 9.6 / 16.7 ms | 9.6 / 12.5 ms |
| 64×1080p | 162.4 Mbps | 19.7% | 156 MB | 10.6 / 18.2 ms | 10.1 / 22.7 ms |
| 68 cameras (4×4K included) | 226.5 Mbps | 24.2% | 171 MB | 8.7 / 15.0 ms | 11.1 / 27.6 ms |
Both curves are near-linear: each added 1080p camera costs about 3 MB of memory, and CPU climbs roughly 2 percentage points per 40 Mbps (16-thread basis). The API probe recorded zero failures the whole way (also zero during the 32-viewer stage), and health returned to 1.2–1.9 ms once the test ended.
On top of the full 68-camera recording load, 32 concurrent playback sessions (12 FLV + 12 LL-HLS + 8 RTSP, 5 minutes):
| Protocol | Result |
|---|---|
| FLV | 12/12 sessions returned 200, each at a precise full 2.50–2.51 Mbps; TTFB min/p50/p95 = 65/167/256 ms |
| LL-HLS | 12/12 streamed normally; startup min/p50/p95 = 438/1414/2095 ms; median live-edge latency 1.3 s |
| RTSP | 8 cameras, three rounds of pulls, zero errors |
Entry ARM NAS: The Bottleneck Is Storage, Not CPU
Recording Ramp (8 → 28 cameras, 20 → 124 Mbps)
| Load | CPU | iowait | NVR memory | health p50 | recordings p50 |
|---|---|---|---|---|---|
| 8×1080p (20 Mbps) | 20.1% | 59% | 130 MB | 24 ms (drifting into saturation late in the tier) | 20 ms |
| 16×1080p (40 Mbps) | 20.0% | 73% | 162 MB | 2.0 s | 22 ms |
| 24×1080p (60 Mbps) | 22.6% | 72% | 203 MB | 2.0 s | 21 ms |
| 24×1080p+4×4K (124 Mbps) | 19.7% | 73% | 197 MB | 2.0 s | 22 ms |
The CPU curve flattens at 20–23%, with the NVR process peaking at 41.8% of a single core. The saturation point is storage: iostat shows the volume at 100% utilization, write waits of 3.4–6.9 s, and actual writes of about 6–7 MB/s (RAID1 amplification included). Behavior after saturation is layered:
- Recording and per-camera accounting are unaffected — per-camera bitrate holds and segments keep flushing;
- Recording queries (the SQLite read path) are unaffected, p50 steady at 20–22 ms throughout;
- The health check degrades to 2 s — internally it runs a statfs against the storage volume, which blocks when the device saturates, making it a surprisingly good "storage is gasping" indicator;
- Playback under extreme saturation: LL-HLS delivered data to 8/8 clients but startup stretched to 88 s and live edge ran 35–53 s; RTSP kept streaming; FLV could not establish new sessions within 15 s (106–135 ms when healthy).
Recovery after load removal is quick: health p50 returned to 17–23 ms within the first minute, though background IO kept individual requests at 2 s for about 6 minutes, after which max < 31 ms. No crashes, no process restarts at any point.
Clean Low-Load Tier (4 cameras, 10 Mbps)
| CPU | iowait | NVR memory | health p50/p95 | recordings p50/p95 |
|---|---|---|---|---|
| 16.7% | 44% | 77 MB | 24 ms / 1485 ms | 21 ms / 188 ms |
Even at 10 Mbps the odd statfs spike appears — this tier's bottleneck is entirely the eMMC volume itself, not CPU and not the software.
Pure Forwarding Tier (30 live cameras, no recording, 75 Mbps)
With 30 cameras set to live-only (recording disabled):
| CPU | iowait | NVR memory | health p50/p95/max | recordings p50/p95 |
|---|---|---|---|---|
| 30.5% | 15% | 95 MB | 21 ms / 40 ms / 62 ms | 21 ms / 30 ms |
With the write path out of the picture, the API distribution over 30 ingests plus distribution is perfectly clean; pulling FLV from a live camera gave 200, 2.52 Mbps, 135 ms TTFB. As a pure live gateway (preview / cascade / re-publish scenarios) the RK3399 has ample margin — roughly 1,555 MP/s of throughput on three tenths of the CPU.
fsync Policy Comparison: strict vs relaxed
durability controls the fsync policy for raw segments: the default strict fsyncs on every finalized segment; relaxed skips explicit fsyncs and leaves it to the filesystem commit interval (a power cut may lose the last few seconds of raw segments; merged/timelapse artifacts are unaffected). To quantify what the switch buys on slow storage, I ran a paired test at identical load and background (rolling-merge jobs running in both cases):
| Tier | durability | health p50 / p95 | iowait | CPU | NVR memory |
|---|---|---|---|---|---|
| 16×1080p (40 Mbps) | strict | 2024 / 3045 ms | 72.5% | 20.0% | 162 MB |
| 16×1080p (40 Mbps) | relaxed | 2019 / 2027 ms | 73.2% | 21.1% | 171 MB |
| 24×1080p (60 Mbps) | strict (first run) | 2022 / 2033 ms | 72.0% | 22.6% | 203 MB |
| 24×1080p (60 Mbps) | strict (paired re-run) | 2020 / 2035 ms | 68.4% | 24.9% | 217 MB |
| 24×1080p (60 Mbps) | relaxed | 2023 / 2040 ms | 54.7% | 31.7% | 131 MB |
Instantaneous write-path metrics did improve: under strict steady state the volume showed write waits of 3.4–3.9 s and constant 100% utilization; switching to relaxed brought write waits down to 0.2–0.3 s and turned utilization from "continuously saturated" to "bursty". But across the whole window, health latency and the saturation threshold did not move — this NAS's storage pressure is the sum of several sources (2× write amplification from the single-member RAID1, background bulk IO from rolling merges, the system's own background reads), and dropping recording fsyncs merely converted sustained pressure into bursty pressure, not enough to pull the device back out of saturation.
Conclusion: on eMMC-class slow storage, relaxed smooths write queuing and is a reasonable option for deployments that tolerate power-cut loss, but it does not change the capacity ceiling of such storage; on NVMe-class fast storage the default tier is never the bottleneck, so there is no need to flip it. This section also confirms that the headline numbers were all measured at the default strict level, with no help from this switch.
A footnote: during the comparison, one docker restart of the container hit a runtime hang (docker ps reported Up while the process was gone; stop/start recovered it; the log showed no panic before exit and the kernel logged no OOM — a host-environment event, not an NVR crash). SQLite's unchecked-pointed stream keys were lost as a result (cameras and recordings intact); I recreated them and continued. In production, restart via the normal mechanisms (systemd / app store) rather than docker restart.
Production SBC: Two Observation Windows
A Banana Pi M5 (S905X3 ×4 / 4GB / 2.7TB SATA) runs 18 real cameras (a mix of ONVIF / GB28181 / Xiaomi / MJPEG) 7×24, with 2.1 TB of recordings on disk (a single camera at up to 467 GB / 2014 segments). Read-only observation, no changes made:
| Metric | Calm window (14 min) | Heavy-IO window (46 min)¹ |
|---|---|---|
| health p50 / p95 / max | 12.2 / 25.3 / 1015 ms | 13.5 / 53 / 2232 ms |
| recordings p50 / p95 | 252 / 306 ms | 253 / 338 ms |
| Ingest bitrate / active cameras | 16.1 Mbps / 12 | 17.9 Mbps / 12 |
| NVR process RSS (min/avg/max) | 182 / 267 / 414 MB | 114 / 194 / 412 MB |
| SoC temperature (avg / max) | 51.2 / 54.5 ℃ | 52.6 / 58.3 ℃ |
| Total CPU (incl. iowait conversion) | 20.1% | 34.5% (mostly iowait) |
¹ The machine was migrating its historical archive during the stress window, with disk iowait around 80%; the NVR process itself used about half a core, and recording continued.
The same SATA device held health p50 at 13.5 ms under 80% iowait, versus 2 s for the eMMC volume under comparable pressure earlier — two orders of magnitude apart. The storage medium sets the experience ceiling of the low tier, and that conclusion reproduced in production. The 250 ms scale of recordings queries matches its 2.1 TB / tens-of-thousands-of-rows library.
Boundaries and Reproducing
Not covered: software transcoding scale (H.265→H.264 software transcoding is CPU-bound; production should avoid it by recording sub-streams), WebRTC end-to-end latency (needs browser-side measurement), 10 GbE, cloud-archive throughput, audio tracks, and the io.budget_bytes_per_sec background IO budget (off by default; background merge load was a constant during the test window, not A/B tested separately). Synthetic streams are video-only.
Reproducing: the tools are the appendix scripts — synthetic camera sources, the pusher fleet, resource samplers, API probe, per-camera snapshots, the three-protocol pull clients, and the ramp orchestration — pointed at any target host via environment variables. A same-hardware cross-product comparison (say, Frigate or ZoneMinder on the same box under the same load) has no third-party public data either; it would be worth doing, but first someone has to put every product under the same methodology.
Repository: Mi-Bee-Studio/MiBeeNvr.
Appendix: The Test Harness
The four-piece setup from the methodology section lands as the scripts below. The original test kit hardcoded LAN addresses, passwords, and machine-specific choreography; before publishing I did two things to all of it: sanitized (addresses and credentials now come from environment variables) and generalized (process names, durations, camera counts, and ramp steps are all parameters — nothing is tied to a specific machine). The tools only do measurement; the product-specific parts reduce to two strings, the API paths and the process name. The pusher and pull clients work against any RTMP / HTTP-FLV / LL-HLS service as-is.
Deployment: the traffic generator runs the pusher fleet and the pull clients; the device under test runs the resource sampler; the API probe can sit on either side. Render the synthetic sources once and reuse them across runs.
Synthetic Camera Sources
testsrc2 renders video-only sources with a 2 s GOP and no B-frames, specced like a real IPC main stream:
#!/bin/bash
# render-cam.sh — render one synthetic camera source (video-only)
# usage: render-cam.sh <WxH> <fps> <bitrate_k> <duration_s> <out.mp4> [profile]
set -e
RES=${1:-1920x1080}; FPS=${2:-25}; BR=${3:-2500}; T=${4:-1800}; OUT=${5:?}; PROF=${6:-main}
ffmpeg -y -loglevel error -f lavfi -i "testsrc2=size=${RES}:rate=${FPS}" -t "$T" \
-c:v libx264 -preset veryfast -profile:v "$PROF" -b:v "${BR}k" \
-maxrate $(( BR * 12 / 10 ))k -bufsize $(( BR * 2 ))k \
-g $(( FPS * 2 )) -bf 0 -pix_fmt yuv420p "$OUT"
# One source per spec (4K uses the high profile):
bash render-cam.sh 1920x1080 25 2500 1800 cam1080.mp4 main
bash render-cam.sh 3840x2160 25 16000 600 cam4k.mp4 high
Pusher Fleet
A single pusher is a reconnect loop — real cameras re-push after a network blip, so the load model can't skip that; the fleet script handles batch start/stop and counting:
#!/bin/bash
# pusher-loop.sh — one pusher: reconnects 2 s after a drop, mimicking a real camera
# usage: pusher-loop.sh <rtmp base> <media file> <stream key>
BASE=$1; FILE=$2; KEY=$3
while true; do
ffmpeg -loglevel error -re -i "$FILE" -c copy -f flv "$BASE/$KEY" >> "pusher_${KEY}.log" 2>&1
sleep 2
done
#!/bin/bash
# pushers.sh — pusher fleet: start <rtmp base> <file> <key...> | stop | count
MODE=$1; PIDF=pushers.pid
case $MODE in
start)
BASE=$2; FILE=$3; shift 3
for KEY in "$@"; do
nohup bash pusher-loop.sh "$BASE" "$FILE" "$KEY" >/dev/null 2>&1 &
echo $! >> "$PIDF"
sleep 0.4 # stagger startups to avoid a connect storm
done
echo "started=$#";;
stop)
pkill -f pusher-loop.sh 2>/dev/null
[ -f "$PIDF" ] && { kill $(cat "$PIDF") 2>/dev/null; rm -f "$PIDF"; }
sleep 1
pkill -9 -f 'rtmp://.*/live/' 2>/dev/null # catch leftover ffmpeg; don't run this near unrelated ffmpeg jobs
echo stopped;;
count)
pgrep -fc 'rtmp://.*/live/' || echo 0;;
esac
Resource Samplers
Reading raw /proc counters instead of pre-computed percentages is what lets iowait be broken out — the bottleneck calls on the low tier hinge on it. The process name is a parameter, so it samples any service:
#!/bin/sh
# sampler.sh — target-host resource sampler: raw counters, one line every 5 s by default
# usage: sampler.sh <out.csv> [interval_s] [process name]
OUT=${1:?}; INT=${2:-5}; PROC=${3:-mibee-nvr}
echo "ts,user,nice,sys,idle,iowait,irq,softirq,steal,mem_avail_kb,proc_rss_kb,pgpgout" > "$OUT"
while true; do
sleep "$INT"
C=$(awk '/^cpu /{print $2","$3","$4","$5","$6","$7","$8","$9}' /proc/stat)
MA=$(awk '/MemAvailable/{print $2}' /proc/meminfo)
PID=$(pgrep -x "$PROC" | head -1)
[ -z "$PID" ] && PID=$(pgrep -f "$PROC" | head -1)
RSS=$(awk '/VmRSS/{print $2}' "/proc/$PID/status" 2>/dev/null)
PG=$(awk '/pgpgout/{print $2}' /proc/vmstat)
echo "$(date +%s),$C,$MA,${RSS:--1},$PG" >> "$OUT"
done
Windows targets get the performance-counter variant:
# sampler.ps1 — Windows target sampler, same idea as the Linux one
param([string]$Out = "sample.csv", [int]$Interval = 5, [string]$ProcName = "mibee-nvr")
"ts,cpu_busy,mem_avail_mb,proc_ws_mb,disk_w_mb_s" | Out-File $Out -Encoding utf8
while ($true) {
Start-Sleep $Interval
try {
$cpu = (Get-Counter '\Processor(_Total)\% Processor Time' -SampleInterval 1 -MaxSamples 1).CounterSamples.CookedValue
$os = Get-CimInstance Win32_OperatingSystem
$p = Get-Process $ProcName -ErrorAction SilentlyContinue | Select-Object -First 1
$ws = if ($p) { [math]::Round($p.WorkingSet64 / 1MB, 1) } else { -1 }
$dw = (Get-Counter '\PhysicalDisk(_Total)\Disk Write Bytes/sec' -SampleInterval 1 -MaxSamples 1).CounterSamples.CookedValue / 1MB
Add-Content $Out ("{0},{1:N1},{2},{3},{4:N1}" -f [DateTimeOffset]::UtcNow.ToUnixTimeSeconds(), $cpu, [math]::Round($os.FreePhysicalMemory / 1KB), $ws, $dw)
} catch { }
}
API Probe
The two paths mean different things: health is unauthenticated and lightweight; recordings carries auth plus a SQLite query — a real read path. The probe only logs timestamps and status codes; the latency distribution is post-processing's job:
#!/usr/bin/env python3
"""apiprobe.py — API latency probe
usage: NVR_API=http://<host>:<port> NVR_USER=... NVR_PASS=... python3 apiprobe.py <out.csv> [interval_s]"""
import os, sys, time, base64, urllib.request
base = os.environ["NVR_API"].rstrip("/")
basic = base64.b64encode(f'{os.environ.get("NVR_USER", "")}:{os.environ.get("NVR_PASS", "")}'.encode()).decode()
out = sys.argv[1]
interval = float(sys.argv[2]) if len(sys.argv) > 2 else 5
def timed_get(path, auth):
req = urllib.request.Request(base + path)
if auth:
req.add_header("Authorization", "Basic " + basic)
t0 = time.perf_counter()
try:
with urllib.request.urlopen(req, timeout=10) as r:
r.read(4096)
code = r.status
except Exception as e:
code = getattr(e, "code", -1)
return (time.perf_counter() - t0) * 1000, code
f = open(out, "a", buffering=1)
f.write("ts,health_ms,health_code,rec_ms,rec_code\n")
while True:
ts = int(time.time())
h = timed_get("/api/health", False)
r = timed_get("/api/recordings?limit=5", True)
f.write(f"{ts},{h[0]:.1f},{h[1]},{r[0]:.1f},{r[1]}\n")
time.sleep(interval)
Per-Camera Snapshots
Periodically recording the /api/streams counters and differencing adjacent snapshots yields each camera's actual ingest bitrate — what the pusher claims doesn't count; what the server receives does:
#!/usr/bin/env python3
"""statsnap.py — per-camera accounting snapshots
usage: NVR_API=... NVR_USER=... NVR_PASS=... python3 statsnap.py <out.jsonl> [interval_s]"""
import os, sys, time, json, base64, urllib.request
base = os.environ["NVR_API"].rstrip("/")
basic = base64.b64encode(f'{os.environ.get("NVR_USER", "")}:{os.environ.get("NVR_PASS", "")}'.encode()).decode()
out = sys.argv[1]
interval = float(sys.argv[2]) if len(sys.argv) > 2 else 30
def get(p):
req = urllib.request.Request(base + p)
req.add_header("Authorization", "Basic " + basic)
with urllib.request.urlopen(req, timeout=15) as r:
return json.loads(r.read())
f = open(out, "a", buffering=1)
while True:
rec = {"ts": int(time.time())}
try:
rec["streams"] = get("/api/streams")
except Exception as e:
rec["streams"] = {"err": str(e)}
f.write(json.dumps(rec) + "\n")
time.sleep(interval)
Pull Clients
The FLV client is small: throughput and time-to-first-byte. The HLS client does more — it follows the master playlist to the variant and computes live-edge latency from PROGRAM-DATE-TIME. Both scripts strip the query string before writing URLs to CSV, so stream tokens never leak into the data files:
#!/usr/bin/env python3
"""flvpull.py — FLV pull client: TTFB, throughput, status code
usage: python3 flvpull.py <stream.flv URL> <out.csv> <duration_s>"""
import sys, time, urllib.request
url, out, dur = sys.argv[1], sys.argv[2], float(sys.argv[3])
t0 = time.perf_counter()
n, ttfb, code = 0, None, 200
try:
with urllib.request.urlopen(url, timeout=15) as r:
code = r.status
while time.perf_counter() - t0 < dur:
b = r.read(65536)
if not b:
break
if ttfb is None:
ttfb = (time.perf_counter() - t0) * 1000
n += len(b)
except Exception as e:
code = getattr(e, "code", -1)
el = time.perf_counter() - t0
mbps = n * 8 / el / 1e6 if el > 0 else 0
with open(out, "a") as f:
f.write(f'{url.split("?")[0]},{ttfb if ttfb is not None else -1:.0f},{n},{el:.1f},{mbps:.2f},{code}\n')
print(f"flv code={code} bytes={n} ttfb={ttfb and round(ttfb)} mbps={mbps:.2f}")
#!/usr/bin/env python3
"""hlspull.py — HLS/LL-HLS pull client: startup time, throughput, live-edge latency (PDT)
usage: python3 hlspull.py <index.m3u8 URL> <out.csv> <duration_s>"""
import sys, time, re, urllib.request
from datetime import datetime, timezone
base, out, dur = sys.argv[1], sys.argv[2], float(sys.argv[3])
t0 = time.perf_counter()
seen = set()
n = 0
startup = None
lags = []
pdt_re = re.compile(r"^#EXT-X-PROGRAM-DATE-TIME:(.+)$", re.M)
var_re = re.compile(r"^#EXT-X-STREAM-INF:", re.M)
url = base
while time.perf_counter() - t0 < dur:
try:
with urllib.request.urlopen(url, timeout=5) as r:
txt = r.read().decode()
if startup is None:
startup = (time.perf_counter() - t0) * 1000
if var_re.search(txt): # master playlist -> follow to the variant
lines = txt.splitlines()
for i, ln in enumerate(lines):
if ln.startswith("#EXT-X-STREAM-INF"):
q = "?" + url.split("?", 1)[1] if "?" in url else ""
url = url.rsplit("/", 1)[0] + "/" + lines[i + 1].split("?")[0] + q
break
continue
m = pdt_re.search(txt) # live edge = local clock - PDT
if m:
try:
pdt = datetime.fromisoformat(m.group(1).replace("Z", "+00:00"))
lag = (datetime.now(timezone.utc) - pdt).total_seconds()
if -5 < lag < 120:
lags.append(lag * 1000)
except Exception:
pass
for ln in txt.splitlines():
if ln and not ln.startswith("#"):
s = ln.split("?")[0]
if s not in seen:
seen.add(s)
u = url.rsplit("/", 1)[0] + "/" + ln
try:
with urllib.request.urlopen(u, timeout=10) as r2:
n += len(r2.read())
except Exception:
pass
except Exception:
pass
time.sleep(0.4)
el = time.perf_counter() - t0
mbps = n * 8 / el / 1e6
lag_avg = sum(lags) / len(lags) if lags else -1
with open(out, "a") as f:
f.write(f'{base.split("?")[0]},{startup if startup is not None else -1:.0f},{n},{len(seen)},{mbps:.2f},{lag_avg:.0f}\n')
print(f"hls segs={len(seen)} bytes={n} startup={startup and round(startup)} mbps={mbps:.2f} live_lag_avg_ms={lag_avg:.0f}")
Ramp Orchestration
The recording ramp is data-driven: edit the TIERS array to change the ramp, and the script itself knows nothing about any particular machine. Each tier pushes, climbs for 60 s, holds steady for 180 s, snapshots per-camera accounting, and logs event timestamps into marks.txt for post-processing:
#!/bin/bash
# ramp.sh — recording ramp orchestration
# usage: NVR_API=http://<host>:<port> RTMP=rtmp://<host>:1935/live \
# NVR_USER=... NVR_PASS=... bash ramp.sh
set -e
: "${NVR_API:?} ${RTMP:?} ${NVR_USER:?} ${NVR_PASS:?}"
TIERS=("cam1080.mp4:16" "cam1080.mp4:16" "cam1080.mp4:16" "cam4k.mp4:4")
M=marks.txt
mark() { echo "$1 $(date +%s)" >> "$M"; }
snap() { curl -s -m 30 -u "$NVR_USER:$NVR_PASS" "$NVR_API/api/streams" > "streams-$(date +%s).json"; }
mark RAMP_START
n=0
for tier in "${TIERS[@]}"; do
n=$((n + 1)); FILE=${tier%%:*}; CNT=${tier##*:}
keys=$(seq 1 "$CNT" | sed "s/^/load-${FILE%.*}-n${n}-/")
mark "S${n}_start"
bash pushers.sh start "$RTMP" "$FILE" $keys
sleep 60; mark "S${n}_steady"
sleep 180; snap; mark "S${n}_end"
done
mark RAMP_DONE
Stacking concurrent playback on top of full load works the same way: resolve camera IDs from /api/cameras by name, hand the ID list to the pull scripts, start a batch each of FLV and LL-HLS, and spot-check RTSP with ffmpeg:
#!/bin/bash
# egress.sh — concurrent playback stacked on full recording load
# usage: NVR_API=... NVR_TOKEN=... RTSP_BASE=rtsp://<host>:8554 \
# bash egress.sh "<camera ID list, space-separated>" [duration_s]
IDS=${1:?"camera ID list, resolved from /api/cameras by name"}; DUR=${2:-300}
: "${NVR_API:?} ${NVR_TOKEN:?} ${RTSP_BASE:?}"
: > egress-flv.csv; : > egress-hls.csv
for id in $IDS; do
nohup python3 flvpull.py "$NVR_API/api/cameras/$id/stream.flv?token=$NVR_TOKEN" egress-flv.csv "$DUR" >/dev/null 2>&1 &
done
for id in $IDS; do
nohup python3 hlspull.py "$NVR_API/api/cameras/$id/stream/index.m3u8?token=$NVR_TOKEN" egress-hls.csv "$DUR" >/dev/null 2>&1 &
done
for id in $(echo "$IDS" | tr ' ' '\n' | head -8); do # RTSP spot-check on the first 8
ffmpeg -loglevel error -rtsp_transport tcp -i "$RTSP_BASE/$id" -t 10 -c copy -f null - >> egress-rtsp.log 2>&1 &
done
wait
echo egress-done
Idle footprint and cold start get their own measurement: start a zero-camera instance with a minimal config, time from process start until health responds, several rounds in a row with the data directory wiped between rounds:
#!/bin/bash
# startup.sh — cold start to API-ready latency, several consecutive trials
# usage: BIN=./mibee-nvr READY_URL=http://127.0.0.1:9199/api/health bash startup.sh
BIN=${BIN:?}; READY_URL=${READY_URL:?}
for i in $(seq 1 ${TRIALS:-10}); do
rm -rf data db.sqlite*
S=$(date +%s%N)
"$BIN" -config config.yaml > nvr.log 2>&1 &
PID=$!
for t in $(seq 1 600); do
curl -s -m 1 -o /dev/null "$READY_URL" && break
sleep 0.05
done
E=$(date +%s%N)
echo "trial $i: $(( (E - S) / 1000000 )) ms"
kill $PID 2>/dev/null; wait $PID 2>/dev/null || true
done
Post-processing has little to explain: the raw data is 5-second CSV/JSONL plus event timestamps in marks.txt; window statistics are percentile queries over timestamp slices, and the charts are matplotlib. Those scripts couple tightly to the data formats so I'll skip them — the idea is exactly the two clauses in this paragraph. Cleanup after a run is a pusher stop plus deleting cameras by name prefix — ordinary API calls, not worth the space.



Top comments (0)