DEV Community

mattleeee
mattleeee

Posted on Originally published at hkcode.dpdns.org

GB28181 Camera Cascading: When SIP INVITE Never Arrived

Heartbeats flowed, the platform showed the device online, but pressing play produced nothing. The SIP INVITE that should have kicked off the media session never reached the NCG. This is a walkthrough of how we localized that failure and what evidence we handed to the firewall engineer.

The Symptom

A school's cameras cascade to a provincial education video platform through a Hikvision NCG (network control gateway). The platform client showed the device as online. Playback requests returned no video. No error dialog, no timeout message on the client side — just silence.

The first instinct is to blame the camera or the NCG. The second instinct is to blame the platform. Both are usually wrong, and in this case both were wrong.

GB28181 Call Flow: Two Different Worlds

GB28181 sits on top of SIP, but the reliability characteristics of its two main message classes are very different. Understanding this split is what makes the diagnosis tractable.

Heartbeat: REGISTER and MESSAGE

The device keeps itself alive with periodic REGISTER and MESSAGE (keepalive) packets. These are small, sent at fixed intervals (typically 60s), and mostly stateful on the device side. If one is lost, the next one arrives within a minute and the platform never notices.

Device                          Platform
  |--- REGISTER (Expires: 3600) --->|
  |<-- 200 OK ----------------------|
  |--- MESSAGE (Keepalive) -------->|
  |<-- 200 OK ----------------------|
  |            ...every 60s...
Enter fullscreen mode Exit fullscreen mode

Because these packets are small (a few hundred bytes) and periodic, they traverse NAT and firewall state tables easily. A NAT entry created by an outbound REGISTER stays alive because the keepalive refreshes it. From the platform's perspective, the device is online. From the network's perspective, a bidirectional path exists.

Media Session: INVITE / ACK / BYE

A playback request is a different animal. The platform sends an INVITE to the device, the device answers with 200 OK (or 486 Busy etc.), the platform sends ACK, and then RTP flows. The INVITE is larger — the SDP body carries media negotiation — and it is unsolicited from the device's point of view. It arrives from the platform, not in response to anything the device just sent.

Platform                        Device (NCG)
  |--- INVITE (SDP, ~950B) ------>|
  |<-- 100 Trying -----------------|
  |<-- 200 OK (SDP) ---------------|
  |--- ACK ----------------------->|
  |<========== RTP ==============>|
  |--- BYE ----------------------->|
Enter fullscreen mode Exit fullscreen mode

Here is the trap. If the device sits behind NAT and the firewall has no pinhole for inbound SIP, the INVITE is dropped before it ever reaches the NCG. The device has no idea a call was attempted. The platform sees its INVITE vanish. The client shows a spinner and then nothing.

Heartbeats work because the device initiates them. INVITEs fail because the platform initiates them. Same SIP port, same session, completely different NAT direction.

Placing tcpdump Correctly

The most common mistake is running tcpdump on the wrong host. If you capture on the camera, you will see heartbeats and nothing else — because the INVITE never got that far. If you capture on the platform, you will see the platform sending INVITEs and no responses. Neither capture tells you where the packet died.

Capture on the NCG server, the box that terminates the SIP signaling and sits between the school LAN and the WAN. That is the boundary. Anything that arrives there is inside the school. Anything that does not arrive there but was seen on the platform is stuck in the firewall or an upstream device.

# On the NCG server, capture SIP signaling both directions.
# -i any so we don't guess the interface.
# -s 0 so the SDP body is not truncated.
# -w to a file so we can analyse offline.
sudo tcpdump -i any -s 0 -w /tmp/sip_capture.pcap \
    'port 5060 or port 5061'
Enter fullscreen mode Exit fullscreen mode

Run it for at least 3.5 minutes. Why 3.5? Because the platform's retry logic for INVITE typically fires at 500ms, 1s, 2s, 4s, 8s, 16s, 32s, 64s... and the client may retry the whole call. Three and a half minutes gives you two full retry ladders and enough heartbeat cycles to confirm the baseline is healthy. In our case, 16 retry bursts of roughly 950 bytes each were visible on the platform side. Zero arrived at the NCG.

Reading the Capture

The capture told a clean story. Heartbeats in both directions, REGISTER/200 OK pairs every 60 seconds, no gaps. Then, at the moment the operator pressed play on the client, nothing on the NCG side. No INVITE. No 100 Trying. No 950-byte burst.

# Count SIP methods seen at the NCG.
tshark -r /tmp/sip_capture.pcap -Y sip \
    -T fields -e sip.Method | sort | uniq -c

# Expected output on a healthy NCG:
#   14 REGISTER
#   14 200
#   2  MESSAGE
#   2  200
#   0  INVITE   <-- this is the smoking gun
Enter fullscreen mode Exit fullscreen mode

A healthy capture would show at least one INVITE per playback attempt, plus the corresponding 100 Trying and 200 OK. Zero INVITEs at the NCG, while the platform confirms it sent them, means the packet was dropped upstream. The failure is not in the camera, not in the NCG, not in the platform. It is in the path.

Building the Evidence Pack

Firewall engineers are busy. Do not send them a pcap and a paragraph. Send them a table.

1. The Capture

One pcap from the NCG, one from the platform side if you can get it. Name them with timestamps and hostnames: ncg01_20240315_1420.pcap, platform_edge_20240315_1420.pcap. The platform-side capture is what proves the INVITE left the platform — without it, the firewall engineer can argue the platform never sent anything.

2. The Five-Tuple Table

Extract the exact tuples from both captures and diff them.

import subprocess
import csv
from collections import Counter

def extract_sip_tuples(pcap_path):
    """Pull src/dst IP, ports, and SIP method from a pcap."""
    cmd = [
        "tshark", "-r", pcap_path,
        "-Y", "sip",
        "-T", "fields",
        "-e", "ip.src", "-e", "ip.dst",
        "-e", "udp.srcport", "-e", "udp.dstport",
        "-e", "sip.Method",
        "-e", "frame.time_epoch",
        "-E", "separator=|",
    ]
    out = subprocess.run(cmd, capture_output=True, text=True, check=True)
    rows = []
    for line in out.stdout.strip().splitlines():
        parts = line.split("|")
        if len(parts) >= 6:
            rows.append({
                "src": parts[0], "dst": parts[1],
                "sport": parts[2], "dport": parts[3],
                "method": parts[4], "ts": float(parts[5]),
            })
    return rows

ncg = extract_sip_tuples("/tmp/ncg01.pcap")
plat = extract_sip_tuples("/tmp/platform_edge.pcap")

def tuple_key(r):
    return (r["src"], r["dst"], r["sport"], r["dport"])

ncg_keys = Counter(tuple_key(r) for r in ncg)
plat_keys = Counter(tuple_key(r) for r in plat)

print("Tuples seen on platform but NOT on NCG:")
for k, count in plat_keys.items():
    if k not in ncg_keys:
        print(f"  {k}  x{count}")
Enter fullscreen mode Exit fullscreen mode

For our case, the output was a single line: the platform's SIP signaling IP and port, destined for the NCG's public IP and port 5060, with roughly 16 occurrences. That is the exact tuple the firewall needs to permit.

3. The Timeline

A short timeline matters more than raw packet dumps. Something like:

14:22:01  REGISTER from NCG -> platform, 200 OK. Baseline healthy.
14:25:33  Operator presses play on platform client.
14:25:33  Platform sends INVITE to NCG public IP:5060 (seen on platform capture).
14:25:33  No INVITE arrives at NCG. No 100 Trying, no 200 OK.
14:25:34  Platform retries INVITE. Same result.
14:25:36  Platform retries. Same result.
...
14:26:37  Final retry. Operator gives up.
Enter fullscreen mode Exit fullscreen mode

This tells the firewall engineer exactly what to look for in their logs: dropped packets on the WAN interface, destination port 5060, source IP of the platform, around 14:25:33.

The Actual Fix

In our case, the school firewall had a stateful rule permitting outbound SIP from the NCG to the platform, but no corresponding inbound rule for the platform's SIP signaling IP. The outbound REGISTER created a NAT pinhole, but the pinhole was UDP and tied to the source port the NCG used. The platform's INVITE came from a different source port on its side, so it did not match the pinhole and was dropped.

The fix was a static inbound rule: permit UDP from the platform's SIP signaling IP (a small CIDR, not a single host — platforms often have a signaling cluster) to the NCG's public IP on port 5060. RTP ports 30000-30500 were opened in the same rule set. After that, INVITEs arrived, 200 OKs went back, and playback worked.

A note on why this is easy to miss: the platform showed the device as online because heartbeats were flowing. That status is derived from REGISTER state, not from INVITE reachability. Online in the platform UI does not mean the call path is open. Treat them as two separate reachability questions.

A Reusable Check

Before you escalate to the firewall team, run this on the NCG to confirm the split between outbound-initiated and inbound-initiated traffic:

import subprocess

def sip_method_counts(pcap_path):
    """Return a dict of SIP method -> count."""
    cmd = [
        "tshark", "-r", pcap_path,
        "-Y", "sip",
        "-T", "fields", "-e", "sip.Method",
    ]
    out = subprocess.run(cmd, capture_output=True, text=True, check=True)
    counts = {}
    for m in out.stdout.strip().splitlines():
        counts[m] = counts.get(m, 0) + 1
    return counts

counts = sip_method_counts("/tmp/ncg01.pcap")
print(counts)

if counts.get("REGISTER", 0) > 0 and counts.get("INVITE", 0) == 0:
    print("Outbound path OK, inbound path blocked. Escalate to firewall.")
elif counts.get("INVITE", 0) > 0 and counts.get("200", 0) == 0:
    print("INVITE arrives but device does not answer. Check NCG/camera.")
else:
    print("Both paths appear active. Look at RTP or SDP.")
Enter fullscreen mode Exit fullscreen mode

The heuristic is simple: if REGISTER is present but INVITE is absent, the problem is directional. Outbound works, inbound does not. That is almost always a firewall or NAT rule, not a device fault.

Takeaways

  • Heartbeats prove outbound reachability, not inbound. They are independent questions.
  • Capture at the boundary (the NCG), not at the endpoints, to localize where the packet dies.
  • Run the capture long enough to cover the platform's retry ladder — 3.5 minutes is a safe floor.
  • Hand the firewall engineer a five-tuple table and a timeline, not a pcap and a hunch.
  • The fix is usually a static inbound rule for the platform's SIP signaling IP and RTP port range.

More notes like this ship every week on this site.


Daily Picks

The following pairs are selected from the multi-timeframe trend scanner (Gate.io futures) and are for technical-analysis study only — not investment advice.
Data updated: 2026-10-08 08:46:21

Short

Pair Signal Price Take Profit Stop Loss R/R
SUE 弱做空 $0.0006 $0.0006 $0.0006 1:1.2

1 picks selected. Scanner runs every 15 minutes.

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.