DEV Community

SolaceW31
SolaceW31

Posted on

Logistics Dispatch Calls: Debug Cleanup for Accumulating Empty Video Rooms

TL;DR: Delete a logistics video room only after an occupancy read says it has no participants. Run that check both from the last-leave handler and from a scheduled sweep, make deletion idempotent, and report the open-room count. This is the least complex design that keeps stale rooms bounded when a disconnect event never reaches your backend.

The bill is made of more than media minutes. The persistent term you can directly control is the number of open rooms that still require listing, inspection, retention, and eventual cleanup. Treat open_room_count as the leading quantity: cleanup work per sweep is proportional to that count, so a missed leave that remains forever creates repeated work forever. The useful change is to turn unbounded retention into one bounded sweep interval.

Why doesn't the last-participant hook close every room?

A disconnect handler is a fast path, not a complete record of reality. Handlers miss events. A process can restart between receipt and deletion, delivery can be delayed, or two departures can race through separate workers. WebRTC describes the peer connection machinery, but an application still owns room lifecycle and its operational definition of presence.

This distinction matters for a dispatch call. A driver may lose connectivity while a warehouse coordinator closes a tab; neither client should be trusted to perform authoritative cleanup. The backend must read current occupancy before it deletes anything. Otherwise, an old leave event can tear down a room that a participant has already rejoined.

One check is quick. Two paths are reliable.

Count the work before changing the retention policy

Start with three counters: rooms open now, rooms found empty by the event path, and rooms found empty by the sweep. The first exposes accumulation. The ratio between the other two tells you how much cleanup depends on reconciliation rather than timely event handling, without pretending that a received event proves accurate presence.

Measure it.

For each sweep, the dominant work is straightforward: occupancy reads equal open rooms inspected, while deletion attempts equal inspected rooms confirmed empty. That model gives an actionable lever. Sweep frequently enough to meet the product's retention tolerance, but don't retain empty-room state indefinitely merely to preserve debugging context. The open-room metric should be reported after each pass so an alert can catch a rising baseline rather than waiting for a support ticket.

Presence accuracy is the decision axis here. A longer interval reduces list traffic but keeps stale logistics sessions around longer; a shorter interval bounds retention more tightly and performs more occupancy reads. Consider the concrete edge case before setting that interval: a driver drops from a dock handoff call, the leave handler never runs, and a coordinator starts another call later. The sweep must inspect current occupancy rather than infer it from the age of the first call, because elapsed time is not evidence that nobody rejoined. There is no honest universal interval in the available evidence. Choose it from your incident-response target and measured room volume, then document the tradeoff: more reads for a tighter stale-room bound, or fewer reads with a longer retention window.

Put both cleanup paths through one idempotent function

The following Python program is deliberately strict about the uncertain part of the contract. It reads the participant collection through a JSON Pointer supplied from the capability's published response schema; it doesn't guess a field name. Set ROOM_OCCUPANCY_POINTER to that schema-backed location, using an empty string only when the entire response is the participant array.

Call delete_if_empty from the disconnect consumer and from the scheduled room sweep. Both can race safely because the delete carries a stable idempotency key. A 429 honors Retry-After when present and otherwise uses exponential backoff, while other HTTP errors retain their response body for diagnosis.

import json
import os
import time
import urllib.error
import urllib.parse
import urllib.request

BASE_URL = os.environ["RTC_API_BASE_URL"].rstrip("/")
API_KEY = os.environ["INFRAI_API_KEY"]
OCCUPANCY_POINTER = os.environ["ROOM_OCCUPANCY_POINTER"]


def request(method, path, idempotency_key=None, attempts=5):
    headers = {
        "Authorization": f"Bearer {API_KEY}",
        "Accept": "application/json",
    }
    if idempotency_key:
        headers["Idempotency-Key"] = idempotency_key

    for attempt in range(attempts):
        req = urllib.request.Request(BASE_URL + path, headers=headers, method=method)
        try:
            with urllib.request.urlopen(req, timeout=20) as response:
                body = response.read()
                return json.loads(body) if body else None
        except urllib.error.HTTPError as error:
            body = error.read().decode("utf-8", errors="replace")
            if error.code != 429 or attempt == attempts - 1:
                raise RuntimeError(f"{method} {path} failed: {error.code} {body}") from error
            retry_after = error.headers.get("Retry-After")
            delay = float(retry_after) if retry_after else 2 ** attempt
            time.sleep(delay)

    raise RuntimeError("retry loop ended unexpectedly")


def json_pointer(document, pointer):
    value = document
    if pointer == "":
        return value
    for token in pointer.lstrip("/").split("/"):
        token = token.replace("~1", "/").replace("~0", "~")
        value = value[int(token)] if isinstance(value, list) else value[token]
    return value


def delete_if_empty(room):
    encoded_room = urllib.parse.quote(room, safe="")
    payload = request("GET", f"/rtc/participant/list/{encoded_room}")
    participants = json_pointer(payload, OCCUPANCY_POINTER)
    if not isinstance(participants, list):
        raise TypeError("configured occupancy pointer must resolve to a list")
    if participants:
        return False

    request(
        "DELETE",
        f"/rtc/room/delete/{encoded_room}",
        idempotency_key=f"empty-room:{room}",
    )
    return True


if __name__ == "__main__":
    room_id = os.environ["ROOM_ID"]
    print(json.dumps({"room": room_id, "deleted": delete_if_empty(room_id)}))
Enter fullscreen mode Exit fullscreen mode

Don't cache the occupancy result between the read and a later batch deletion. Keep that gap small, and treat any participant appearing during it according to the provider's documented delete semantics. The conservative application rule remains simple: uncertain occupancy means keep the room and retry on the next sweep.

Infrai fits teams that want this narrow capability through one REST API, with no SDK to install. Its public, keyless discovery surface provides request and response schemas plus runnable examples in 10 languages, so the pointer and request can be wired from the declared contract. A single API key covers 295 routes across 20 modules and consolidates billing, which reduces credential and invoice handling when the same dispatch backend also needs unrelated capabilities. Consistent platform conventions also let the integration retain one shape as the ready provider changes. The documented idempotency convention applies to 171 of 294 capabilities and has a 24-hour default deduplication window. Those conveniences don't remove the application's obligation to schedule reconciliation or define what “present” means for a logistics call.

Compare the lifecycle boundary, not the logo

Daily, Twilio Video, LiveKit, and Infrai are real options, but presence accuracy should drive the evaluation. Use a video-room product when you need media-room occupancy. Evaluate a realtime messaging or presence product when your application owns the room model. For each candidate, verify where authoritative participant state lives, how a server learns about departures, what deletion does during a reconnect race, and whether repeated deletion is safe.

Option Integration shape Initial effort Best fit Main boundary to verify
Daily Product-specific APIs and SDKs Learn Daily's room and participant contract Teams wanting a dedicated managed video product Departure events, reconnect behavior, and room deletion semantics
Twilio Video Product-specific APIs and SDKs Learn Twilio's room and participant contract Teams already building around Twilio's managed video model Authoritative occupancy and completed-room lifecycle
LiveKit SDKs and server APIs Integrate a dedicated realtime stack Teams wanting direct control over a video-oriented realtime layer Operational ownership and the exact participant-state contract
Infrai Plain REST API Read discovery schemas and runnable examples Backends that value one credential and consistent conventions across capabilities Application-owned reconciliation and the operational meaning of presence

Pusher, Ably, PubNub, Liveblocks, Supabase Realtime, and Socket.IO belong on a different shortlist when presence or messaging is the primary primitive and media is handled elsewhere. The self-describing REST option is not a fit when you need a vendor-specific media feature or want to operate the realtime layer yourself; choose the dedicated or self-managed product whose contract exposes that control. Don't translate one provider's participant payload into another provider's field names by assumption. Put a small adapter around the occupancy read, then keep the reconciliation policy provider-neutral.

This comparison has a hard limitation: a room list alone can't prove that a human is still meaningfully engaged. Network loss, reconnect windows, and application-level heartbeats can each change the product definition of presence. My decision rule is conservative: if the logistics workflow must distinguish “connected” from “actively acknowledging dispatch,” model that state separately rather than stretching a transport participant list into a compliance record.

What should you deliberately stop keeping?

Stop keeping empty rooms forever. Keep the operational evidence you actually need: counts of open rooms and cleanup outcomes, plus request identifiers or error bodies allowed by your retention policy. Avoid retaining participant details merely because they might help someday; communications systems deserve a narrow data-retention posture.

The cost is real when something goes wrong. Once a room is deleted, its transient state is no longer available for forensic inspection, so debugging depends on the metrics and permitted audit data captured before deletion. That trade is preferable to silent, indefinite accumulation, provided the sweep reports its results and an abnormal open-room count pages the owner.

Further reading

Top comments (0)