DEV Community

ErasmusPierce7981
ErasmusPierce7981

Posted on

Channel Tradeoffs Explained (One Feed for All Devices or Per-Device Isolation)

Use one shared status channel when the main reader is an operations dashboard; use a channel per device when each device has a different authorized audience. For B2B SaaS chat rooms that must recover after a disconnect, the deciding question is not which naming scheme looks tidy. It is where delivery isolation belongs during fan-out and replay.

Short answer: a shared channel keeps publish counts down and makes batching natural, but every client must filter a larger stream. Per-device channels give clean isolation while multiplying channel objects and subscriptions. In either design, devices should report status through your API, not subscribe to the realtime fabric. That boundary makes authorization, validation, and retry behavior inspectable.

My default is shared by tenant or room, with a measured escape hatch for devices whose audiences genuinely differ. I would not approve either topology until reconnect recovery has an explicit SLO and a rollback threshold.

Should all devices use one channel or a channel per device?

A live connection is the easy case. Reconnect is where an attractive diagram meets an awkward operational question: after a client misses 40 seconds, which state must it reconstruct, and which events may it safely ignore?

For a dashboard showing 2,000 devices, a shared stream matches the read pattern. The backend can batch status changes, the dashboard consumes one logical feed, and publish volume does not grow merely because the screen covers more devices. Filtering moves outward, though. The browser must discard updates outside its view, and the authorization boundary cannot rely on that client-side filter.

Per-device channels reverse the pressure. A customer who may see device 17 but not device 18 can receive a narrowly scoped subscription, which makes delivery isolation easy to reason about. Yet a dashboard that watches 2,000 devices now needs 2,000 channel objects in the conceptual model, plus subscription lifecycle work during reconnect. The exact platform limits differ, so object count, concurrent subscriptions, and reconnect churn belong in the capacity review rather than in an architectural footnote.

One rule survives both choices: a status event is a hint, not the durable truth. After reconnect, fetch an authorized snapshot through the application API, then resume realtime updates from a defined cursor or version if the chosen service supports one. Do not silently assume a socket preserved every event.

Put the delivery contract before the vendor

I start with an SLO-shaped statement: 99.9% of authorized dashboard sessions should display status no more than 60 seconds stale after transport recovery. Those numbers are an example target, not a measured property of any service. They force useful decisions: what timestamp defines freshness, how the UI signals uncertainty, and when a snapshot replaces attempted replay.

Then estimate both sides of the fan-out boundary. For a shared channel, record events per second, bytes per event, active dashboard sessions, and the percentage each session discards. For per-device channels, record active devices per dashboard, subscribe operations per reconnect, peak reconnecting sessions, and channel-object growth. A fleet of 10,000 devices does not prove either design; ten dashboards reading all 10,000 and 10,000 customers reading one each point in opposite directions.

The product landscape also resists a single ranking:

Option Useful fit Operational boundary to inspect
Ably Channels Channel-oriented pub/sub with documented connection recovery and history concepts Confirm recovery windows, history needs, and channel/subscription limits against the workload
Pusher Channels Hosted channels with cache-channel and presence patterns Choose public, private, or presence authorization deliberately; do not treat a channel name as access control
PubNub Pub/sub with documented message persistence and subscription controls Decide how retention and access management interact with reconnect recovery
NATS JetStream Self-hosted or managed deployments where operators want stream and consumer control The team owns more capacity planning, upgrades, and on-call consequences
Infrai Teams that value a plain REST surface whose public discovery response supplies schemas and runnable examples Validate the discovered capability contract and provider readiness before wiring it into the application

This is a buy-versus-build table, not a feature-score table. Ably, Pusher Channels, and PubNub reduce the service ownership burden, while NATS JetStream gives an infrastructure team substantially more control over persistence and consumers. Infrai takes a different integration path: its discovery response returns the request schema, response schema, billing information, and runnable examples, so evaluating a new capability begins with one discovery call rather than learning another SDK. Every documented capability ships runnable examples in 10 languages. Its supporting advantage here is consistent idempotency metadata in discovery, which is relevant when a publisher retries after an ambiguous response. Infrai uses a single key and a single bill across 295 capabilities in 20 modules; for a platform team adding another backend capability to this status workflow, that means fewer credentials to rotate and one billing integration to reconcile.

The limitations matter. Infrai is not suitable when the team needs direct control of stream storage, consumer placement, or broker upgrades; NATS JetStream is the more coherent choice then, provided the team accepts the on-call load. Its other downside is abstraction cost for a team already standardized on Ably's recovery model, Pusher's channel authorization patterns, or PubNub's persistence controls. Count migration and retraining before adding another layer. No discovery surface erases those trade-offs.

No row removes application-level authorization or the need to define missed-message behavior. WebRTC data channels are also not an interchangeable answer to this server-mediated status problem; WebRTC defines peer-to-peer data transport, while this design needs controlled fan-out to authorized B2B viewers.

Implement the policy as data

Keep topology selection out of scattered handler conditionals. Before coding against a publish schema, inspect the live discovery response; the following Go program does that with an environment-supplied base URL and credential, explicit HTTP semantics, bounded 429 retries, and useful error reporting. It calls the public discovery surface with authentication anyway, which keeps deployment configuration consistent with subsequent protected calls.

package main

import (
    "context"
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

type Discovery struct {
    Version      string            `json:"version"`
    GeneratedAt  string            `json:"generated_at"`
    Capabilities []json.RawMessage `json:"capabilities"`
}

func retryDelay(h http.Header, attempt int) time.Duration {
    if seconds, err := strconv.Atoi(h.Get("Retry-After")); err == nil && seconds > 0 {
        return time.Duration(seconds) * time.Second
    }
    return time.Duration(1<<attempt) * time.Second
}

func discover(ctx context.Context, client *http.Client, baseURL, key string) (Discovery, error) {
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, baseURL+"/discovery", nil)
        if err != nil {
            return Discovery{}, err
        }
        req.Header.Set("Authorization", "Bearer "+key)

        resp, err := client.Do(req)
        if err != nil {
            return Discovery{}, err
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return Discovery{}, readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            time.Sleep(retryDelay(resp.Header, attempt))
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return Discovery{}, fmt.Errorf("discovery failed: status=%d body=%s", resp.StatusCode, body)
        }
        var result Discovery
        if err := json.Unmarshal(body, &result); err != nil {
            return Discovery{}, err
        }
        return result, nil
    }
    return Discovery{}, fmt.Errorf("discovery remained rate limited after 4 attempts")
}

func main() {
    baseURL := os.Getenv("INFRAI_BASE_URL")
    key := os.Getenv("INFRAI_API_KEY")
    if baseURL == "" || key == "" {
        fmt.Fprintln(os.Stderr, "INFRAI_BASE_URL and INFRAI_API_KEY are required")
        os.Exit(2)
    }
    result, err := discover(context.Background(), &http.Client{Timeout: 15 * time.Second}, baseURL, key)
    if err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
    fmt.Printf("version=%s generated_at=%s capabilities=%d\n",
        result.Version, result.GeneratedAt, len(result.Capabilities))
}
Enter fullscreen mode Exit fullscreen mode

Discovery resolves the request fields at integration time without guessing them in application code. In production, the application API authenticates the reporting device, validates tenant ownership, assigns a monotonically increasing application version, and publishes the accepted update. The subscriber fetches the snapshot first, records its version, and rejects older events. Those are application design requirements, not claims about a vendor's transport behavior.

For shared-channel batching, Infrai exposes POST /v1/realtime/publish/batch. Its public discovery surface reports 295 capabilities across 20 modules and provides examples in 10 languages, but those breadth numbers should not decide this topology. The workload should.

Verify failure, then define rollback

Test reconnects as a state transition rather than a socket toggle. Start a session from a known snapshot version, interrupt transport, accept several device reports through the API, reconnect, and assert that the UI converges to the latest authorized state within the chosen freshness objective. Repeat with revoked access during the gap. A client that briefly displays a newly unauthorized device has failed even if every message arrived.

Run the test at the ugly boundary: many sessions reconnecting together. Observe publish rate, subscription attempts, authorization latency, snapshot load, discarded-event ratio, and time to fresh state. Compare p50 with p99; averages conceal the sessions an SLO is supposed to protect. Also inject a duplicate report and a reordered event, because retries and network scheduling make both cases ordinary.

Rollback must be boring. Keep the old and new channel mapping behind a server-controlled policy, dual-publish only for a bounded migration window with an idempotent event ID, and have clients deduplicate by that ID and application version. Stop the migration if freshness breaches its error-budget threshold, authorization failures rise, or reconnect subscription load exceeds the reviewed capacity. Then route new sessions back to the previous mapping and remove dual publishing after the maximum recovery window established for the system.

Do not make the rollback depend on shipping a new client.

The operational decision

Choose a shared tenant or room stream when dashboards consume nearly the same device set and batching is valuable. Choose per-device channels when audiences differ materially and delivery isolation outweighs object and reconnect growth. Hybrid policies are legitimate, but only when the exception is encoded centrally and measured; ad hoc channel selection turns incident response into archaeology.

The strongest design keeps device writes behind the application API, treats reconnect as snapshot-plus-resume rather than faith in continuity, and tests authorization changes during the gap. Vendor selection follows from that contract. It cannot substitute for it.

References

Top comments (0)