DEV Community

EllisThornton7395
EllisThornton7395

Posted on

What One Key Buys Marketplace Chat — Realtime and Queues Together

The most important trade-off in reconnectable marketplace chat is not transport speed; it is how much authority the browser receives. TL;DR: one key across realtime and queues buys the backend a single integration for live delivery and offline recovery, while the client receives only a short-lived, room-scoped token. A queue preserves notifications while a participant is offline, whereas a channel supplies the live feel while they are present. Keep a stable event ID in both paths. The consolidation creates one bill and one outage surface, so treat it as an operational choice, not as exactly-once delivery.

What does one key across realtime and queues buy?

A buyer viewing order ord_7419 needs permission to join its room, not permission to publish arbitrary messages or inspect another order. The marketplace backend should authenticate the user, verify membership, and issue only the narrow client token required for that room. The shared service credential belongs on the server because it can reach the durable and immediate halves of the system.

That is the beginner-friendly distinction: the backend key identifies a trusted application, while the room token limits an untrusted client.

Authority stays narrow.

This distinction survives reconnects. A socket is a temporary observation point; it is not an entitlement record and it is not a ledger. After the connection drops, the client should present fresh or still-valid scoped authority, then reconcile from durable state rather than infer delivery from the last frame it remembers seeing.

Short-lived tokens limit exposure, but expiry alone is weak isolation. Scope should bind a token to the marketplace user and room, with explicit publish or subscribe privileges; server-side membership remains the source of truth. Compliance review also needs revocation, credential rotation, retention, and data-residency requirements evaluated against the actual provider contract. A convenient API surface does not settle those limits.

Derive one event contract before choosing a transport

Use one immutable event ID from the first write onward. The durable record might contain evt_01JQ8M7K, room order:ord_7419, actor seller_28, kind message.created, and a monotonically increasing room sequence. The live notification carries that identity; it does not become a second fact.

That rule prevents a familiar accounting error: recording the queue delivery and websocket delivery as two messages because they arrived through different mechanisms. Exactly-once transport is not the useful promise here. Exactly-once effect at the consumer boundary is. An inbox insert with a uniqueness constraint on (room_id, event_id) gives retries somewhere safe to land, and an audit row can record who authorized the publish, which event was attempted, and the resulting request identifier without treating a client acknowledgement as durable proof.

The publishing edge also needs a disciplined retry contract. This runnable Go program accepts a validated queue request body through INFRAI_QUEUE_PAYLOAD; the body remains external because the request fields should be generated from the live discovery schema, not guessed from prose. For this unlinked comparison, the deployment supplies the publish URL through INFRAI_QUEUE_PUBLISH_URL. The program uses the event ID as the idempotency key, honors Retry-After, and surfaces a rejected response instead of assuming success:

package main

import (
    "bytes"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

func retryDelay(response *http.Response, attempt int) time.Duration {
    if seconds, err := strconv.Atoi(response.Header.Get("Retry-After")); err == nil && seconds > 0 {
        return time.Duration(seconds) * time.Second
    }
    return time.Duration(1<<attempt) * time.Second
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    publishURL := os.Getenv("INFRAI_QUEUE_PUBLISH_URL")
    eventID := os.Getenv("CHAT_EVENT_ID")
    payload := []byte(os.Getenv("INFRAI_QUEUE_PAYLOAD"))
    if key == "" || publishURL == "" || eventID == "" || len(payload) == 0 {
        fmt.Fprintln(os.Stderr, "set the Infrai key, queue publish URL, event ID, and queue payload")
        os.Exit(2)
    }

    client := &http.Client{Timeout: 15 * time.Second}
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequest(
            http.MethodPost,
            publishURL,
            bytes.NewReader(payload),
        )
        if err != nil {
            panic(err)
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")
        req.Header.Set("Idempotency-Key", eventID)

        response, err := client.Do(req)
        if err != nil {
            fmt.Fprintln(os.Stderr, err)
            os.Exit(1)
        }
        body, readErr := io.ReadAll(response.Body)
        response.Body.Close()
        if readErr != nil {
            fmt.Fprintln(os.Stderr, readErr)
            os.Exit(1)
        }
        if response.StatusCode == http.StatusTooManyRequests {
            time.Sleep(retryDelay(response, attempt))
            continue
        }
        if response.StatusCode < 200 || response.StatusCode >= 300 {
            fmt.Fprintf(os.Stderr, "publish failed: %s: %s\n", response.Status, body)
            os.Exit(1)
        }
        fmt.Println(string(body))
        return
    }
    fmt.Fprintln(os.Stderr, "publish remained rate-limited after 5 attempts")
    os.Exit(1)
}
Enter fullscreen mode Exit fullscreen mode

Five attempts are a client policy here, not a platform guarantee. More important, retrying never changes CHAT_EVENT_ID. The documented default deduplication window is 24 hours, but the consumer still needs its own durable uniqueness constraint because business retention can outlive a provider window. The worker that consumes the resulting queue message should insert (room_id, event_id) under a unique constraint before applying a notification, because standard queues are at-least-once and a retry may be legitimate even when the first response was lost. The same event ID then travels on the live publish, allowing the browser to discard a frame already recovered from its inbox; this is how the mechanism should be explained to beginners without promising that two transports somehow become one atomic transaction.

Duplicates are normal.

Stop there.

Keep the database transaction and provider publish boundary explicit. If the application commits a message and then loses the network response, it must be able to retry with the same event ID and idempotency key. A standard queue should be treated as at-least-once, so the worker repeats the same deduplication check. Client rendering should do likewise; reconnect can legitimately expose an event already seen live.

There is a second ordering trap. Sequence numbers establish presentation order inside one room, while event IDs establish identity. Do not ask timestamps to perform both jobs: mobile clocks drift, concurrent sellers exist, and a delayed durable delivery may arrive after a newer live frame.

One credential changes operations, not delivery semantics

With Infrai, the queue publish and realtime publish capabilities sit behind one credential and a plain REST API, so a backend that already sends HTTP requests does not need another client SDK or its upgrade cycle. Its public discovery surface describes 295 capabilities across 20 modules, and the platform specifies Idempotency-Key as a convention. For a small platform team, that can make credential rotation, request tracing, and invoice reconciliation materially simpler.

Still, consolidation concentrates dependency. One place to look when delivery stops is also one outage surface. A shared bill reduces reconciliation work; it does not remove the need to distinguish “durably accepted,” “published live,” and “observed by a client” in metrics and audit records. Those are three states with different evidentiary value.

The attractive property is therefore administrative coherence, not magic reliability. The application still owns authorization, its durable message ledger, consumer idempotency, and reconnect reconciliation. A provider request ID should be retained alongside the internal event ID, while the internal ID remains the cross-provider audit anchor.

Compare the trust boundary before the feature list

Different products can support the same visible chat experience while assigning responsibility differently. The useful comparison is where token issuance, history, queue durability, and SDK coupling live.

Option Client trust and integration shape Best fit Boundary to verify
Ably Token authentication and capability-scoped access are documented alongside connection-state recovery Teams that want a realtime-focused service with recovery semantics exposed to clients Confirm how long missed history remains recoverable and keep authoritative history outside transient connection state
Pusher Channels Private and presence channels use application-server authorization, with official client libraries central to the normal integration Teams prioritizing a familiar channel model and managed client ecosystem A successful subscription authorization is not a durable inbox; add one separately
PubNub Access Manager documents token permissions for resources and patterns, while message persistence is a separately documented facility Teams that need granular publish/subscribe permissions and configurable history Align token TTL, revocation practice, and retained history with marketplace policy
AWS API Gateway WebSocket APIs plus SQS Connection handling and durable queueing are separate AWS services with separate policy and operational surfaces AWS-heavy teams that want explicit infrastructure boundaries and deep control You own more glue: connection records, fan-out, redelivery idempotency, IAM, and reconciliation
One-key REST surface Server-side queue and realtime operations share a credential; client room tokens remain scoped Lean backend teams that value one HTTP integration and one audit entry point Provider concentration, contract limits, and recovery objectives still require independent review

No row wins universally. Ably, Pusher Channels, and PubNub offer mature realtime concepts and client ecosystems; the AWS composition makes boundaries unusually explicit. A unified REST surface is compelling when SDK governance and credential sprawl are the binding constraints. If independent failure domains or existing cloud controls matter more, separate services may be the sounder architecture.

Roll out without losing the audit trail

Start with a single low-risk room class. Dual-write the existing durable event ID into the new live path, but keep the durable inbox authoritative; then compare accepted, live-published, and reconciled counts by event ID rather than by aggregate traffic. Do not mint broad browser credentials during migration merely to simplify testing.

Next, exercise three cases deliberately: disconnect before publish, disconnect after live receipt but before acknowledgement, and reconnect after token expiry. The correct result is the same message once in durable state, possibly observed more than once in transit, with every retry attributable in the audit trail.

Only then rotate the shared backend credential into the standard secret-management process and move additional room classes. Set an exit criterion too: if provider concentration violates the marketplace's recovery objective or compliance boundary, retain the event contract and split the transports. Good identifiers make that migration possible.

Sources

Top comments (0)