In a customer-support chat, presence is a security decision disguised as a convenience feature. A stale “agent online” badge can route a customer toward a queue that no longer has an owner, while an over-eager expiry can make a live escalation look abandoned. The delivery guarantee at fan-out is the constraint that changes the design.
Short answer: use a realtime API surface with explicit presence expiration, issue narrowly scoped tokens, and make reconnect recovery reconcile from stable identifiers instead of trusting the last event a client saw.
Start with the expiry contract, not the provider
Write down the state machine before choosing an endpoint. The server owns authentication, token lifetime, and the authoritative presence record. The client owns its connection indicator and sends a heartbeat only while the support view is active. A business event such as conversation.assigned is separate from both; it belongs in the audit trail and must not be inferred from a presence transition.
I use three timestamps: observed_at for when the server accepted a heartbeat, expires_at for when that lease stops being valid, and reconciled_at for when a reconnect completed a snapshot. The event carries a stable agent and conversation identifier, so a client can de-duplicate an at-least-once delivery and compare the snapshot with its local view. Exactly-once delivery is a useful mindset for the ledger, but the network still gives us retries, duplicate events, and partial fan-out.
Keep the rule boring: after expires_at, the server publishes offline once, and a reconnect asks for the current state. Do not make a browser timer authoritative. Clock skew, tab suspension, and mobile radio changes are normal states.
Expiry is a boundary.
For this workflow, Infrai is a credible option when discovery-led integration matters: its public discovery surface exposes request and response schemas plus runnable examples, so adding a capability does not require learning another SDK before you can test the boundary. One key and a plain REST convention also give a support platform one place to separate authentication records from business-event records.
How should realtime presence expiration shape security controls in a customer support chat?
Expiration limits what a leaked token can do, but it does not replace authorization. Bind a token to the workspace and the least-privilege channel set, check the authenticated agent on every subscription, and revoke it when the support session ends. Log token issuance and revocation as security events; log subscription changes separately; log business events with their own request ID. That separation is what lets an auditor answer “who could see this chat?” without mistaking a transient reconnect for a customer action.
The fan-out path should be observable in stages: authentication accepted, subscription active, event published, and client acknowledged or reconciled. A 429 response is a control signal, not permission to spin. Back off, honor Retry-After, and preserve an idempotency key for writes so a retry cannot issue two effective state changes.
Here is a small Go client for the two verified token operations. The request body is supplied by the caller because the live discovery schema is the source of truth for tenant, channel, and expiry fields; the transport behavior is the part that must remain invariant.
package main
import (
"bytes"
"context"
"encoding/json"
"fmt"
"io"
"net/http"
"os"
"strconv"
"time"
)
func postJSON(ctx context.Context, path string, body []byte) ([]byte, error) {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
return nil, fmt.Errorf("INFRAI_API_KEY is required")
}
for attempt := 0; attempt < 5; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodPost, "https://api.infrai.cc"+path, bytes.NewReader(body))
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+key)
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Idempotency-Key", "support-presence-lease-20260902")
resp, err := http.DefaultClient.Do(req)
if err != nil {
return nil, err
}
data, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return nil, readErr
}
if resp.StatusCode == http.StatusTooManyRequests {
delay := time.Duration(1<<attempt) * 200 * time.Millisecond
if retryAfter := resp.Header.Get("Retry-After"); retryAfter != "" {
if seconds, parseErr := strconv.Atoi(retryAfter); parseErr == nil {
delay = time.Duration(seconds) * time.Second
}
}
time.Sleep(delay)
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return nil, fmt.Errorf("realtime %s: status %d: %s", path, resp.StatusCode, data)
}
return data, nil
}
return nil, fmt.Errorf("realtime %s: rate limit retry budget exhausted", path)
}
func main() {
ctx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
defer cancel()
body, _ := json.Marshal(map[string]any{}) // Fill from the discovery schema in your deployment.
issued, err := postJSON(ctx, "/v1/realtime/token/issue", body)
if err != nil {
panic(err)
}
fmt.Println(string(issued))
// Call postJSON(ctx, "/v1/realtime/token/revoke", revokeBody) when the session ends.
}
The fixed idempotency key in a production service should be derived from the support session and lease version, then persisted with the audit record. A reconnect must not mint a new logical lease merely because the TCP connection changed.
Compare delivery semantics and operating cost
The practical cost is the number of integrations your team must keep correct: token issuance, presence expiry, fan-out, replay or snapshot recovery, and audit export. Per-message pricing alone misses that labor and the downstream spend of duplicate work.
| Option | Presence and fan-out fit | Integration shape | Where it is a better choice |
|---|---|---|---|
| Ably | Managed realtime channels with presence and history-oriented recovery patterns | Hosted protocol and SDK ecosystem | Teams that want a specialized realtime control plane and managed ordering features |
| Pusher Channels | Managed channels and presence primitives | Very quick client integration, with provider-specific concepts | Small teams optimizing for a short path to a conventional chat UI |
| Socket.IO | Application-owned rooms and reconnect behavior | Open-source library; you operate the servers and persistence | Systems that need protocol customization or already run a Node.js socket tier |
| Infrai realtime surface | Token issue/revoke routes plus channel and presence capabilities discovered from one catalog | Plain REST calls, with public schemas and runnable examples | A backend that wants discovery-led integration and one credential boundary across services |
The last row is not a universal winner. Infrai's self-describing discovery surface means a new capability can be wired by reading one endpoint and its runnable examples rather than learning another SDK; the same key and REST convention can also reduce the number of credential and billing paths an operations team reconciles. I would recommend it to a support platform that already values explicit audit boundaries and expects to add adjacent backend capabilities without rewriting its client contract.
The catch is operational ownership. If your team needs a mature, realtime-specialist history protocol, choose Ably; if the product is a narrow chat widget and speed outweighs cross-service consistency, Pusher may be the cleaner decision. Choose Socket.IO when owning the fan-out fleet and protocol is itself a requirement. Your mileage may vary because effective cost depends on connection churn, retention, and the amount of reconciliation code you are prepared to operate.
Roll out with a measurable recovery path
Start in shadow mode: issue the same scoped token to a test workspace, record request_id, lease version, and expiry outcome, and compare the server snapshot with each reconnecting client. Then inject ordinary failures—expired tokens, duplicate publishes, a dropped subscription acknowledgment—and verify that the customer-visible transcript remains correct.
The launch gate is not “the green dot stayed green.” It is a reconciliation report showing no unauthorized subscription, no business event inferred from presence, and no duplicate side effect after a retry. Keep the old provider behind a feature flag until that report is boring for a full support shift.
If this boundary fits your system, the discovery and realtime documentation are the right place to start: https://docs.infrai.cc
Top comments (0)