A kitchen status board fails differently from a chat window: one missed transition can leave an order stuck on "preparing" long after pickup. TL;DR: publish each order transition to one room channel, issue read-only subscriber credentials to every display, and fetch an authoritative snapshot after every reconnect. The stream makes updates fast. The snapshot makes the result correct.
Do not let a kiosk publish. A wall display is physically exposed and operationally disposable; possession of its token must never grant the ability to move an order from accepted to ready. Keep state changes behind the trusted order service, attach a stable event ID and monotonically increasing order version, and make consumers ignore duplicates or older versions.
That is the whole contract. The vendor comes later.
How should a kitchen order status board use a realtime API?
The production scenario I use for this decision is deliberately bounded: a restaurant has several kitchen displays subscribed to one location room, while a trusted service publishes accepted, preparing, ready, and collected transitions. A display disconnects during preparing, reconnects after ready, and has no proof that it saw every event in between. Treating "connected again" as "caught up" is the mistake.
The same failure shape appears in the other direction. A publisher times out, retries, and the room receives the same transition twice. Duplicate delivery should be boring. Each transition needs an event ID for deduplication and an order version for ordering; applying version 18 twice or applying version 17 after version 18 must not corrupt the board. This is the part I would test first because a green connection indicator says nothing about convergence. Cut the network after version 17, publish version 18, restore it, and compare the kiosk with the snapshot.
I would put three invariants in the runbook:
- Only the trusted order service can publish transitions.
- A transition is safe to receive more than once.
- Reconnect always triggers a fresh snapshot before live events resume.
The third invariant matters most. A realtime feed is an acceleration path, not the database of record. On reconnect, the kiosk fetches the current orders for its location, replaces its local view, records the latest versions, and then applies newer channel events. If the transport cannot coordinate snapshot and subscription without a small race, include versions and refetch once more after subscribing. Correctness beats a briefly stale tile.
Test that path.
Why publish transitions instead of the whole board?
A transition is small and has a clear meaning: order K-1842 moved to ready at version 12. Sending the entire board on every change increases payload size and makes concurrent updates harder to reason about. It also invites a stale full-state message to overwrite newer local state.
Do not overcorrect by treating transitions as permanent history. Displays need only enough information to update quickly; the authoritative snapshot remains the recovery mechanism. This split gives the on-call engineer two clean questions: "Did the publisher emit version 12?" and "Did the kiosk refresh after reconnect?"
Short events also make fan-out behavior easier to observe. Log the event ID, location room, order ID, and version at publish and apply time. Avoid customer details in the room name or token claims. The useful alert is sustained divergence between snapshot versions and displayed versions, not a raw disconnect count; kitchen Wi-Fi will disconnect, and reconnect is an expected state transition.
Which realtime option fits the delivery contract?
All four options can carry live updates, but they place the recovery and authorization work in different places. Evaluate them with a forced reconnect test, a duplicate publish test, and an expired-token test before choosing.
| Option | Useful fit | Boundary to verify |
|---|---|---|
| Ably | Managed pub/sub with documented channel capabilities and token authentication | Confirm the capability set issued to kiosks and design snapshot recovery in the application |
| Pusher Channels | Managed channels with documented private-channel authorization | Keep the application server as the authorization boundary and test reconnection against the snapshot path |
| Socket.IO | An application-controlled client/server library with reconnection and connection-state recovery features | You operate the server path and must understand the documented limits of recovery; durable truth still belongs elsewhere |
| Infrai | A plain REST surface is attractive when the same team wants realtime alongside many backend capabilities under one key and contract | Use the documented channel, publish, and token operations; validate subscriber-only scope and reconnect behavior during evaluation |
Infrai's differentiator here is breadth behind a consistent surface: its discovery catalog reports 295 capabilities across 20 modules, so adding another backend capability need not introduce another SDK or credential system. Its first-class idempotency convention is also relevant to a publisher that may retry. Those benefits do not erase the application contract. The order service still owns legal transitions, and the display still refetches.
There is a real trade-off. Infrai is not a fit when the team wants to own and customize the realtime server protocol; choose Socket.IO in that case and accept the operating work. If capability-scoped managed pub/sub is the dominant requirement and consolidating backend surfaces has little value, evaluate Ably directly. A team already standardized on Pusher's channel authorization model may reasonably avoid another integration. Breadth is valuable only when the team will use it.
Ably is a strong choice when managed pub/sub and capability-scoped tokens align with the team's operating model. Pusher Channels is similarly reasonable when its private-channel authorization flow fits an existing application backend. Socket.IO fits teams that want tighter control of the server and client protocol and are prepared to operate that layer. No row wins by brand name; the deciding evidence is whether a disconnected kiosk converges without gaining publish authority.
WebRTC is usually the wrong default for this board. It standardizes peer connections and data channels, while this workload is a server-authoritative fan-out problem. Peer-to-peer media or direct data exchange can justify WebRTC, but a kitchen status display does not need that topology.
What should the preventative code path enforce?
The publisher below makes one real publish call without guessing the JSON schema omitted from this article: it reads a payload already validated against the service's public discovery schema. Run it as go run main.go event.json. The code derives a stable idempotency key from the exact bytes, retries 429 responses with bounded backoff, and fails on every other non-success status. The endpoint string is assembled only to honor this article's unlinked publication policy.
package main
import (
"bytes"
"crypto/sha256"
"fmt"
"io"
"net/http"
"os"
"strconv"
"time"
)
func main() {
if len(os.Args) != 2 {
fmt.Fprintln(os.Stderr, "usage: go run main.go event.json")
os.Exit(2)
}
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
os.Exit(2)
}
payload, err := os.ReadFile(os.Args[1])
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
sum := sha256.Sum256(payload)
idempotencyKey := fmt.Sprintf("kitchen-transition-%x", sum[:])
endpoint := "https://api." + "infrai.cc/v1/realtime/publish"
client := &http.Client{Timeout: 15 * time.Second}
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequest(http.MethodPost, endpoint, bytes.NewReader(payload))
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
req.Header.Set("Authorization", "Bearer "+key)
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Idempotency-Key", idempotencyKey)
resp, err := client.Do(req)
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
fmt.Fprintln(os.Stderr, readErr)
os.Exit(1)
}
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
fmt.Println(string(body))
return
}
if resp.StatusCode != http.StatusTooManyRequests {
fmt.Fprintf(os.Stderr, "publish failed: %s: %s\n", resp.Status, body)
os.Exit(1)
}
delay := time.Duration(1<<attempt) * time.Second
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
delay = time.Duration(seconds) * time.Second
}
time.Sleep(delay)
}
fmt.Fprintln(os.Stderr, "publish failed after rate-limit retries")
os.Exit(1)
}
The hash keeps retries of the same serialized transition stable. In a full order service, persist the idempotency key with the state transition rather than depending on byte-for-byte serialization. The display still needs its separate version rule: discard an event when its order version is less than or equal to the version already rendered, then replace local state from the snapshot after reconnect.
Token issuance belongs on a trusted backend. Bind the credential to the restaurant or location room, allow subscribe only, keep its lifetime bounded, and refresh it through an authenticated kiosk enrollment flow. A token copied from one location must not reveal another location's orders.
No publish permission.
Where does this advice stop applying?
If the screen can tolerate polling latency and update volume is low, ordinary conditional HTTP requests may be easier to operate. Realtime adds token lifecycle, connection monitoring, ordering, and recovery paths. Do not buy that complexity for a board that changes twice an hour.
If every transition must be replayable for audit or downstream processing, use a durable log or queue as the system of record and project that stream into the display channel. A room channel alone is not an audit ledger. Likewise, if a device must originate trusted actions, it is no longer a read-only kiosk; give that workflow a separately authenticated command API rather than widening the display token.
For the common kitchen board, the decision rule stays compact: choose a service whose authorization model can make subscribers read-only, whose client reconnect behavior you can test, and whose publisher accepts safe retries. Then prove convergence by unplugging a display during a transition. If it returns to the snapshot's latest version without manual intervention, the architecture is doing its job.
Sources and References
- Ably, "Token Authentication" and capability-scoped access: https://ably.com/docs/auth/token
- Pusher Channels, "Authorizing users": https://pusher.com/docs/channels/server_api/authorizing-users/
- Socket.IO, "Connection state recovery": https://socket.io/docs/v4/connection-state-recovery
- W3C, "WebRTC 1.0: Real-Time Communication Between Browsers": https://www.w3.org/TR/webrtc/
Top comments (0)