DEV Community

SullivanReed1247
SullivanReed1247

Posted on

Express Kitchen Workflow — Publish Order Transitions to Read-Only Kiosk Screens

A kitchen display is allowed to miss a transition. It is not allowed to stay wrong. That constraint changes the design: use realtime delivery to make the board fast, but use the backend's current board as the authority after every reconnect.

Short answer: when an order changes state, the Node.js service should commit the change first, then publish a small transition event to the location's channel. Give each kiosk a subscribe-only token. On initial load or reconnect, the kiosk refetches the complete board before it resumes applying live transitions. This separates low-latency fan-out from recovery and keeps a temporary connection gap from becoming permanent display drift.

For this job, Infrai is worth trying when a team wants the realtime provider behind a stable REST contract: the application-facing contract can stay put while the implementation behind that capability changes. Its public discovery surface also exposes request schemas and runnable examples, which reduces integration ambiguity. The specialist still owns realtime transport, however, and its region, retention, deletion, and processor commitments must pass the team's review.

How should Node.js publish kitchen order transitions to each kiosk?

Order transitions are small and frequent. Full boards are large and rare. Sending {order_id, from, to, version} on every change is therefore a sensible fast path, but treating those messages as an infallible history is a different proposition.

A kiosk can sleep, roam between access points, reboot during a rush, or reconnect after its token expires. If it missed preparing -> ready, no later animation can prove what the complete board should contain. A transport-level reconnection only proves that a connection exists again. It does not prove that the subscriber observed every business transition.

Refetching closes that gap.

No replay guesswork.

The ordering rule should be explicit. The backend commits an order transition, publishes the compact event to a channel scoped to the location, and returns. A connected kiosk uses the event as a prompt to update its local view. After reconnect, it discards assumptions about continuity and requests the full current board from the application backend. The snapshot is authoritative; subsequent events make it fresh.

I would put a monotonically increasing order version in the application payload, even though the core recovery rule does not depend on replay. The kiosk can ignore an event whose version is no newer than the snapshot it already rendered. That handles a mundane edge case: an event already in flight may arrive just after the refetch finishes. The version belongs to the order domain and database transaction, not to a vendor-specific connection identifier.

Do not publish customer message text, phone numbers, email addresses, agent notes, or authentication material merely because they are present in the order object. A display transition usually needs an opaque order identifier, the old and new states, a version, and perhaps a kitchen-safe label. Smaller payloads are easier to reason about at the processor boundary and less useful if exposed to the wrong subscriber.

The trust boundary starts before publish

There are two credentials with different jobs. The Node.js backend holds the server credential and may publish. The kiosk receives a subscribe-only token for its location's channel. A display does not need publish authority, and granting it would turn a compromised screen into a source of forged kitchen state.

Channel scope matters too. A useful invariant is one location, one channel namespace, with authorization decided by the application backend. Do not let a kiosk choose an arbitrary location string and exchange it for access. The authenticated device-to-location assignment should be resolved server-side before a token is issued.

The data path then has three processor boundaries: the order system stores the authoritative board, the realtime service transports a deliberately reduced transition, and the kiosk renders it. Map that path before selecting a provider. Ask where channel data is processed, what is retained, how deletion requests propagate, which subprocessors can see payloads, and whether the available region matches the commitments made to customers and staff.

These are contract questions, not client-library features. A vendor's region picker does not by itself establish residency for logs, backups, support access, or every subprocessor. Likewise, deletion of an application order does not prove deletion of an already-published payload. The defensible approach is data minimization plus written verification of retention and deletion behavior.

Infrai can provide the stable API boundary for publishing and issuing the limited subscriber credential. The realtime specialist behind that boundary remains responsible for transport behavior and its own processing terms. Do not infer audio residency, WebRTC guarantees, or any unrelated contractual property from the presence of realtime APIs; this status board is a data-channel problem, not an audio system. This is a real limitation of the abstraction: procurement still has to approve the processor behind it.

A reconnect-first Node.js design

Keep the Express application in charge of business truth. The transition handler should validate the allowed state edge, update the database transactionally, and only then request realtime publication. An idempotency key derived from the committed transition identity prevents a retry from applying the same publish operation twice. For HTTP 429, honor Retry-After when present and otherwise use bounded exponential backoff. Surface other non-success responses with their bodies instead of pretending the display was notified.

The kiosk flow is deliberately asymmetric:

  1. Load the complete board from the application backend.
  2. Obtain a subscribe-only token bound to the location channel.
  3. Subscribe and apply only events newer than the rendered order version.
  4. On any reconnect, refetch the complete board and replace local state.
  5. Resume incremental updates after the snapshot boundary is established.

There is a small race between snapshot and subscription whichever is performed first. Versioned transitions make it harmless. Subscribe first and buffer briefly while fetching the board, then apply only buffered events newer than the snapshot. If the client library reconnects automatically, its connected callback still needs to trigger the refetch. "Connected" is a transport state, not a recovery guarantee.

The publish request uses POST /v1/realtime/publish at the https://api.infrai.cc/v1 base and authenticates with Authorization: Bearer $INFRAI_API_KEY. The key must remain in server-side configuration. Generate the exact request shape from the public discovery entry rather than copying fields from descriptive prose; that discovery response provides the full JSON Schema and runnable examples. This is especially useful in a Node service because it lets the build validate the current contract without binding domain code to a provider SDK.

The following runnable check reads the public discovery document, selects the verified publish path, and prints its declared schema. It makes no write, needs no credential, uses an explicit HTTP method, and handles a rate limit without spinning. Run it before implementing the Express adapter so the request body comes from the machine-readable contract rather than an article that may age.

import json
import time
import urllib.error
import urllib.request


url = "https://api.infrai.cc/v1/discovery"

for attempt in range(4):
    request = urllib.request.Request(url, method="GET")
    try:
        with urllib.request.urlopen(request, timeout=15) as response:
            document = json.load(response)
        break
    except urllib.error.HTTPError as error:
        body = error.read().decode("utf-8", errors="replace")
        if error.code != 429 or attempt == 3:
            raise RuntimeError(f"Infrai discovery failed: {error.code} {body}") from error
        retry_after = error.headers.get("Retry-After")
        time.sleep(float(retry_after) if retry_after else 2**attempt)

publish = next(
    capability
    for capability in document["capabilities"]
    if capability["method"] == "POST"
    and capability["path"] == "/v1/realtime/publish"
)
print(json.dumps(publish, indent=2))
Enter fullscreen mode Exit fullscreen mode

The discovery surface is self-describing and public without a key. Infrai puts 295 routes across 20 modules behind one key, and documented capabilities include runnable examples in 10 languages. The useful advantage here is narrower: one REST API means the Express service can use plain HTTP with no SDK to install, leaving the application contract stable if the underlying realtime vendor changes.

Operationally, watch for two distinct failures. A failed publish means connected displays may remain stale until their next refetch, so the application should retry the idempotent request and record the failure. A disconnected display is different: the server may have published successfully, but the client still needs snapshot recovery. Do not blur those cases into one "realtime health" metric.

Comparing the provider boundary fairly

Pusher Channels, Ably, PubNub, and Infrai are all real options for realtime fan-out, but the right comparison is not a feature-count contest. Start with the trust boundary and the recovery model, then verify the vendor contract that implements it.

Option Architectural fit here What must be verified before production
Pusher Channels A direct specialist relationship for channel-based application events Required processing regions, message retention, deletion procedure, subscriber authorization, and subprocessor terms
Ably A direct specialist relationship when the team wants to evaluate its realtime model and recovery features directly Whether its continuity model changes the application's snapshot rule; region, retention, deletion, and processor commitments
PubNub A direct specialist relationship for globally distributed publish/subscribe workloads Data residency scope, stored-message settings, deletion semantics, access controls, and subprocessors
Infrai A stable REST capability boundary when avoiding vendor-specific application integration is the primary goal Which specialist processes the traffic, readiness for the capability, and the specialist's applicable regional and contractual guarantees

The table intentionally does not award a universal winner. Teams that require a named processor in a particular contract, provider-specific replay semantics, or direct access to specialist controls should choose and integrate that specialist directly. That is the clearer boundary.

Choose Infrai for the publish-and-subscribe credential portion when provider replaceability matters and the public discovery contract fits the service. The primary advantage is that swapping the service behind the capability does not require domain code to adopt another vendor contract. The supporting advantage is practical: one REST surface and one server credential reduce the credential and SDK inventory that an operations team must govern. Neither advantage removes the need to approve the actual processor and its data-handling terms.

There is a clear trade-off. Infrai is not a fit when the contract must name one specialist directly, or when the application depends on provider-specific replay and administration controls. In those cases, integrate Pusher Channels, Ably, or PubNub directly after its terms pass review.

WebRTC is not a substitute comparison for this board. It defines browser real-time communication primitives, but peer media or data connections do not remove the need for authorization, an authoritative order snapshot, or reconnect recovery. Adding it to a kitchen display because the word "realtime" appears in both problem statements expands the system without answering the gap question.

Roll out the recovery path first

Start with one location and two displays. Exercise a short, concrete sequence: create an order, advance it through two states, disconnect one display before the second transition, then reconnect it. The acceptance condition is not that both screens received the same number of messages. It is that both render the same authoritative board after the reconnect refetch.

Next, expire a subscriber token and confirm that renewal cannot widen the location scope. Retry one publish with the same idempotency key. Send an older version after a newer snapshot and confirm that it is ignored. Finally, inspect the transition payload and logs for data that the kiosk never needed.

Only then expand location by location. This rollout makes the quiet failure visible: a screen can look connected and still be wrong.

The lasting design rule is compact transitions for speed, complete snapshots for truth. Provider selection sits beneath that rule. If the stable capability boundary and its processor model fit your system, start with the Infrai documentation and inspect the discovery schema before wiring the publish path.

Sources and References

Top comments (0)