Short answer: publish a versioned snapshot of the whole customer queue, let each client compute its own position from that snapshot, and make the snapshot version the consistency boundary. Express can deliver the snapshot over a normal endpoint and push only a new version hint over a realtime channel; the client must fetch the full state again when it detects a gap.
This is a display protocol, not a promise that a position remains true between two paints. A customer who sees position: 4 is seeing the queue at revision 1842. The server should be able to explain that revision later.
What should a Node.js queue snapshot contain for client-computed position?
Start with an ordered list, a stable item identifier, and a revision. Do not publish a bare integer called position; it has no meaning without the ordering rule and the state it came from. For a marketplace pickup queue, a compact snapshot might contain waiting ticket IDs, their status, and the moment at which the server created the revision. The client can then count eligible items before its own ticket.
The invariants I would put in the architecture record are deliberately narrow:
- One revision describes one complete ordering.
- A ticket appears at most once in that ordering.
- Items marked
servedorcancelledare excluded by the same rule on every client. - A client never applies revision 1841 after revision 1842.
- A gap causes resynchronization, not guesswork.
Here is the data shape, expressed in Python because the important part is the state transition rather than Express syntax:
from dataclasses import dataclass
from typing import Literal
Status = Literal["waiting", "serving", "served", "cancelled"]
@dataclass(frozen=True)
class Ticket:
ticket_id: str
status: Status
@dataclass(frozen=True)
class QueueSnapshot:
revision: int
generated_at: str
tickets: tuple[Ticket, ...]
def position_for(self, ticket_id: str) -> int | None:
eligible = [t.ticket_id for t in self.tickets
if t.status in ("waiting", "serving")]
try:
return eligible.index(ticket_id) + 1
except ValueError:
return None
snapshot = QueueSnapshot(
revision=1842,
generated_at="2026-09-14T08:30:00Z",
tickets=(
Ticket("A-104", "serving"),
Ticket("A-105", "waiting"),
Ticket("A-106", "waiting"),
),
)
assert snapshot.position_for("A-106") == 3
The list is the source of truth for the display. A separate position field can be included as a convenience, but it should be treated as derived data and checked against the revision.
How can Express publish whole queue state while clients compute position?
Use two paths with different jobs. A snapshot request returns the complete state and its revision. A realtime message carries a newer revision number, or a small invalidation event; it does not try to carry every intermediate mutation. The client compares the number it has with the number announced. If the next revision is not available, it requests a fresh snapshot.
That distinction matters during reconnects. Suppose a customer loses connectivity while revisions 1843 and 1844 are committed, then receives a push for 1845. Applying only 1845 is safe if the message means “fetch revision 1845”; applying a patch that assumes 1843 is present is not. Whole-state publication trades bandwidth for a much easier recovery story.
In an Express service, make the snapshot read and the revision announcement observe the same committed source. A cache filled before the transaction commits can produce a newer announcement with an older body, which is the exact kind of race that makes a queue display appear to move backward. Include an ETag or revision token so a client can ask whether its copy is still current, and log the revision with every render-related request.
I once treated a websocket reconnect as a harmless transport event. It was not. The browser reconnected successfully, but its last acknowledged revision had been evicted from the server's short replay window. The UI showed a plausible position for nearly a minute. Plausible is not correct. The recovery rule became: after reconnect, request a complete snapshot unless the server can prove the missing revisions are available.
Three words: version first, paint second.
Trade-offs in whole-state publication
| Design | What it buys | Failure boundary | Appropriate use |
|---|---|---|---|
| Full snapshot on every change | Simple client logic and deterministic recovery | Bandwidth grows with queue size | Small or moderate customer queues |
| Snapshot plus revision hint | Lower push volume while retaining a clear resync path | Extra read after most hints | Queues with frequent updates |
| Per-event position patches | Small messages in the happy path | Missed events and ordering bugs become client state bugs | Only when a durable replay log is guaranteed |
| Server-computed position only | Thin clients | Every reorder requires trusted server output and careful cache invalidation | Restricted clients that cannot hold queue state |
The rejected option here is an unversioned “current position” event. It looks efficient, but it cannot tell a stale browser whether position 3 came before or after a cancellation, and it gives support staff nothing stable to reference. It is valid for a disposable animation where temporary jumps do not matter; a customer queue display has a stronger expectation.
The catch is payload size. Whole-state publication is not suitable for a queue with tens of thousands of entries or privacy-sensitive records that every browser must not receive. In that case, keep the authoritative ordering server-side, publish a bounded window plus a cursor, or use a durable event log with explicit replay semantics. Stick with snapshots when recovery simplicity and presence accuracy matter more than shaving every byte.
Failure tests and operational boundaries
Test the protocol as a state machine. Start a client at revision 10, deliver 12 before 11, duplicate 12, reconnect after the replay window, and remove the client's ticket between snapshots. The expected behavior is monotonic revision handling, one resync, and a clear “not in queue” state; it is not a fabricated position.
Measure snapshot age, queue length, resync rate, and the percentage of renders backed by a revision newer than the last user action. Alert on a rising resync rate rather than hiding it with retries. A retry that returns the same stale cache is still stale.
Presence accuracy also has a human boundary: a displayed position is advisory while payment, reservation, or fulfillment decisions remain authoritative on the server. State that in the UI and in the API contract. Your mileage may vary with clock skew and replication lag; the revision, not a browser timestamp, should settle disagreements.
Top comments (0)