DEV Community

Nitin Doyal
Nitin Doyal

Posted on

Designing a Real-Time Token Queue: What We Learned Building Appointment Systems

A token queue looks like a trivial problem until you build one. It is a sorted list with a counter — surely the simplest thing you can ship?

It is not. The failure modes are all in the edges: two people joining in the same second, a staff member serving out of order, a phone that reconnects after four minutes underground, and a customer staring at a number that has not moved in eleven minutes wondering if it is broken.

We have spent time working through how token-based queue and appointment systems behave in production for Indian clinics, salons and service counters. Here is what actually matters when you design one.

1. The queue is not a list — it is an event log

The tempting model is an array of tokens with a currentIndex. Do not do it.

If your source of truth is "the current position", you cannot answer the questions that matter:

  • How long did this customer actually wait?
  • Which counter is the bottleneck?
  • How many joined and never got served? (Your walk-away rate.)
  • What was the queue state at 4:12 pm last Tuesday?

Model every transition as an event instead: joined, called, served, skipped, left. The current position becomes a fold over the log. Analytics become a query over the log. Disputes become a lookup.

This is the same reason billing systems keep ledgers rather than balances. You will be asked "why was I skipped?" and you need a defensible answer.

type QueueEvent =
  | { type: 'joined';  tokenId: string; at: number; channel: 'qr' | 'counter' | 'api' }
  | { type: 'called';  tokenId: string; at: number; counterId: string }
  | { type: 'served';  tokenId: string; at: number; counterId: string }
  | { type: 'skipped'; tokenId: string; at: number; reason: string }
  | { type: 'left';    tokenId: string; at: number }
Enter fullscreen mode Exit fullscreen mode

2. Ordering must be decided at join time, not read time

The classic race: two clients hit POST /queue/join concurrently, both read max(position), both write max + 1. One customer's place silently disappears.

Fix it at the database, not in application code:

  • A serialised counter row (SELECT ... FOR UPDATE), or
  • A monotonic sequence / SERIAL column with the position derived from it, or
  • An append-only table where position is assigned by the insert transaction.

Whatever you choose, the invariant is: position is assigned atomically with the join, and no read path ever computes it. Under load, the bug appears maybe one shift in fifty — which is exactly the frequency that destroys trust in the product.

3. Real-time delivery needs a fallback that is boring

Push updates (WebSocket or SSE) make the customer's screen feel alive. They also fail constantly on mobile networks.

Design for it from the start:

  • Snapshot first, then deltas. On connect, send the full state. Only then stream changes. A client that reconnects mid-queue must be able to rebuild without replaying history.
  • Sequence numbers on every message. Gap detected → refetch snapshot. Do not try to be clever about resuming.
  • Polling as a first-class path. A 15-second poll is an acceptable degraded mode. Many customers are on patchy 4G; the experience must not depend on a persistent socket.
  • Heartbeats with a short timeout. Detect dead sockets quickly so you stop counting phantom viewers.

Honest note: for most counter queues, a well-built 10-second poll outperforms a fragile WebSocket implementation. Ship the poll, add push when you have measured the need.

4. Serving out of order is a feature, not a bug

Reality at any counter: token 14 is a no-show, token 15 has to leave for a school pickup, and the receptionist serves token 17 because it is a five-minute form fill.

A rigid FIFO queue forces staff into workarounds — which means the software stops matching reality within a day.

Support explicit, auditable transitions instead:

  • skip (customer not responding) with a required reason
  • recall (bring them back)
  • priority (genuine cases: accessibility, emergencies at a clinic)
  • reorder with a permission check

Every one of these must be an event with a timestamp and an actor. If you cannot reconstruct why someone was served out of order, you will lose the argument at the counter.

5. Merge booked appointments with the walk-in queue

Booking systems and walk-in queues are usually built as two separate features. In practice they are one problem.

A customer who booked 3:00 pm and a customer who walked in at 2:55 pm are competing for the same chair. If you keep two lists, your double-book rate climbs until staff start managing on paper again.

The workable model:

  1. Booked appointments occupy slots on a resource (a stylist, a room, a counter).
  2. Walk-ins fill gaps between slots, estimated by real historical service duration.
  3. Late arrivals degrade into the walk-in queue rather than silently breaking the slot.
  4. One ordered view is presented to staff; customers only see their own position and estimate.

That last point matters: customers do not need global ordering, they need an honest estimate and a signal when they are close.

6. Waiting time beats position as the UX

"you are 7th" is a weak message. Seven at a quiet clinic is four minutes; seven on a Monday morning is forty.

Two improvements:

  • Show an estimate with a range — "about 20–30 minutes" — and refresh it as observed service times change. Ranges survive variance; precise numbers do not.
  • Notify on distance, not on position — "your turn is next" when they are 2 away, so they can come back from the car park.

Never promise a time you cannot hit. An estimate that runs long is worse than no estimate at all, because it converts a patient wait into a complaint.

7. Analytics is the real product

The token display is what the customer sees. The reporting is what the business owner pays for.

Worth having from day one:

Metric Why it matters
Median wait by hour Flat the peaks; staff to the real load
Service time p50 / p95 per staff Separates slow service from slow systems
Joined vs served Your walk-away rate — the money leak
Booking share vs walk-in Whether the link is actually working
No-show rate Whether reminders are earning their cost

p95 matters more than the mean. One 40-minute colour service will drag the average and hide the fact that everything else ran at 12 minutes.

8. Operational details nobody thinks about

  • Idempotency keys on join. Mobile connections retry. Duplicate tokens are embarrassing and common.
  • Timezone in storage and in display. Store UTC, render in the business's timezone, and show it. Indian businesses operate IST; daylight-saving bugs are rare here but off-by-one-day bugs are not.
  • Offline-tolerant staff UI. The receptionist's screen must keep working through a 30-second drop, then reconcile.
  • A kill switch per counter. Sometimes you just need to pause intake because the doctor stepped out.

The takeaway

The queue itself is easy. The hard parts are concurrency, honesty about time, and letting staff break the order without breaking the data.

Get those three right and the system survives contact with a busy Saturday — which is the only benchmark that counts.

If you want to see a production version of this in action, SWIQ runs a queue and booking system built for Indian clinics, salons and service counters — live token tracking, a QR booking page, and WhatsApp turn alerts, with a free demo at swiq.services/landing.

Further reading: this guide to appointment booking with live tokens covers the customer-facing flow, and this queue management system guide covers the operational side.


Top comments (0)