DEV Community

Daniel Pertu
Daniel Pertu

Posted on

I put Redis on the delivery path of my WebSocket server, and it cost every phone in the room seconds

A while ago I wrote up the fan-out layer of our quiz server as a small, pleasant piece of code: each instance keeps its own socket map, broadcasting publishes to a per-session Redis channel, and every instance delivers to its local clients when the message comes back round, including the instance that published it.

I liked that symmetry. One code path, no special case for "my own clients", trivially correct across any number of instances.

It was the worst decision in the codebase, and the version of it I shipped had a failure mode I did not see coming.

The flaw in the symmetry

Read the sentence again: including the instance that published it.

On a single instance, which is what we actually run most nights, every broadcast took a network round trip to Redis and back before reaching a phone that was connected to the very process doing the publishing. The socket was in a Map three stack frames away. The message went to another machine and came back to be delivered to it.

Latency on a healthy Redis is a millisecond or two, so for a long time this looked like elegance I was getting for free. It was not free, and the bill arrived in the one configuration I had been proud of supporting.

The part that actually hurt

The fan-out was built to be optional. Run one instance and you need no Redis at all. Except that ioredis, handed no URL, does not refuse to work. It connects to localhost:6379, because that is a sensible default for a library and a terrible one for this.

So on a deployment with no Redis configured, every broadcast went to a connection that was retrying against a port with nothing behind it. The publish did not fail fast. It worked through its retry policy, and only when that was exhausted did the local fallback run and the phones get their question.

That is up to several seconds of added latency on every question, on every phone, on exactly the deployments that chose the simplest possible setup. The feature flag was not off. It was on, pointed at nothing, and slow.

Worse, it was invisible to me, because my development machine had Redis running.

The shape that replaced it

The fix is to notice that cross-instance fan-out and local delivery are not the same operation and were never the same operation. One is an optimisation for a configuration we rarely run. The other is the product.

export function broadcastToSession(sessionId, event, opts = {}): void {
  const envelope: Envelope = { ...opts, event, audience: opts.audience ?? 'all' };
  deliverLocally(sessionId, envelope);
  publishEnvelope(sessionId, envelope);
}
Enter fullscreen mode Exit fullscreen mode

Local delivery first, synchronously, before anything touches the network. Then publish, for whoever else is out there. The comment on it is the requirement:

/**
 * Local delivery happens synchronously, before anything touches the network,
 * so the latency from notify to phone is one WebSocket write.
 */
Enter fullscreen mode Exit fullscreen mode

And Redis becomes genuinely optional rather than nominally optional:

const REDIS_URL = process.env.REDIS_URL;
export const redisEnabled = Boolean(REDIS_URL);

export function initRedis(handler: Handler): void {
  if (!redisEnabled || publisher) return;
  ...
}
Enter fullscreen mode Exit fullscreen mode

No URL, no clients, no connection attempts, no retry budget being spent on a port that is not listening. publishEnvelope returns immediately when there is no publisher.

The loop you have to break

Delivering locally and publishing means the publishing instance will also receive its own message back through the subscriber, and deliver everything twice. So the wire format carries an origin:

type Wire = { origin: string; sessionId: string; envelope: Envelope };
const INSTANCE_ID = randomUUID();

subscriber.on('message', (_channel, message) => {
  const wire = JSON.parse(message) as Wire;
  if (wire.origin === INSTANCE_ID) return;
  onRemote?.(wire.sessionId, wire.envelope);
});
Enter fullscreen mode Exit fullscreen mode

A UUID per process, discarded on restart, compared on the way in. This is the one piece of complexity the old symmetric design genuinely saved, and it is six lines. I traded six lines for seconds of latency and did not realise I had.

The two connections now disagree about queueing, on purpose

The other thing that changed is subtler and I think more generally useful. The publisher and the subscriber are configured differently:

publisher = createClient('publisher', { offlineQueue: false });
// The subscriber keeps its offline queue so subscribe calls made during a
// reconnect are not lost; ioredis re-subscribes known channels on reconnect.
subscriber = createClient('subscriber', { offlineQueue: true });
Enter fullscreen mode Exit fullscreen mode

The publisher has its offline queue disabled:

// A publish while disconnected fails at once instead of queueing: the local
// clients already have the event, and stale cross-instance events are worse
// than missing ones (the snapshot on reconnect fills the gap).
Enter fullscreen mode Exit fullscreen mode

A queued publish is a promise to deliver a game event later. Later is useless and actively harmful: a question-launched event arriving after the question has closed would push the wrong state onto phones that had already moved on. Failing immediately is correct, and the cost is that clients on another instance miss one event, which the snapshot they get on their next reconnect repairs anyway.

The subscriber keeps its queue, because a subscribe call is not an event with a shelf life. It is a standing intention, and losing one during a reconnect means an instance silently stops receiving a session's messages for the rest of the night. Different queueing policy, same connection pool, and the reason is entirely about what the queued thing means.

Failure is a warning, not a crash

Unchanged from the first version, and still the part I would keep:

publisher.publish(channelFor(sessionId), JSON.stringify(wire)).catch((err: Error) => {
  console.warn('[redis] publish failed (other instances miss this event):', err.message);
});
Enter fullscreen mode Exit fullscreen mode

The log line names the exact consequence. Not "publish failed", which tells an on-call engineer nothing about whether the product is broken, but "other instances miss this event", which says precisely what is degraded and lets them decide whether it matters at 9pm on a Tuesday.

What I would tell myself

  1. "Including the instance that published it" is a smell. If a message leaves the process to come back to the same process, you have chosen symmetry over the thing the system exists to do.
  2. Optional infrastructure has to be optional at the point of configuration, not at the point of fallback. A library's sensible default becomes your worst outage when the absence of configuration is a supported mode.
  3. Test the degraded configuration, not just the degraded network. "No Redis configured" was a supported setup that nobody had ever actually run, including me.
  4. Pick queueing policies per message, by asking what a late delivery would mean. Sometimes late is worse than never.

Where this is pointed

  • pub-trivia.app/features/live-leaderboard is the thing whose latency this was: standings on the big screen and on every phone the moment a question closes, which does not survive a round trip it did not need.
  • pub-trivia.app/features/multi-venue is the only reason the cross-instance path exists at all: several hosts running separate quizzes at the same time, possibly on different instances behind a load balancer.
  • pub-trivia.app/solutions/pubs is the room where a two-second delay between the host pressing Next and the tables seeing a question is the difference between a quiz and an argument.
  • Start a free session if you want to time it yourself: press Next on a laptop and watch a phone next to it.

Top comments (0)