Every real-time app has a line like this somewhere:
void notifyWsServer(sessionId, 'question_launched')
return { success: true }
Do the work, tell the WebSocket server about it, do not make the caller wait on a message bus. On a long-lived Node process this is correct and has been correct for twenty years. The event loop is still there after your handler returns, so the fetch finishes.
On a serverless platform it is a coin toss, and we found out the way you always find out.
The symptom
A host presses Next. The question is written to the database, the response goes back, the host view advances. Forty phones show the previous question's result screen and stay there. Nobody gets the question that, as far as the database is concerned, started four seconds ago.
Not always. Perhaps one launch in ten, more on a cold start, essentially never in local development.
The reason is in the platform's execution model rather than in the code. A serverless function may be frozen or torn down as soon as it has sent its response. A promise you did not await is not part of the response. So question_launched was a fetch racing a freeze, and the race was decided by how warm the container happened to be.
What made it unpleasant to diagnose is that the database was always right. Replay the state and everything is consistent: the question is launched, the timestamp is there. The only thing missing is that nobody was told. There is no error, no exception, no failed row. The side effect simply never happened.
Two kinds of side effect
The fix is not "await everything". Awaiting a message-bus call on the hot path means the player who just tapped an answer waits for our fan-out before their screen acknowledges the tap, and that cost lands on every answer from every phone.
So we split the notifications by a single question: can the game proceed without this message?
For the transitions, no. If question_launched is lost, the room is stuck:
/**
* Await it for events the game cannot progress without (launch, close, end):
* a serverless function can be frozen as soon as its response is sent, and an
* un-awaited fetch for question_launched was liable to die with it, leaving
* every phone waiting for a question that had already started.
*/
Launch, close and end are awaited. Three calls per question, on a path the host triggers, where a hundred extra milliseconds costs nothing anyone can perceive.
For joins and answers, yes. A lost answer_received means the host's answered-count is briefly stale, which the next event corrects. Those are not worth a player's latency, but they are also not worth losing. That is what after() is for:
export function notifyWsServerAfterResponse(
sessionId: string,
event: NotifyEvent,
payload?: NotifyPayload,
): void {
try {
after(() => notifyWsServer(sessionId, event, payload))
} catch {
void notifyWsServer(sessionId, event, payload)
}
}
after from next/server is the explicit version of the thing void was pretending to be. It schedules work to run once the response has been sent and tells the platform to keep the invocation alive until that work is done. The response is not delayed; the work is not abandoned.
The whole app has exactly two callers of it, both on genuinely hot paths: a player joining, and a player answering. Everything else awaits.
The catch clause is not defensive padding
after() throws outside a request scope. Tests, scripts and anything calling the same function from a non-request context would get an exception from the notification helper, which is an absurd place for a test to fail.
Hence the fallback: inside a request, defer; outside one, run it now. In a test or a script there is no response to get out of the way of and no platform waiting to freeze the process, so running immediately is not a compromise, it is the correct behaviour for that environment.
Failures are logged, never thrown
The other half of the policy, which predates the freeze bug and survived it:
if (err instanceof Error && err.name === 'TimeoutError') {
console.error(`[notifyWsServer] Notify timed out after ${NOTIFY_TIMEOUT_MS}ms for event '${event}' on session ${sessionId}`)
}
The notification has an eight-second timeout and never propagates its failure to the caller. The state change already committed. Turning a broadcast failure into a failed action would mean a host whose Next click reported an error for a question that had, in fact, started: the worst of both outcomes, because now the UI disagrees with the database as well.
A dropped broadcast is survivable because of the layer underneath. The WebSocket server re-derives what should be scheduled from the database every three seconds, and every reconnecting client is sent a full state snapshot rather than a replay. A missed message costs a few seconds of staleness, not a stranded room. That is the property that lets this code log and move on instead of retrying.
What generalises
Three things, in order of how often I now say them out loud:
- On any platform that can freeze a process after a response, an un-awaited promise in a request handler is not deferred work. It is work that may not happen, with no diagnostic when it does not.
- Sort your side effects by whether the system can make progress without them. The ones that must happen get awaited. The ones that merely should happen get an explicit deferral primitive, not a bare
void. - Analytics and telemetry hide inside this same trap. We turned off the rate limiter's analytics partly because the write happened on a path that never awaited it, which makes it not data collection but a coin toss that costs money.
One last note on testing this. None of it reproduces locally, because a local Next.js server never freezes after a response, so the un-awaited version passes every test you can write on a laptop. The only honest check is a preview deployment and a stopwatch.
Where this shows up in the product
- pub-trivia.app/features/live-leaderboard is the promise this code has to keep: standings move the moment a question closes, on the big screen and on every phone. "Usually arrives" is not a feature you can write that sentence about.
- pub-trivia.app/features/qr-code-quiz-joining is one of the two paths that uses the deferred notification, because a player who has just scanned a code should see the lobby before our servers finish talking to each other.
- pub-trivia.app/features/big-screen-display is the host view on the venue's TV, which is where a lost broadcast is most visible: it is the screen the whole room is looking at.
- You can watch it work on a free session: sign up, open the host view on a laptop and join from two phones, and the transitions you see are the awaited ones.
Top comments (0)