TL;DR: In Nodejs, queue each health-chat event, commit a durable per-user inbox row, and only then publish a live hint. Offline users still get notified after reconnect because they fetch rows after their last acknowledged cursor. Presence may decide whether to send the hint, but it must never decide whether the notification is stored.
That rule separates two jobs. The queue protects work while a process is busy or restarting. The inbox protects a user's notification while their browser is asleep, their phone changes networks, or their room connection has not been re-established. A live channel then makes an already committed row appear quickly.
Store first.
How can Nodejs queue then publish so offline users still get notified?
For a clinical chat room, a connected socket is weak evidence of reachability. A tab can stop responding before its connection is declared dead, and a reconnect can overlap the previous session. If the server turns that fuzzy signal into a delivery verdict, a stale presence record can suppress the only notification a user needed.
Presence lies.
Use presence for prioritization. Give every session a short lease, refresh it with heartbeats, and expire it unless the server observes another heartbeat. Multiple sessions should coexist because a clinician may have a workstation and a phone open. The useful question is not “does this user have one socket?” but “is there at least one current session to which a hint is worth publishing?”
Accuracy has a cost. Shorter leases detect departed clients sooner, but create more heartbeat traffic and make transient pauses look like disconnects. Longer leases reduce churn, but increase false-online time. Choose the lease from the product's tolerance for a delayed hint, then measure false-online and false-offline decisions rather than treating the heartbeat interval as a magic constant. Storage still happens either way.
Consider the two recipients in the example below. clinician-7 has a laptop session whose lease is current when the worker starts, while patient-19 has no current session. The worker writes both inbox rows before it asks presence anything. It then publishes one hint to the clinician. If that publish fails, the row remains; if the laptop disconnects a millisecond later, the row remains; if the worker restarts and claims the same job again, the stable event-and-user key resolves to the existing row. The patient gets no hint at this moment, which is fine. On reconnect, both clients ask for rows after their own saved cursors. This sequence has two recipients, one attempted live hint, and one recovery rule. There is no branch in which an inaccurate presence answer deletes or skips durable work. That is the property to protect when adapters, databases, and transports replace the in-memory pieces.
A runnable queue-then-publish example
The flow is small enough to audit. An application transaction creates the chat event and an outbox job. A worker claims that job, inserts one inbox row per recipient with a stable idempotency key, and only then checks presence. Online recipients receive a lightweight hint containing the cursor. Offline recipients receive no live hint, but their inbox row waits for the next sync. The worker may retry without creating duplicates.
This TypeScript example keeps the adapters in memory so the ordering and failure behavior are visible. Replace them with transactional storage and a durable job system in production; preserve the interfaces and invariants.
import { randomUUID } from "node:crypto";
type Job = {
id: string;
roomId: string;
eventId: string;
recipientIds: string[];
preview: string;
};
type InboxRow = {
cursor: number;
userId: string;
eventId: string;
roomId: string;
preview: string;
};
interface Inbox {
putOnce(key: string, row: Omit<InboxRow, "cursor">): Promise<InboxRow>;
after(userId: string, cursor: number): Promise<InboxRow[]>;
}
interface Presence {
hasLiveSession(userId: string, now: number): Promise<boolean>;
}
interface Realtime {
publish(userId: string, message: {
type: "inbox.updated";
cursor: number;
}): Promise<void>;
}
class MemoryInbox implements Inbox {
private nextCursor = 1;
private readonly byKey = new Map<string, InboxRow>();
async putOnce(
key: string,
row: Omit<InboxRow, "cursor">,
): Promise<InboxRow> {
const existing = this.byKey.get(key);
if (existing) return existing;
const saved = { ...row, cursor: this.nextCursor++ };
this.byKey.set(key, saved);
return saved;
}
async after(userId: string, cursor: number): Promise<InboxRow[]> {
return [...this.byKey.values()]
.filter((row) => row.userId === userId && row.cursor > cursor)
.sort((a, b) => a.cursor - b.cursor);
}
}
async function handleJob(
job: Job,
inbox: Inbox,
presence: Presence,
realtime: Realtime,
now = Date.now(),
): Promise<void> {
for (const userId of job.recipientIds) {
const row = await inbox.putOnce(`${job.eventId}:${userId}`, {
userId,
eventId: job.eventId,
roomId: job.roomId,
preview: job.preview,
});
if (await presence.hasLiveSession(userId, now)) {
try {
await realtime.publish(userId, {
type: "inbox.updated",
cursor: row.cursor,
});
} catch {
// The committed inbox row is recovered by the next sync.
}
}
}
}
async function reconnect(
userId: string,
lastCursor: number,
inbox: Inbox,
): Promise<InboxRow[]> {
return inbox.after(userId, lastCursor);
}
const job: Job = {
id: randomUUID(),
roomId: "care-team-42",
eventId: randomUUID(),
recipientIds: ["clinician-7", "patient-19"],
preview: "A new care-team message is available",
};
The example catches publish failure only after putOnce succeeds. Reversing those lines creates a loss window: the client can observe a transient hint, disconnect before fetching, and leave no durable notification to recover. It also matters that the hint carries a cursor rather than the full clinical message. The client uses the authenticated sync path as the source of truth, where authorization can be checked again.
One caveat: the in-memory cursor is global and putOnce is illustrative. A production implementation needs an atomic uniqueness constraint on (event_id, user_id) and a monotonic cursor with a defined scope. Per-user cursors make reads simple. Global cursors make ordering across recipients observable but can expose gaps; gaps must not be interpreted as missing user-visible messages.
Failure semantics before transport choices
Aim for at-least-once job processing plus idempotent inbox insertion. Exactly-once language is tempting, but the useful contract lives at the boundary the user sees: one logical inbox item per event and recipient, even if a worker runs twice or a live hint arrives twice. The client should merge by event ID and acknowledge the highest cursor it has durably applied.
Order needs a written definition. If a single room requires strict display order, assign the ordering key when the room event is committed, not when a worker happens to run. If independent rooms can interleave, a per-user inbox cursor can represent notification order without pretending it is the clinical conversation's canonical order. Those are different clocks.
Backpressure belongs in the design. A worker should claim a bounded batch, limit concurrent fanout, and retry failed jobs with delay. Large rooms can otherwise monopolize workers while small conversations wait. Keep job payloads compact: identifiers and a safe preview are enough. Fetch mutable details during authorized inbox reads instead of duplicating a full health record through every queue and log.
WebRTC data channels can carry application data once peers establish the required connection, but they do not replace this durable inbox contract. A browser that was absent had no peer connection on which to receive the event. The recovery path still needs server-side state, regardless of whether the live hint travels over a data channel, a bidirectional socket, or another push mechanism.
Transport comes second.
Test the gaps, not the happy path
Start with four deterministic tests. Process the same job twice and assert one inbox row. Fail publish after insertion and assert reconnect returns the row. Mark presence stale while an old session remains registered and assert storage still occurs. Finally, connect two sessions for one clinician, deliver duplicate hints, and assert the client applies the inbox item once.
Then exercise awkward timing: pause a client after it receives a cursor but before it persists the acknowledgment; reconnect it with the old cursor; restart a worker after the inbox insert but before job completion. The expected result is duplication at internal boundaries and no duplicated logical notification. This is the trade: a little deduplication buys a much smaller loss surface.
Retries aren't exceptional here. They're part of the contract.
Observe the same boundaries in deployment. Track outbox age, job attempts, inbox insert latency, publish attempts, publish failures, active presence leases, reconnect sync size, and acknowledgment lag. Do not collapse “published” into “delivered.” Published means the transport accepted a hint; acknowledged means a client reports applying an inbox row. Even that acknowledgment is product telemetry, not proof that a human read the message.
For cost control, retention and fanout dominate the discussion more than the choice of wire protocol. Set inbox retention from the application's notification policy, archive or delete expired rows predictably, and cap reconnect pages so a long-absent account cannot trigger an unbounded response. Sampling routine transport logs while retaining error and lag aggregates can also prevent observability from becoming a second copy of sensitive message content.
The operational decision rule
Before release, verify the invariant in staging with forced disconnects: every accepted room event creates the intended recipient rows, retries keep their cardinality unchanged, and a reconnect from an older cursor catches up in order. Confirm lease expiry under suspended tabs and network changes, then check that false presence decisions affect latency only. They must never affect durability.
Roll out with bounded worker concurrency and alerts tied to age and acknowledgment lag, not raw connection count. Connection count is capacity information. It says little about whether a particular clinician can recover a missed notification.
The final choice is straightforward: queue the work, commit the inbox, then publish the hint. If presence is wrong, the experience may be slower for one reconnect cycle. If persistence is conditional on presence, the notification can disappear. In health chat, that is the wrong bargain.
Top comments (0)