DEV Community

Thalion51
Thalion51

Posted on

Nodejs Queue Publish for Offline Users: 12-Minute Live Poll Notifications

Nodejs Queue Publish for Offline Users: 12-Minute Live Poll Notifications

Short answer: in a Nodejs-style service, queue the result before you publish it so offline users still get notified. The live channel is an acceleration path; the queue and durable record are the delivery contract. A B2B SaaS session may have 4,000 participants answering within 12 minutes, including browsers that disappear during the session.

That ordering is the whole design.

What must remain true when a browser vanishes?

The useful question is not “can the server publish?” It is “what evidence proves that each intended recipient can obtain the event after reconnecting?” The invariants are a stable event id, deduplication, a reconnect cursor or durable inbox, and retries that cannot delete the only copy.

Option Guarantee Failure boundary
Direct request publish Low latency Crash can lose the event
Queue first, then fan out Replayable work Needs idempotency and retention
Shared session transcript Easy replay No per-user unread state
Per-recipient inbox Auditable offline recovery More writes and cleanup

For a user-specific result, I choose the per-recipient inbox. A shared session log can supplement it, but should not carry the entire delivery promise.

How should Nodejs queue and publish events for offline users?

The poll-closing request validates the result, assigns an event id, creates inbox entries, and enqueues a fan-out job in one transactional boundary. A worker publishes to connected clients and records attempts. The broker is replaceable; ordering and idempotency are not.

def close_poll(tx, queue, event, recipient_ids):
    tx.insert_event(event.event_id, event.session_id, event.payload)
    for recipient_id in recipient_ids:
        tx.insert_inbox(recipient_id, event.event_id, "pending")
    tx.commit()
    queue.enqueue({"event_id": event.event_id})

def fan_out(store, live_connections, job):
    for row in store.pending_inbox_rows(job["event_id"]):
        envelope = {"id": row.event_id, "type": "poll.closed",
                    "data": store.event_payload(row.event_id)}
        if live_connections.send(row.recipient_id, envelope):
            store.mark_delivered(row.recipient_id, row.event_id)
        else:
            store.record_attempt(row.recipient_id, row.event_id)
Enter fullscreen mode Exit fullscreen mode

Marking delivery after transport acceptance still is not proof that a human saw it. Add a client acknowledgement or durable cursor when the product needs stronger evidence.

How do reconnects avoid duplicates, gaps, and false delivery claims?

On reconnect, the client presents its last applied event id or cursor. The server returns pending inbox rows after that cursor, then resumes the live stream. Apply each (recipient_id, event_id) once. This is deliberately boring data modeling, which is why it survives retries.

Keep the envelope small and stable: poll id, result version, event id, and a link to fetch details. WebRTC provides realtime connection and data-channel behavior, but it does not replace durable application storage or an offline inbox.

Inject failure after the inbox transaction, after a queue lease, after transport acceptance, and during reconnect pagination. A worker crash, delayed consumer, full connection pool, or old cursor should cause repeated work without repeated user-visible effects.

Retention needs a number. For a 12-minute poll, one hour may cover ordinary reconnects, but the correct value depends on session policy and compliance requirements. Expired rows should be a terminal state, with metrics distinguishing expiry from success.

Observe intent and outcome separately: events created, pending rows, attempts, accepted pushes, acknowledgements, duplicate suppressions, and expirations. Alert on the oldest pending row and retry growth, not only websocket count.

The rejected design is direct publish from the poll-closing request. It is valid for ephemeral hints where loss is acceptable, such as typing indicators. It is the wrong boundary for a poll result that must survive a sleeping laptop.

The practical limitation is operational weight. Per-recipient rows increase storage churn, queue leases need monitoring, and replay can surprise a user if the client cursor is stale. A team with no durable job runner should first narrow the promise to “available on reconnect” and document the retention window, rather than bolt on a best-effort queue and call it guaranteed. The trade-off is deliberate: stronger delivery evidence consumes write capacity and requires cleanup, while direct publish stays simpler but knowingly drops offline users. During a 12-minute poll, I would rather expose a pending notification than silently report success. That choice also makes load tests honest: fan-out pressure appears as queue age and inbox depth, not as an unexplained websocket timeout. Support can then distinguish a missing recipient row, exhausted retries, expired retention, and a client that failed to advance its cursor. Keep the event id in logs, metrics, and acknowledgements; without that shared key, a small notification bug becomes a search across timestamps and connection ids.

There is a useful boundary for teams sharing ownership. The poll service owns event creation and recipient intent. The delivery worker owns retries and transport adapters. The client owns idempotent application and cursor advancement. Storage operations own retention and compaction. Write those responsibilities down before load testing, because a queue that everyone can enqueue but nobody owns will become a second, undocumented database. This design values a recoverable record over a perfect realtime illusion, and it makes each delayed notification explainable.

Decision rule

Use a durable per-recipient inbox and an idempotent fan-out worker when a notification must survive disconnection. This costs more writes and cleanup work, so it is a poor fit for disposable typing hints or telemetry where loss is acceptable. Use direct realtime publish only when the product explicitly accepts loss. Document whether evidence means attempted, transport-accepted, client-acknowledged, or replayable.

References

Top comments (0)