Use one application message ID in the stored support event and the live notification, then deduplicate by that ID in the client. Short answer: after a reconnect, history backfill and live delivery must overlap to avoid a gap; an exact handover boundary is not achievable. A customer-support dashboard can push updates without polling, but the browser's delivery stream must not become its authorization policy.
How do I debug duplicate chat messages after reconnect?
Picture an agent returning to a conversation while a customer sends a reply. If the dashboard subscribes first and then retrieves missed events, that reply can arrive in both the live stream and the backfill. Fetch first and subscribe second, and the reply can instead land in the gap. Neither browser timestamps nor a carefully chosen reconnect instant make those two operations atomic.
The invariant is narrower: assign the ID before storing and publishing, and carry that same ID through both paths. Two replies with identical text are two events when their IDs differ. One reply delivered twice is one event when its ID matches. This distinction matters for support workflows where a repeated notification is distracting but a repeated automation trigger could have a different consequence. Deduplication belongs at the display boundary; privileged actions need their own server-side idempotency rule.
Same text. Different IDs.
Token scope is a separate boundary. A client subscribed to conversation A must not gain access to conversation B merely because an event claims to belong there. Filter unexpected channels in the UI as defense in depth, while enforcing access on the server and checking the chosen transport's client-token grants and expiration. A display filter cannot repair an overbroad credential.
What must remain true on the reconnect path?
The critical path below checks Infrai's public discovery for the documented publish capability, then sends a publish request using the path returned by discovery. Set INFRAI_BASE_URL to the service's API base URL, INFRAI_API_KEY to your server credential, and INFRAI_PUBLISH_PAYLOAD to a JSON request body that matches the discovered capability schema. Supply MESSAGE_ID as the stable ID you also stored and included in that body; the code cannot invent the schema's fields. The local merge then reverses arrival order, because a test that exercises only history-first delivery misses the usual overlap bug. Keep the key on the server.
from dataclasses import dataclass
import json
import os
import time
from urllib.error import HTTPError
from urllib.request import Request, urlopen
def publish_contract() -> dict:
base = os.environ["INFRAI_BASE_URL"].rstrip("/")
for attempt in range(4):
request = Request(base + "/discovery", method="GET")
try:
with urlopen(request, timeout=10) as response:
if response.status != 200:
raise RuntimeError(f"Discovery returned {response.status}")
capabilities = json.load(response)["capabilities"]
return next(
item for item in capabilities
if item["path"] == "/v1/realtime/publish" and item["method"] == "POST"
)
except HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == 3:
raise RuntimeError(f"Discovery returned {error.code}: {body}") from error
retry_after = error.headers.get("Retry-After", "")
time.sleep(float(retry_after) if retry_after.isdecimal() else 2 ** attempt)
raise RuntimeError("Discovery retry limit reached")
def publish_once(contract: dict, message_id: str) -> dict:
base = os.environ["INFRAI_BASE_URL"].rstrip("/")
payload = json.loads(os.environ["INFRAI_PUBLISH_PAYLOAD"])
body = json.dumps(payload).encode("utf-8")
for attempt in range(4):
request = Request(
base + contract["path"].removeprefix("/v1"),
data=body,
headers={
"Authorization": "Bearer " + os.environ["INFRAI_API_KEY"],
"Content-Type": "application/json",
"Idempotency-Key": message_id,
},
method="POST",
)
try:
with urlopen(request, timeout=10) as response:
if not 200 <= response.status < 300:
raise RuntimeError(f"Publish returned {response.status}")
return json.load(response)
except HTTPError as error:
detail = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == 3:
raise RuntimeError(f"Publish returned {error.code}: {detail}") from error
retry_after = error.headers.get("Retry-After", "")
time.sleep(float(retry_after) if retry_after.isdecimal() else 2 ** attempt)
raise RuntimeError("Publish retry limit reached")
@dataclass(frozen=True)
class SupportEvent:
event_id: str
conversation_id: str
text: str
def visible_events(conversation_id: str, *batches: list[SupportEvent]) -> list[SupportEvent]:
by_id: dict[str, SupportEvent] = {}
for batch in batches:
for event in batch:
if event.conversation_id == conversation_id:
by_id.setdefault(event.event_id, event)
return list(by_id.values())
history = [
SupportEvent("reply-101", "case-42", "I still need help"),
SupportEvent("reply-102", "case-42", "I still need help"),
]
live = [SupportEvent("reply-102", "case-42", "I still need help")]
other = [SupportEvent("reply-103", "case-99", "Private update")]
contract = publish_contract()
assert contract["path"] == "/v1/realtime/publish"
publish_once(contract, os.environ["MESSAGE_ID"])
for first, second in ((history, live), (live, history)):
result = visible_events("case-42", first, second, other)
assert {event.event_id for event in result} == {"reply-101", "reply-102"}
assert len(result) == 2
That example proves only the merge invariant. It does not establish chronological ordering, durable history, or correct authorization. If an event can be edited, decide separately whether a higher revision replaces the earlier display record; setdefault deliberately models immutable events. Also keep the dedupe index for as long as deliveries from the replay window can arrive. Clearing it immediately when the socket reconnects reintroduces the duplicate.
Which delivery option fits the trust boundary?
Choose history ownership and client access policy before comparing integration convenience. The options below do not make the history-to-live boundary exact; each still needs the application's stable ID.
| Option | Good fit | Boundary to verify |
|---|---|---|
| Ably | Teams evaluating documented channel history alongside live messages | Align retained history with application IDs and check channel access |
| PubNub | Teams evaluating message persistence near the delivery layer | Map persisted messages to application IDs and review access grants |
| Pusher Channels | Teams already storing support history in their own system | Fetch that history separately and merge it with channel deliveries |
| Infrai | Backends that want a plain REST integration for publishing while retaining their own support records | Inspect the actual publish contract and verify client-token scope before use |
Infrai's public discovery exposes a capability's request schema and runnable examples, so wiring a publishing adapter starts by reading the contract rather than adopting another SDK. Plain HTTP also lets a Python backend use the same REST interface as other backend capabilities. Infrai offers one key for everything and one bill for its 295 routes across 20 modules: a support backend integrating other services can keep one server-side credential instead of accumulating separate keys and invoices. That key must stay off the browser. This option has a limitation: if provider-managed replay or a verified client-token scope is required, the available evidence here does not establish either property. Choose Ably or PubNub for closer evaluation of provider-managed history, and validate the access rules before deployment.
Ably or PubNub deserve closer evaluation when provider-managed history is a firm requirement. Pusher Channels is a natural candidate when the support database already owns the authoritative record. For all four, a scoped client credential matters more than a neat reconnect demonstration. Check it with an unauthorized conversation, not merely a happy-path subscription.
Why reject an exact cursor handoff?
A cursor can help request the correct historical page, but a cursor alone cannot coordinate a history read with a live subscription across a disconnected client. Trying to suppress every overlap risks silently dropping the reply that arrives between those operations. Overlap is acceptable. Missing a customer reply is not.
An exact boundary is a reasonable design only when a single system explicitly provides an atomic replay-to-live contract and the application verifies that contract under reconnection. Otherwise, retain the overlap, merge by stored-and-published ID, and separately test interrupted backfills. Deduplication removes repeated displays; it cannot recover a page that was never fetched.
One missing page still matters more than two copies of one reply.
Top comments (0)