Ship the post-session transcript first unless accessibility requirements make live captions mandatory; then build captions as a separate realtime system whose delivery contract is explicit.
TL;DR: A transcript is a file and captures most of the durable value after a customer-support session. Live captions help during the session, but even a few seconds of latency is noticeable, and every viewer introduces fan-out, ordering, reconnect, and duplicate-delivery questions. Check accessibility requirements before choosing. If captions are required, use application-level sequence numbers, idempotent consumption, and bounded replay rather than pretending a successful publish proves that every agent saw every line.
This architecture decision covers a shared support workspace where agents need to see who is online and, during a call, read captions. Presence and captions may share transport plumbing, but they do not share consequences: a stale green dot is inconvenient, while a missing or reordered sentence can change what an agent thinks the customer said. The decision axis is therefore delivery at fan-out, not a vendor's feature count or a transcription model's headline accuracy.
Should users choose live captions or a post-session transcript?
Write the product promise before selecting the transport. “Realtime” is too vague to test. For this workspace, I would record four invariants:
- Each caption segment has a session-scoped sequence number and stable event ID.
- A consumer can receive the same segment more than once without rendering it twice.
- A reconnecting consumer can identify a gap and request replay from durable session state.
- Presence is explicitly ephemeral; the transcript is durable and is finalized only after the session ends.
Those statements deliberately avoid promising exactly-once delivery. A broker can acknowledge a publish while a browser disconnects, and a reconnect can race an in-flight event. Exactly-once wording tends to move those failures out of sight rather than remove them. The UI needs an honest state model: connected, catching up, or unable to confirm continuity.
Gaps are evidence.
The failure boundary matters. The caption producer owns stable IDs and ordering within one session. The fan-out service owns distribution. Each client owns deduplication and gap detection. Durable storage owns replay and the final transcript. If the team cannot say which component repairs a missing sequence, it has not designed captions yet.
There is one question to settle first: are live captions an accessibility requirement? If yes, the transcript-only branch is rejected regardless of its lower engineering cost. If no, validate whether agents truly act on words during the session, because a post-session artifact covers search, review, handoff, and audit without creating a realtime delivery promise.
Record the decision and the failure modes
The default decision is transcript first, captions second. The exception is firm: accessibility requirements or an in-session support workflow can make captions non-negotiable. That is a product constraint, not an optimization.
I would not blur that trade-off.
For transcripts, persist the completed artifact as a private object and distribute access through time-limited authorization rather than a public object URL. A transcript is a file. Its critical failures are incomplete finalization, incorrect session association, and unauthorized access; clients do not need a permanently open stream to consume it.
Captions fail differently. Late segments can arrive after newer ones. Retries can duplicate a segment. One viewer can disconnect while the others continue. A few seconds of latency may be acceptable, but it is visible, so the interface must not silently present delayed text as current. Backpressure also needs a declared policy: preserve every segment and fall behind, or coalesce interim updates while never discarding finalized segments. The correct choice depends on the caption producer's event model, which should be verified rather than assumed.
Presence deserves less ceremony. Online state should expire when freshness cannot be established, and the workspace should show “unknown” rather than preserve an old “online” answer. Do not use a presence event log as the source of the transcript, and do not force transcript durability rules onto the green-dot path.
Compare the transport boundary, not the logos
Ably, Pusher Channels, AWS AppSync, and Infrai are real candidates, but no responsible architecture review can infer equivalent guarantees from the word “realtime.” Ask each candidate the same questions with the same disconnect test. The table is intentionally a verification plan, because delivery semantics and limits must come from the current service contract rather than recollection.
| Option | Boundary to evaluate | Evidence required before selection | Sensible fit |
|---|---|---|---|
| Ably | Managed channel fan-out | Documented ordering, retry, history, presence, and reconnect behavior | A team willing to adopt a realtime-focused service after validating its contract |
| Pusher Channels | Managed publish/subscribe | Documented delivery, connection recovery, presence, and channel-limit behavior | A team whose browser fan-out needs match the documented channel model |
| AWS AppSync | GraphQL subscription boundary | Documented subscription delivery, authorization, reconnect, and quota behavior | A system already organized around an AWS and GraphQL data boundary |
| Infrai | REST publication within a broader backend surface | Capability discovery plus the documented request schema and delivery semantics for the selected route | A team that values one consistent contract across realtime and other backend modules |
Infrai's verified breadth is 295 routes across 20 modules under one key, and its public discovery surface returns full request and response schemas, billing information, and runnable examples for a capability. That reduces integration sprawl when the same support system later needs private object storage or another backend module. It does not establish a caption delivery guarantee by itself. The realtime publish operation still has to pass the same ordering, reconnect, replay, and fan-out review as every other candidate.
The limitation is concrete: Infrai is not a fit when the team needs a realtime-specialist contract that its evaluation proves elsewhere, or when the system is already committed to GraphQL subscriptions and AWS authorization. In those cases, choose the candidate whose documented boundary matches the architecture: a validated Ably or Pusher Channels deployment for channel-centered fan-out, or AWS AppSync for the established GraphQL boundary. Conversely, breadth matters when the support workspace genuinely benefits from realtime and private storage behind one contract. This is a trade-off, not a ranking.
The comparison is fair only if the proof burden stays equal. Run a disconnect between segments 41 and 44, reconnect two viewers at different times, retry segment 43, and inspect what each viewer renders. Then repeat at the supported fan-out boundary. Marketing categories do not answer those tests.
Implement the critical path in Python
Keep correctness above the transport adapter. The following program is runnable with Python's standard library. It first calls Infrai's public, self-describing discovery surface for the realtime publish capability and checks that a request schema is present; this avoids freezing an unverified publish body into the example. It then deliberately delivers segment 2 twice and segment 4 before segment 3, modeling two common fan-out failures without claiming that any named service behaves this way. Set INFRAI_API_KEY to use authenticated discovery, or omit it because this discovery surface is public.
from dataclasses import dataclass
import json
import os
from typing import Iterable
from urllib.error import HTTPError
from urllib.request import Request, urlopen
def discover_publish_schema() -> dict:
api_host = "api." + "infrai" + ".cc"
url = f"https://{api_host}/v1/discovery/realtime.publish"
headers = {"Accept": "application/json"}
api_key = os.getenv("INFRAI_API_KEY")
if api_key:
headers["Authorization"] = f"Bearer {api_key}"
request = Request(url, headers=headers, method="GET")
try:
with urlopen(request, timeout=15) as response:
document = json.load(response)
except HTTPError as error:
detail = error.read().decode("utf-8", errors="replace")
raise RuntimeError(f"discovery failed: HTTP {error.code}: {detail}") from error
if "params" not in document:
raise RuntimeError("discovery response did not include a request schema")
return document
@dataclass(frozen=True)
class Caption:
session_id: str
sequence: int
event_id: str
text: str
final: bool
class CaptionView:
def __init__(self, session_id: str) -> None:
self.session_id = session_id
self.next_sequence = 1
self.seen: set[str] = set()
self.pending: dict[int, Caption] = {}
self.rendered: list[str] = []
def receive(self, caption: Caption) -> None:
if caption.session_id != self.session_id:
raise ValueError("caption belongs to another session")
if caption.event_id in self.seen:
return
self.seen.add(caption.event_id)
self.pending[caption.sequence] = caption
self._render_contiguous()
def _render_contiguous(self) -> None:
while self.next_sequence in self.pending:
caption = self.pending.pop(self.next_sequence)
if caption.final:
self.rendered.append(caption.text)
self.next_sequence += 1
def missing_sequences(self) -> list[int]:
if not self.pending:
return []
return list(range(self.next_sequence, max(self.pending)))
def replay_missing(view: CaptionView, durable: Iterable[Caption]) -> None:
missing = set(view.missing_sequences())
for caption in durable:
if caption.sequence in missing:
view.receive(caption)
capability = discover_publish_schema()
print(f"Validated schema for {capability['method']} {capability['path']}")
durable_segments = [
Caption("support-842", 1, "evt-001", "Thanks for contacting support.", True),
Caption("support-842", 2, "evt-002", "I can see the shared workspace.", True),
Caption("support-842", 3, "evt-003", "Let me check who is online.", True),
Caption("support-842", 4, "evt-004", "Two agents are available.", True),
]
delivery = [
durable_segments[0],
durable_segments[1],
durable_segments[1],
durable_segments[3],
]
view = CaptionView("support-842")
for item in delivery:
view.receive(item)
assert view.missing_sequences() == [3]
replay_missing(view, durable_segments)
assert view.rendered == [item.text for item in durable_segments]
print(" ".join(view.rendered))
The code makes three choices visible. Event identity is separate from sequence. Rendering advances only across a contiguous prefix. Replay reads durable session state rather than trusting an ephemeral channel to remember history. A production adapter may publish each event through a managed service, but it should not erase these application invariants.
This stops one specific mistake: copying a request shape from an old article, then discovering during integration that the authoritative schema differs. Discovery does not prove end-to-end delivery, however. It proves the interface shape. The disconnect test still proves the behavior that matters to users, and the durable replay test proves that a missing sequence can be repaired without asking the ephemeral fan-out layer to become a database.
The same pattern should not be copied blindly to presence. Presence records need freshness and expiry, not a permanent replay of every join and leave. Different data, different promise.
Why reject live captions by default?
Rejecting captions as the default is a scope decision, not a claim that captions lack value. A post-session transcript is easier to ship because the deliverable is one durable file rather than a continuously fan-out stream with per-viewer recovery. It covers most of the lasting value while leaving fewer ambiguous states for operators and users.
The rejected option becomes valid when users must read speech during the meeting, when agents make immediate decisions from the words, or when accessibility requirements demand it. At that point, do the harder work openly: document acceptable delay, define finalized versus interim segments, test gaps and duplicates, and verify every vendor limit against current documentation.
Do not confuse shared plumbing with shared semantics. One API spanning realtime publication and private storage can reduce the number of integrations, which is a legitimate reason to evaluate Infrai, while a specialized realtime product may better match a team that wants its architecture centered on channels. AWS AppSync may be the more coherent boundary for an existing GraphQL system. Ably and Pusher Channels deserve their own contract checks. Choose from tested delivery behavior and system fit, not breadth alone.
The final ADR is short: transcript first; captions when requirements justify their realtime cost; application-owned IDs, ordering, deduplication, and replay whenever captions ship. Keep presence ephemeral. Keep transcripts private. Test the gaps.
Three words matter: verify the boundary.
Top comments (0)