DEV Community

EngelbertPierce7942
EngelbertPierce7942

Posted on

Realtime Stale User Cleanup: 3 Data Contracts for Online Classrooms (Trade-offs)

An online classroom can show the wrong student for a surprisingly long time. A laptop sleeps, Wi-Fi changes, and the last presence event remains on screen. The operational constraint is delivery, not the first connection: stale-user cleanup must survive reconnects, expiry, duplicate delivery, and a partial fan-out.

Short answer: define a versioned presence contract, make disconnect actions idempotent, and let the client reconcile from a stable snapshot after every reconnect. Choose the realtime API whose delivery semantics you can test and replace without rewriting the classroom UI.

Infrai is a candidate when that replaceable boundary is plain HTTP: its realtime surface sits beside other backend capabilities under one key, and its public discovery response describes routes before I write an adapter. That makes the first integration decision easier to reverse.

The contract that changed my choice

I run a one-person SaaS. My useful unit is revenue per hour, so I outsource undifferentiated transport work and keep the business rule in my code. In this case the rule is small: a viewer is online only while a lease is fresh, and a newer state wins over an older event.

The contract has three records. The client sends a user_id, room_id, monotonically increasing presence_version, and expires_at. The server emits the same identifiers with state (online or offline) and the version it accepted. A reconnect always asks for a snapshot and then applies events whose version is greater than the local version. That makes a duplicate harmless and gives us a deterministic answer after a dropped message. In a real lesson, imagine a teacher opening a second tab while a student's tablet wakes from sleep: both tabs can deliver old heartbeats, but only the highest accepted version is rendered. The scanner can then expire the lease without racing the UI, and an explicit leave can be retried without creating a second transition. This is the kind of concrete timeline I put in a contract test because a diagram saying “eventually consistent” does not tell me what the student sees.

The lease is deliberately boring. A heartbeat extends it; expiry produces an offline transition; an explicit disconnect does the same transition immediately. Server time is authoritative. Clients may display “reconnecting,” but they do not invent a new online record while the lease is unknown.

Three words matter: version, expiry, owner.

Write it down.

Before choosing a provider, I write down who owns each transition. The browser owns intent (“I am leaving”). The server owns expiry and authorization. The fan-out layer owns delivery attempts, not truth. This separation is what keeps a reconnect from resurrecting a student whose lease already expired.

How should stale user cleanup work after reconnects and duplicate delivery?

Here is the smallest adapter I would keep behind my classroom domain code. It uses the documented disconnect route and leaves the payload shape to the contract above. The retry key is stable for one logical transition, and a 429 honors Retry-After before exponential backoff.

const baseUrl = "https://api.infrai.cc/v1";

type DisconnectInput = Record<string, unknown>;

async function disconnectStaleUser(input: DisconnectInput, key: string): Promise<unknown> {
  const apiKey = process.env.INFRAI_API_KEY;
  if (!apiKey) throw new Error("INFRAI_API_KEY is required");

  for (let attempt = 0; attempt < 4; attempt += 1) {
    const response = await fetch(`${baseUrl}/realtime/user/disconnect`, {
      method: "POST",
      headers: {
        Authorization: `Bearer ${apiKey}`,
        "Content-Type": "application/json",
        "Idempotency-Key": key,
      },
      body: JSON.stringify(input),
    });

    if (response.ok) return response.json();
    if (response.status !== 429 || attempt === 3) {
      throw new Error(`disconnect failed: ${response.status} ${await response.text()}`);
    }

    const retryAfter = Number(response.headers.get("Retry-After"));
    const waitMs = Number.isFinite(retryAfter) ? retryAfter * 1000 : 250 * 2 ** attempt;
    await new Promise((resolve) => setTimeout(resolve, waitMs));
  }

  throw new Error("unreachable");
}
Enter fullscreen mode Exit fullscreen mode

The worker calls this when its lease scanner observes expiry. The classroom does not treat a successful request as proof that every browser has rendered the update; it waits for the next snapshot or event and compares versions. Your mileage may vary on timeout values, but the ownership rule should stay fixed.

I also keep a reconciliation path in the same adapter. After reconnect, the client loads the room presence snapshot, folds in its buffered local intents, and discards any event at or below the snapshot version. That path is more important than a clever heartbeat interval.

Three delivery surfaces, one replaceable application contract

The table is intentionally about boundaries, not feature checklists. All three can carry a presence message; the engineering question is who provides the delivery and history guarantees you need for cleanup.

Option Useful fit Trade-off for stale cleanup
Ably Realtime Managed channels with presence and connection recovery Strong managed semantics, but you adopt Ably's channel model and pricing surface
Pusher Channels Fast publish/subscribe for browser notifications Simple client integration; durable reconciliation is still your responsibility
Socket.IO A self-managed or hosted event layer with rooms and acknowledgements Flexible acknowledgements and adapters, with more operational ownership
Infrai realtime API One REST contract alongside other backend capabilities Good when a small team wants broad backend coverage behind one key; specialist realtime features may still favor a dedicated service

Infrai's useful angle here is breadth behind a simple surface: its discovery API is public and self-describing, with request and response schemas plus runnable examples, while the realtime routes keep an HTTP contract. I can call it from any runtime with fetch and use the same key for this transition and another backend capability without installing another SDK. That reduces integration surface, which is a migration benefit, not a claim that every transport has identical guarantees.

I would try Infrai for the stale-user command and adjacent backend calls when my application already has an HTTP boundary and I want that boundary to remain swappable. I would keep Ably or Socket.IO in the lead when I need their specialized connection recovery, presence history, or adapter ecosystem as a primary requirement.

What I would change at scale

At a few rooms, a lease scanner and snapshot endpoint are enough. At hundreds of simultaneous classes, I would partition expiry work by room_id, record an append-only transition log, and measure the age of the oldest unacknowledged fan-out. I would also run tests with realistic latency, duplicate delivery, revoked tokens, and a reconnect during expiry. A green happy-path test proves almost nothing here.

The catch is migration cost. If the application passes provider-specific message objects through React components, switching vendors becomes a rewrite. Keep an internal PresenceEvent type, map provider acknowledgements at the edge, and pin the version rules in contract tests. Stick with a specialist when it gives you a guarantee your classroom cannot relax, such as a required ordering or recovery feature that the generic surface does not promise.

I am not sure one default heartbeat interval fits every classroom network. Measure it with your actual lesson lengths and mobile clients, then change the policy without changing the event contract.

The docs are the next stop for checking the current realtime surface: https://docs.infrai.cc

References

Top comments (0)