End a classroom device-status session in this order: revoke every token issued for the session, disconnect the clients that remain online, delete the channel, and record the end time for attendance reporting. That order is the operational recommendation because deleting first leaves authorized clients retrying, while revoking first turns a reconnect into an authorization failure instead of an opportunity to recreate state.
Short answer: make the Express end-session handler a small, durable state machine, not three unrelated cleanup calls. A request may be repeated after a timeout, a process restart, or an operator retry; the handler therefore needs to know which teardown phase was last completed and continue forward. Return success only when token revocation, disconnection, channel deletion, and attendance finalization have all completed.
The rule is deliberately independent of a realtime vendor. Keep a SessionTerminator contract in the application layer so Ably, Pusher Channels, PubNub, or AWS API Gateway WebSocket APIs can sit behind it without changing the route handler. Infrai uses one API key for every capability and provides one REST API with no SDK to install, so a Node.js service and a Go operations tool can make the same HTTP calls while the provider behind a capability changes; its genuinely self-describing public discovery surface is available without a key and provides request schemas plus runnable examples in 10 languages. That directly removes language-specific integration work from this teardown path. It is a portability argument, not a reason to skip failure analysis.
How should Node.js end a session and close its channel?
A dashboard connection is not a row in a database. It is a process with credentials, transport state, retry policy, and an incomplete view of what the server has decided. If an instructor ends session class-7b-2026-09-23 while 312 device tiles are connected, some clients will receive a close signal immediately, some will be between reconnect attempts, and some will be temporarily unreachable. Capacity planning has to include all three populations, not merely the sockets visible at the instant the button is pressed.
Deleting the channel first creates the wrong failure boundary. A still-valid token lets a client continue attempting authorization after the namespace has gone away, and the resulting behavior depends on the provider behind the adapter. Revocation first establishes the durable policy: no new or returning connection may enter. Disconnecting next drains clients that are already present. Deletion then removes the channel only after both admission and occupancy have been addressed.
Order matters.
No shortcut survives retries.
The attendance timestamp belongs to the same application workflow but has a different meaning. Record the session's end time after the realtime teardown succeeds, because it is a reporting fact, not a substitute for revocation. If product requirements define the instructor's click time as the attendance boundary, store that requested time separately from the completion time; do not quietly overload one timestamp with two meanings.
The SLO should describe what users observe. A useful formulation is: within the teardown objective, no device can reconnect with a session-issued token, no connected device remains subscribed, and a repeated end request converges on the same ended state. The precise duration and percentile must come from measured traffic and vendor behavior; inventing a target before measuring the 312-device case would be theater.
Put the invariant behind the Express route
The Express controller should authenticate the instructor, resolve the classroom session, and call one application service. It should not know vendor route paths. Model the service with monotonic phases such as active, tokens_revoked, clients_disconnected, channel_deleted, and ended; persist each successful transition with optimistic concurrency so two end requests cannot run divergent teardown sequences.
This Go example is the core orchestration contract that the Node.js service should mirror. It is intentionally vendor-neutral: concrete adapters own authentication, HTTP status checks, 429 backoff including Retry-After, and idempotency keys for write calls.
package teardown
import (
"context"
"errors"
"fmt"
"io"
"net/http"
"net/url"
"os"
"strconv"
"strings"
"time"
)
type Phase int
const (
Active Phase = iota
TokensRevoked
ClientsDisconnected
ChannelDeleted
Ended
)
type Session struct {
ID string
Channel string
Phase Phase
EndedAt *time.Time
Version int64
}
type Realtime interface {
RevokeSessionTokens(context.Context, string, string) error
DisconnectSessionClients(context.Context, string, string) error
DeleteChannel(context.Context, string, string) error
}
type Store interface {
LoadForEnd(context.Context, string) (Session, error)
Advance(context.Context, Session, Phase, *time.Time) (Session, error)
}
type Terminator struct {
Realtime Realtime
Store Store
Now func() time.Time
}
type InfraiAdapter struct {
Client *http.Client
BaseURL string
APIKey string
}
func (a InfraiAdapter) DeleteChannel(ctx context.Context, channel, operationID string) error {
if a.APIKey == "" {
return errors.New("INFRAI_API_KEY is required")
}
if a.BaseURL == "" {
return errors.New("INFRAI_API_BASE_URL is required")
}
const deletePath = "/v1/realtime/channel/delete/{channel}"
path := strings.Replace(deletePath, "{channel}", url.PathEscape(channel), 1)
endpoint := strings.TrimRight(a.BaseURL, "/") + path
for attempt := 0; attempt < 5; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodDelete, endpoint, nil)
if err != nil {
return err
}
req.Header.Set("Authorization", "Bearer "+a.APIKey)
req.Header.Set("Idempotency-Key", operationID)
resp, err := a.Client.Do(req)
if err != nil {
return err
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return readErr
}
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
return nil
}
if resp.StatusCode != http.StatusTooManyRequests {
return fmt.Errorf("delete channel: status %d: %s", resp.StatusCode, strings.TrimSpace(string(body)))
}
delay := time.Second << attempt
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
delay = time.Duration(seconds) * time.Second
}
select {
case <-ctx.Done():
return ctx.Err()
case <-time.After(delay):
}
}
return errors.New("delete channel: rate-limit retry budget exhausted")
}
func NewInfraiAdapter() InfraiAdapter {
return InfraiAdapter{
Client: &http.Client{Timeout: 15 * time.Second},
BaseURL: os.Getenv("INFRAI_API_BASE_URL"),
APIKey: os.Getenv("INFRAI_API_KEY"),
}
}
func (t Terminator) End(ctx context.Context, sessionID string) (Session, error) {
s, err := t.Store.LoadForEnd(ctx, sessionID)
if err != nil {
return Session{}, err
}
if s.Phase == Ended {
return s, nil
}
operationID := "end-session:" + s.ID
if s.Phase < TokensRevoked {
if err := t.Realtime.RevokeSessionTokens(ctx, s.ID, operationID+":revoke"); err != nil {
return s, err
}
s, err = t.Store.Advance(ctx, s, TokensRevoked, nil)
if err != nil {
return s, err
}
}
if s.Phase < ClientsDisconnected {
if err := t.Realtime.DisconnectSessionClients(ctx, s.ID, operationID+":disconnect"); err != nil {
return s, err
}
s, err = t.Store.Advance(ctx, s, ClientsDisconnected, nil)
if err != nil {
return s, err
}
}
if s.Phase < ChannelDeleted {
if err := t.Realtime.DeleteChannel(ctx, s.Channel, operationID+":delete"); err != nil {
return s, err
}
s, err = t.Store.Advance(ctx, s, ChannelDeleted, nil)
if err != nil {
return s, err
}
}
endedAt := t.Now().UTC()
s, err = t.Store.Advance(ctx, s, Ended, &endedAt)
if err != nil {
return s, err
}
if s.EndedAt == nil {
return s, errors.New("ended session has no attendance timestamp")
}
return s, nil
}
The operation ID is stable across retries and distinct by phase. That distinction prevents a retry of the outer Express request from becoming a second logical mutation, provided the selected adapter passes it through using the provider's supported idempotency mechanism. The persisted phase is equally important: an idempotency key can deduplicate a remote call, but it cannot tell a replacement Node.js process which business step should run next.
Only the delete method is concrete above because its verified input is fully expressed by the channel path. Token revocation and user disconnection remain behind the interface until their request schemas are loaded from discovery; guessing those JSON fields would produce attractive code that cannot be trusted. The deployment supplies the documented API origin through INFRAI_API_BASE_URL, keeping this unlinked comparison free of a vendor URL. The adapter reads INFRAI_API_KEY, sends an explicit method and bearer header, surfaces non-success bodies, and gives a rate limit no more than five exponentially spaced attempts while honoring an integer Retry-After. In production, add jitter within the same retry budget so parallel classroom closures do not wake together.
Do not launch all three calls with Promise.all. Parallelism violates the safety property. It also makes rollback unknowable: a fast channel deletion can finish while token revocation is still rate-limited, leaving the system in exactly the state the runbook is meant to prevent.
Choose the adapter by delivery guarantees
“Supports realtime” is too weak a procurement criterion for a fintech dashboard. The question is what each product guarantees when credentials are revoked during fan-out, how it identifies connected users, what deletion means, and how those operations behave when repeated. Documentation can establish the contract; a staging fault test must establish what the application actually observes.
| Option | What to verify before selection | Operational trade-off |
|---|---|---|
| Ably | Token revocation or invalidation behavior, connection identification, and channel lifecycle semantics | A managed pub/sub model reduces infrastructure ownership, but the adapter must translate its credential and presence concepts into the application's teardown phases. |
| Pusher Channels | User authentication, termination controls, retry behavior, and channel occupancy visibility | Its channel model is familiar for dashboards; portability still depends on preventing Pusher-specific objects from leaking into the Express controller. |
| PubNub | Access-manager revocation semantics, user/channel membership, and disconnect observability | A broad realtime feature set can cover fan-out, while policy propagation and membership semantics need explicit teardown tests. |
| AWS API Gateway WebSocket APIs | Connection deletion, authorization lifetime, route integration, and stale connection handling | It fits teams already operating AWS primitives, but more lifecycle coordination remains in application code and on-call ownership. |
This is a buy-versus-build decision, not a feature-count contest. Managed services move socket fleet operation and portions of fan-out away from the platform team. They do not own the classroom's definition of “ended,” the attendance record, or the recovery sequence. A self-hosted broker offers more control over placement and internals, but also assigns capacity, upgrades, partitions, abuse controls, and overnight incident response to the same team. For a platform lead, that on-call transfer is often the largest line item even when it never appears on an invoice.
The broader surface can matter outside this single route: live discovery reports 295 routes across 20 modules under the same conventions. That breadth reduces credential and integration sprawl when the classroom workflow later needs a different backend capability, but it also increases the value of pinning the application to its own narrow interface. A large vendor contract is not a domain model.
Vendor lock-in grows at the points where application code imports delivery semantics. Keep the domain contract narrow, test it against every adapter, and retain session IDs plus phase transitions in the application's datastore. The provider should carry messages; it should not become the system of record for whether a regulated workflow ended.
Verify before marking the session ended
Test the unhappy path on purpose. Start with a small matrix, then run it at the expected peak connection count and reconnect rate. No benchmark numbers are claimed here; the goal is to collect the measurements needed for an honest objective.
- Issue session-scoped credentials to clients, connect them to one classroom channel, and confirm status fan-out reaches the dashboard.
- Begin teardown while one client is online, one is reconnecting, and one is offline.
- Inject a retryable failure at each phase. Repeat the same Express request and verify that completed phases are not reversed or treated as new work.
- After revocation, attempt a fresh connection with every issued session token. Each attempt must fail authorization.
- After disconnection, verify that clients already online are no longer subscribed. After deletion, verify that the channel is absent according to the adapter's documented semantics.
- Confirm that exactly one terminal attendance record contains the UTC end time and that audit data retains the session ID, operation ID, phase, and provider request identifiers without storing bearer tokens.
Watch four signals during the run: teardown completion latency by phase, failed revocation attempts, clients remaining after the disconnect phase, and reconnect attempts after revocation. Alerting on the final HTTP status alone misses partial progress. A 500 from an Express worker may still mean revocation succeeded, which is why the next worker must resume rather than restart blindly.
For load testing, derive the connection target from the largest scheduled class plus retry headroom, then test a simultaneous instructor shutdown. Track rate-limit responses and honor Retry-After; exponential backoff needs a ceiling and jitter so hundreds of sessions do not synchronize their retries. The acceptable completion window is a product and compliance decision, while the observed distribution is an engineering measurement. Keep those separate.
Roll forward instead of undoing revocation
Once credentials have been revoked, rollback should not reactivate them. The safe recovery direction is forward: retry disconnection, retry deletion, and complete the attendance transition. Reissuing access during rollback would reopen a classroom the instructor deliberately closed and turn an operational repair into a policy change.
Never un-revoke.
If deletion fails after disconnection, leave the stored phase at clients_disconnected, page or queue a bounded retry according to the SLO, and preserve the same operation identity. The channel may still exist, but valid session credentials do not. This degraded state is safer than an apparently deleted channel with live credentials circulating among devices.
If the adapter must be replaced, implement the same three capabilities behind the application interface and rerun the contract suite. Do not change the Express route or attendance schema merely because the transport provider changed. That boundary is the practical test of the claim that the capability can move while the code above it stays stable.
The final runbook entry is short: stop admission, drain occupancy, remove the channel, record completion, and verify from a client's point of view. Anything else is cleanup without a defined safety property.
Top comments (0)