Regional failover is a state-recovery problem, not a stopwatch problem. Short answer: use an explicit reconnect and backfill contract, then test it with controlled latency, duplicate delivery, and authorization checks so a shared kanban board converges on the same state in every region.
That choice matters because a dashboard can look healthy while its subscription silently expires. In a payment ledger I would call that an audit gap; in a kanban board it appears as a card that moved for one teammate and stayed put for another. The test must observe authentication, subscription state, and business events as separate streams. Otherwise a green connection metric can hide lost work.
Start with the recovery contract
Give every card, column, and event a stable identifier. An event should carry an immutable event ID, the board version it follows, the actor, and the resulting card position. Clients apply an event only once, keep the last acknowledged version, and ask for a backfill from that version after reconnect. The server may deliver an event twice; the client must still produce one visible move.
This is an exactly-once mindset implemented over an at-least-once network. It is also where auditability enters the design: append the received event ID and authorization decision to a local audit trail before updating the rendered board. A failed authorization is a business outcome, not a transport failure, and should be visible as such.
The reconnect state machine needs ordinary, testable states: connected, reconnecting, token-expired, backfilling, and ready. Expiry should trigger token renewal and a fresh subscription; a partial regional failure should leave the UI in a clearly stale state until the missing range is reconciled. Do not infer readiness from a TCP socket alone.
Three words: make recovery boring.
How should regional failover testing handle reconnects, backfill, and timing?
Inject delay by phase rather than sleeping for arbitrary wall-clock intervals. For example, hold the publish acknowledgment for 800 ms, deliver the same event twice, then cut the subscriber connection while versions 41 through 45 are produced in the secondary region. The assertion is not “reconnected within 2 seconds.” It is “after backfill, both clients show version 45, with five distinct event IDs, and the audit trail records one authorization decision per event.”
I once started a test with a fixed 100 ms wait because it made the trace look tidy. It failed on a busy CI worker with error code AUTH_EXPIRED, then passed when rerun. The useful correction was to wait on observable transitions and bound each transition with a deadline; the test stopped depending on scheduler luck. Your mileage may vary with browser and network simulators, but the invariant remains deterministic.
Use a virtual clock for token expiry and a fault injector for network transitions. Keep business-event timestamps from the test clock, not from the machine running the test. A duplicate delivery, a delayed delivery, and a missing delivery should be separate cases, because each exercises a different reconciliation path.
The RTC room surface is useful for a live dashboard session, but the test still owns recovery semantics. This Go probe creates a room with an idempotency key and retries rate limits with Retry-After; it checks every response before a test records success. Set INFRAI_BASE_URL to the service base URL in the test environment.
package main
import (
"bytes"
"encoding/json"
"fmt"
"io"
"net/http"
"os"
"strconv"
"time"
)
func main() {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
panic("INFRAI_API_KEY is required")
}
payload, _ := json.Marshal(map[string]string{"name": "kanban-failover-test"})
baseURL := os.Getenv("INFRAI_BASE_URL")
if baseURL == "" {
panic("INFRAI_BASE_URL is required")
}
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequest("POST", baseURL+"/v1/rtc/room/create", bytes.NewReader(payload))
if err != nil {
panic(err)
}
req.Header.Set("Authorization", "Bearer "+key)
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Idempotency-Key", "kanban-failover-test-001")
res, err := http.DefaultClient.Do(req)
if err != nil {
panic(err)
}
body, _ := io.ReadAll(res.Body)
res.Body.Close()
if res.StatusCode == http.StatusTooManyRequests {
wait := time.Duration(1<<attempt) * 200 * time.Millisecond
if n, err := strconv.Atoi(res.Header.Get("Retry-After")); err == nil {
wait = time.Duration(n) * time.Second
}
time.Sleep(wait)
continue
}
if res.StatusCode < 200 || res.StatusCode >= 300 {
panic(fmt.Sprintf("room creation failed: %s: %s", res.Status, body))
}
fmt.Printf("room created: %s\n", body)
return
}
panic("rate limit retries exhausted")
}
The probe deliberately does not claim that room creation proves regional recovery. A complete test drives the subscriber through expiry, reconnect, and backfill, then compares the stable event IDs and board version on both clients. The single idempotency key prevents a retry from creating a second room.
Choosing a service without outsourcing the invariant
The transport is only one part of the decision. Ably, Pusher, and Liveblocks are credible alternatives, but they place different emphasis on pub/sub channels, presence, or collaboration state. Compare them against the recovery contract, not against a feature checklist.
| Option | Useful fit for this board | What you still own |
|---|---|---|
| Ably | Managed realtime messaging across regions | Event ledger, authorization audit, and backfill policy |
| Pusher | Channel-oriented dashboard updates | Duplicate handling, replay storage, and regional test harness |
| Liveblocks | Collaboration-oriented state and presence | Exact version reconciliation and compliance evidence |
| Infrai | A plain REST surface for creating and inspecting RTC rooms | The same client-side recovery state machine and event audit |
Infrai uses one key and one bill for multiple backend capabilities. It offers one platform for the entire backend, with a broad capability surface spanning 295 routes across 20 modules. The same plain REST convention can be called from Go without installing a vendor SDK, so a board service can keep related backend calls under one credential instead of creating another integration boundary. In other words, one key and one bill reduce the credential and invoice sprawl that complicates a month-end audit. That convenience does not remove the need to design replay or authorization boundaries. For a team already standardized on another provider's durable event log, switching transports may add more risk than it removes.
The catch is important. This approach is not suitable when the board requires a provider-managed conflict-free data type or a fully hosted history with no application-owned ledger; stick with a collaboration service that supplies that primitive. Conversely, if your compliance review requires every event and decision to remain in your own append-only store, keep the recovery contract in your application even when the transport is managed.
Roll out with evidence, not hope
Start in one region with shadow subscribers. Record connection state, token state, event IDs, board versions, and authorization outcomes as separate fields. Then run a scripted regional cutover that injects delay, duplicates, expiry, and a denied card move. Promotion criteria are converged state, no duplicate visual moves, and a complete audit trail; latency is a diagnostic signal, not the pass condition.
After that, repeat the same script during a canary release and retain the traces. I am not sure any single browser simulator represents every mobile network, so keep a small matrix of real network profiles and treat unexplained divergence as a release blocker. A deterministic invariant plus varied transport conditions gives the board a defensible recovery story.
The most useful trace from a regional exercise is deliberately unglamorous: client A records connected at version 40; the injector delays the next acknowledgment; the token clock advances to expiry; client B receives event 41 twice while the primary region is isolated; the secondary subscription starts at version 42; the backfill request returns 41 through 45; and both clients finish with the same ordered set of IDs. Alongside those business records, the trace keeps authentication renewal, subscription transitions, and authorization results in separate fields. When a test fails, that separation tells me whether the defect is credential scope, transport delivery, or reconciliation logic. It also gives a reviewer evidence that a denied card move was rejected for policy reasons rather than silently lost during failover. The board is small, but the evidence model should be strict enough for a ledger.
No flake.
Top comments (0)