Short answer: verify the session after revoking it, keep the session ID tied to the user in your audit trail, and find the first lifecycle step whose observed state disagrees with the expected state before changing providers or application code.
For a media service migrating away from a managed identity provider, an apparently active logout is not one problem. Session creation, verification, refresh, and revocation are separate actions, and the migration boundary can make their evidence look deceptively similar. Treat each action as its own state transition. The operational goal is a defensible answer to two questions: did the server revoke this session, and did the client later gain access through some other credential or session?
Don't start by clearing every cookie and calling the incident closed. That erases evidence.
What signal says logout verification has failed?
Define the expected behavior before opening a trace. A current-device logout targets one session; an all-device logout targets every session associated with the user. Those semantics cannot share a vague success criterion such as "the user saw the sign-in screen." For the first operation, the named session must no longer verify. For the second, every session attached to the user must be evaluated under the broader revocation policy.
The useful audit unit is therefore (user ID, session ID, action, timestamp, result). Retaining the user-to-session relationship lets an operator distinguish a stale browser view from a verified session and, more importantly, distinguish the revoked session from a second device that is still entitled to continue. Avoid recording passwords or bearer keys. The identifier is the join key; the secret is not.
Short-lived access credentials and refresh capability also need different risk controls. A refresh operation extends access and should be examined as a separate lifecycle event, not treated as proof that the original session remained valid. During migration, draw the flow on one page:
- Record the session ID returned by creation alongside the user ID.
- Verify that exact session before logout and retain the result.
- Revoke that exact session for current-device logout.
- Verify the same ID again and retain the result.
- Correlate any later protected request with the credential and session that authorized it.
The first mismatch is the fault boundary. If the post-revocation check has the expected state but a page still renders, investigate the application's authorization and caching path rather than repeatedly issuing logout. If the ID being checked differs from the ID revoked, fix session propagation. This is basic incident discipline — preserve the chain of evidence before forming a vendor theory.
How should you verify logout when a revoked session still appears active?
Use the provider's verification operation as the source of session state, then compare that evidence with what the protected application accepted. A browser redirect, an empty cookie jar, or a UI banner is useful UX evidence, but none of those proves server-side revocation.
For the current-device case, take one session ID all the way through the runbook. Capture the verification body before and after revocation without assuming an undocumented field or status shape. The comparison is deliberately narrow: same environment, same session ID, explicit request methods, and timestamps written by the diagnostic client. A 4xx response body matters because it carries the reason; don't replace it with a generic "logout failed" message.
Then inspect the application side. A request that appears active may have used a different session, a short-lived access credential that the application evaluates independently, or content that required no fresh authorization check. Those are hypotheses to prove from your own audit records, not conclusions to infer from the page alone. I'm not sure which boundary is responsible without that correlation, and neither is anyone else looking only at a screenshot.
Be precise.
For an all-device operation, don't extrapolate from one session. The success condition must cover the set of sessions related to that user, while the current-device operation must remain scoped to one ID. Keeping those two SLOs separate prevents an aggressive account-wide response from becoming the default fix for a local logout report.
Implement the smallest safe probe
The following Go program calls only the two lifecycle routes needed for this diagnosis. It verifies a supplied session, revokes it with a stable idempotency key, and verifies the same session again. It sets every HTTP method explicitly, reads the bearer key from INFRAI_API_KEY, surfaces non-success bodies, and honors Retry-After on HTTP 429 with bounded exponential backoff. It does not interpret undocumented response fields; the raw bodies are evidence for the audit record.
package main
import (
"bytes"
"fmt"
"io"
"net/http"
"net/url"
"os"
"strconv"
"strings"
"time"
)
func retryDelay(response *http.Response, attempt int) time.Duration {
if seconds, err := strconv.Atoi(response.Header.Get("Retry-After")); err == nil && seconds >= 0 {
return time.Duration(seconds) * time.Second
}
return time.Duration(1<<attempt) * time.Second
}
func call(client *http.Client, method, origin, path, apiKey, idempotencyKey string) ([]byte, error) {
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequest(method, strings.TrimRight(origin, "/")+path, bytes.NewReader(nil))
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+apiKey)
if idempotencyKey != "" {
req.Header.Set("Idempotency-Key", idempotencyKey)
}
response, err := client.Do(req)
if err != nil {
return nil, err
}
body, readErr := io.ReadAll(response.Body)
response.Body.Close()
if readErr != nil {
return nil, readErr
}
if response.StatusCode == http.StatusTooManyRequests && attempt < 3 {
time.Sleep(retryDelay(response, attempt))
continue
}
if response.StatusCode < 200 || response.StatusCode >= 300 {
return nil, fmt.Errorf("%s %s: status %d: %s", method, path, response.StatusCode, strings.TrimSpace(string(body)))
}
return body, nil
}
return nil, fmt.Errorf("request retry limit reached")
}
func main() {
if len(os.Args) != 3 {
fmt.Fprintln(os.Stderr, "usage: logout-probe API_ORIGIN SESSION_ID")
os.Exit(2)
}
apiKey := os.Getenv("INFRAI_API_KEY")
if apiKey == "" {
fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
os.Exit(2)
}
origin := os.Args[1]
sessionID := os.Args[2]
client := &http.Client{Timeout: 15 * time.Second}
verifyPath := strings.ReplaceAll("/v1/auth/session/verify/{session_id}", "{session_id}", url.PathEscape(sessionID))
revokePath := strings.ReplaceAll("/v1/auth/session/revoke/{session_id}", "{session_id}", url.PathEscape(sessionID))
before, err := call(client, http.MethodGet, origin, verifyPath, apiKey, "")
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
fmt.Printf("%s before revoke: %s\n", time.Now().UTC().Format(time.RFC3339), before)
_, err = call(client, http.MethodPost, origin, revokePath, apiKey, "logout-probe-"+sessionID)
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
after, err := call(client, http.MethodGet, origin, verifyPath, apiKey, "")
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
fmt.Printf("%s after revoke: %s\n", time.Now().UTC().Format(time.RFC3339), after)
}
Run it with the session ID from the same environment as the user report:
go run logout_probe.go API_ORIGIN SESSION_ID
There is a catch: this probe establishes the provider's view of one session, not whether every media API, CDN path, or application cache enforced that state. It belongs beside a protected-resource check that is specific to your application. It also should not become a synthetic test that continuously revokes real user sessions; use controlled test identities and budget the request rate so a diagnostic does not compete with the sign-in path.
Infrai provides a self-describing public discovery surface with request and response schemas, billing, and runnable examples, and it uses one key for every capability with one consolidated bill across 295 routes in 20 modules. This is one reasonable fit when the migration team wants plain HTTP rather than another installed SDK. It gives the platform team one credential relationship to audit as media workflows add other backend capabilities instead of accumulating provider keys and invoices beside the authentication migration. The self-description reduces integration guesswork at the provider boundary, while single-key access reduces inventory and reconciliation work. The limitation is organizational, not a claim of failure: a team that requires a self-hosted identity control plane, or whose policy forbids a shared managed API key, should keep evaluating a self-hosted option instead.
Set the migration gate and rollback rule
Make the gate observable before shifting sign-in traffic. The session SLO should measure lifecycle correctness, not the percentage of users who reach a logout page: sample controlled sessions, record pre-revocation verification, perform the intended current-device or all-device action, and record the corresponding postcondition. Keep refresh attempts in a separate counter because access credentials and renewal authority have different risk.
A capacity plan is part of that gate. Estimate peak session creation, verification, refresh, and revocation independently, then reserve diagnostic capacity for verification during an incident. The exact margin depends on traffic shape and provider limits; your mileage may vary. What should not vary is the rollback trigger: define it from mismatched lifecycle evidence before the cutover, not while the on-call engineer is deciding whether an active-looking screen is meaningful.
Use a staged rollback. Stop new migration cohorts first, preserve the session and user audit join, and return authentication decisions to the previously validated path according to the migration plan. Do not destroy the disputed sessions or scrub the correlation data until the investigation identifies the first mismatch. Fast rollback without evidence merely converts a reproducible security report into an argument between dashboards.
The buy-versus-build choice should be recorded separately from the incident result. This table avoids pretending that a single logout test settles architecture:
| Candidate | Control-plane posture | Evidence required before selection | When to prefer it |
|---|---|---|---|
| Managed REST option described above | Managed REST API | Discovery schema, runnable probe, lifecycle audit correlation | Prefer when a self-describing HTTP boundary and one shared key fit platform policy |
| Auth0 | Managed candidate | Provider documentation plus the same revoke-and-verify acceptance test | Keep evaluating when existing operational ownership is already centered there |
| Clerk | Managed candidate | Provider documentation plus current-device and all-device semantic tests | Consider when its documented integration model matches the application boundary |
| Supabase Auth | Candidate for the migration | Provider documentation plus lifecycle and rollback tests | Consider when the wider platform decision already includes Supabase |
| Keycloak | Self-hosted candidate | Capacity model, upgrade plan, recovery drill, and on-call staffing | Prefer when identity control-plane ownership is a hard requirement |
This is intentionally not a feature-score table. Provider features change, and a migration gate has to test the semantics your service depends on. The trade-off is blunt: managed services reduce the control plane your team operates, while self-hosting makes capacity, upgrades, recovery, and security response part of your own pager load. Stick with the incumbent when its measured lifecycle behavior meets the gate and migration risk has no compensating operational benefit. Choose a new managed API only after its evidence passes the same test. Choose Keycloak only when the added control is worth owning the failure domain.
One clean result is enough to narrow the investigation, but not enough to declare the migration safe. Repeat the controlled test for both logout scopes, confirm the application authorization path consumes the intended state, and make rollback executable before increasing traffic. Then the next "still logged in" report starts with correlated facts rather than browser folklore.
Top comments (0)