The bug report that keeps repeating itself
Scan enough MCP issue trackers and you'll notice the same shape of bug filed against a dozen different clients and servers: Copilot CLI, Gemini CLI, Cursor, Open WebUI, Codex, and several homegrown gateways. The symptom is always some version of "it worked, then it silently stopped working, and the only fix was to log in again."
The root cause is almost never "OAuth is broken." It's that two independent clocks — the access token's expiry and the MCP session's lifetime — get treated as one clock, and the recovery logic for one gets applied to the other.
That distinction is the entire subject of this post.
Scope: what this covers and what it doesn't
This post is narrowly about keeping a connection alive after the initial handshake — refreshing tokens and re-initializing sessions without forcing a user back through a browser. It assumes you already have PKCE, discovery, and initial token validation working (covered in an earlier post in this series). It does not cover setting up dynamic client registration or scopes — that's a separate problem with a separate fix.
Phase 1: Understand that you have two clocks, not one
OAuth token expiry and session expiry are unrelated. An expired token gives you 401 and needs a refresh; an expired session gives you 404 and needs a re-initialize. The session itself, per spec, is a correlation and state handle. It carries no lease, no expiry timestamp and no promise about how long it stays valid. Nothing in the lifecycle specification obliges a server to keep it alive for any period at all.
So you're managing two independent failure modes with two independent recoveries, and most of the bugs above come from collapsing them into one.
Phase 2: Correctly classify the failure before you react
The spec's intent is clean: If a server has ended a session, requests carrying the old id must receive HTTP 404. The client's required response is to discard the id and send a fresh initialize. A 401 means something entirely different — the MCP server acts as an OAuth resource server. The client sends a bearer token on each request, the server validates expiry, scope and audience, and an invalid token gets a 401.
In practice, plenty of servers don't respect this. One real-world rmcp-based server was found answering an unrecognized session ID with 401 Unauthorized: Session not found instead of the spec-mandated 404 — and the consequence is severe: OAuth MCP clients (claude.ai, Claude Code) treat a 401 as an invalid credential: they drop the token and demand an interactive re-auth. A perfectly valid, unexpired token gets discarded because the server picked the wrong status code for a session problem.
If you're building the server, this is the single highest-leverage fix in this whole post: return 404 for unknown/expired sessions, 401 (with WWW-Authenticate) only for actual credential problems, and never conflate the two.
If you're building the client, don't treat every 401 as "burn the token and force login." Distinguish an auth failure from a stale session by checking whether WWW-Authenticate is present, and only discard credentials on a genuine 401 with that header.
Phase 3: Refresh the token — but check the provider actually gave you one
This is where a surprising number of teams get stuck before refresh logic even runs. Several identity providers don't issue a refresh token unless you explicitly ask:
- Google: Google OAuth issues a refresh token only when the authorization request includes access_type=offline (typically with prompt=consent). This is a Google-specific query parameter, not part of OAuth scope negotiation. Google does not honor offline_access as a scope, so requesting it via the existing scopes config has no effect.
- Meta/Facebook-backed servers: one production report found refresh_token is advertised but never issued. Meta's OAuth returns a long-lived access token with no refresh token. So once the client decides the session must be refreshed, there is nothing to refresh with, and the only possible outcome is the message above.
- Custom auth servers: a similar bug appeared in the wild where offline_access is required but the engine never tells the client to request it, so every session hard-expires after 7 days.
The fix at the client level: check for a refresh_token field immediately after the initial grant, and fail loudly in development if it's missing — don't wait until a user hits it in production three hours into a session.
Phase 4: Don't let a successful token refresh leave a stale session behind
This is the gotcha most teams never see coming, because the token refresh itself succeeds. A Streamable HTTP transport ties a session ID to the credentials that created it. If you refresh the bearer token but keep reusing the old Mcp-Session-Id, the server sees a mismatch between the session and the token attached to it.
This exact failure was reported against a production MCP integration: after Cursor performs a silent OAuth access-token refresh (near token expiry / refresh flow), the result was Session token does not match bearer token. The only reliable recovery was heavy-handed: Note that only toggling the MCP server off/on, full re-login, or restarting Cursor recovers it. Manual testing confirmed the actual fix is much lighter than a full restart: refreshing the OAuth access token and calling initialize again with the new bearer produces a new Mcp-Session-Id and restores tool calls. Cursor's silent refresh does the token part but skips session re-initialization.
Rule of thumb: a successful token refresh is not a completed recovery on its own. After any refresh, re-run initialize with the new bearer token and accept whatever session ID comes back. Don't assume the old session ID is still valid just because the token now is.
Phase 5: Retry once, with a real circuit breaker — not forever
The last gotcha is about failure discipline. One gateway logged this pattern in production: An MCP server whose credential was never saved is retried forever. Backoff grows to 300 s and then stays there; 18+ failures logged in one session with no circuit breaker, no UI indicator. A 401 caused by a genuinely dead credential is not a transient network blip — retrying it indefinitely just burns cycles and hides the real problem from whoever's on call.
The pattern that actually works, synthesized from several fixes above:
- Catch the 401 vs 404 distinction at the transport layer.
- On 401 with a refresh token available: refresh once, then re-run
initialize, then retry the original call once. - On 404: discard the session ID, re-run
initializewith the existing (still-valid) token, retry once. - If refresh itself fails (invalid_grant, no refresh_token, or repeated 401 after a fresh refresh): stop retrying, surface a "needs reauthorization" state, and don't touch the retry loop again until the credential changes.
- Log which of these four paths fired. When a session breaks in production, that log line is the difference between a five-minute diagnosis and an hour of guessing.
Why this is worth getting right before launch
None of these five phases are exotic — they're each a handful of lines. But they interact, and the bug reports above show that even mature clients (Cursor, Gemini CLI, Copilot CLI) shipped with at least one of these five phases missing or wrong, and the resulting incident always looked the same to the end user: "the MCP server just stopped working."
If you're running MCP servers in production, this class of failure — auth working perfectly for hours, then a session or refresh edge case taking the whole integration down with a vague error — is exactly the kind of incident worth documenting before it happens, not after. Our AI Agent Incident Postmortem & Permission-Scoping Template Pack gives you a structured way to capture root cause, blast radius, and the permission/session boundaries involved, so the next time a token refresh gotcha slips through, you're writing a five-minute postmortem instead of re-deriving the timeline from scratch.
Written with AI assistance and reviewed for accuracy.
Top comments (0)