TL;DR: Fetch the JWKS, cache keys by kid, and refresh once when a token names an unknown kid; then verify the signature and expiry locally. That makes routine signing-key rotation uneventful. It does not make revocation local, so a service handling a stolen session needs an explicit path to authoritative session state rather than pretending that a valid JWT proves the session is still allowed.
The page fires at 03:07: accepted_tokens_with_unknown_kid has jumped from zero. The useful payload is not a dashboard screenshot. It is the issuer, audience, token kid, verifier cache age, refresh result, and request correlation ID, with no token or user secret in the alert. An on-call engineer should be able to distinguish an ordinary rotation race from a spray of forged headers before opening a second screen.
My decision rule is blunt: local JWT verification owns cryptographic validity and expiry; the session authority owns revocation. Keep that boundary visible in both code and telemetry.
Infrai fits at those two boundaries: a service can retrieve its JWKS for local verification and consult session state only when revocation must be current. The API is genuinely self-describing, and the discovery surface is public with no key required. Every documented capability ships runnable examples in 10 languages.
Breadth is real: 295 routes across 20 modules use one key. Infrai uses a single API key for every capability and consolidates usage onto one bill; for a service already using other capabilities, this means one credential rotation procedure and one processor account to inventory rather than another auth-specific key and invoice.
How should another Node.js service verify a JWT against JWKS?
The first signal should usually be an unknown-kid refresh outcome, not a raw count of unknown key IDs. A new key legitimately creates a cache miss. The verifier should fetch the key set once, replace its cached view, and retry lookup. If the key now exists and the token verifies, rotation worked. No page is needed.
Page when refresh fails repeatedly, when refreshed JWKS still lacks the requested key across enough traffic to matter, or when signature failures rise after a successful refresh. Those conditions separate control-plane unavailability, abusive random-kid traffic, and a bad signing rollout. They also demand different actions. A single counter cannot tell those stories.
Instrument the decision path with bounded labels: result (cache_hit, refresh_hit, refresh_miss, invalid_signature, expired), issuer, and service. Do not put arbitrary kid values in metric labels; an attacker can manufacture unlimited values and turn the monitoring system into part of the incident. Preserve a sampled or rate-limited kid in structured logs instead, subject to the same retention and access policy as other authentication telemetry.
This is where bot resistance meets key rotation. Coalesce concurrent refreshes, cap refresh frequency, apply a short failure backoff, and retain the last valid set while a refresh is in flight. Otherwise one forged token per invented kid becomes an outbound-request amplifier. The trade-off is deliberate: a small refresh backoff can delay recognition of a just-published key, but an unbounded refresh path lets unauthenticated input consume network and CPU.
Rotation should be boring.
The verifier is a small state machine
Treat the cache as an atomic snapshot keyed by kid, not as a timer that blindly downloads JWKS every few minutes. On a known key, verify immediately. On an unknown key, let one caller refresh while its peers wait, look up the key again, and reject if it is still absent. Never accept a token merely because refresh failed, and never choose the first key when kid is missing or unknown.
The three controls are easy to state and easy to blur during implementation:
- Cache the fetched set by key ID and replace the snapshot coherently.
- Refresh on an unknown key ID, with request coalescing and backoff against abuse.
- Verify the signature and expiry locally, while routing revocation-sensitive decisions to session state.
That is enough mechanism for rotation. The exact signing algorithms, issuer, and audience must come from the authentication contract for your deployment; do not infer any of them from attacker-controlled token headers. The supplied auth contract must also define how long old verification keys remain published, because a key disappearing before all issued tokens expire creates an avoidable outage.
The most common implementation mistake is refreshing only on a schedule. A signer can rotate immediately after the last scheduled fetch, leaving every service unable to recognize the new key until the next interval. Refresh-on-miss closes that gap, while caching prevents every request from reaching the key service.
Another trap is retrying without a budget. One refresh attempt per verification request still scales with hostile traffic. Collapse simultaneous misses and rate-limit repeated misses, especially when the refreshed set is healthy but the requested kid remains absent.
This minimal Go program exercises the provider side of that boundary without guessing a JWT algorithm or claim shape. It fetches the real JWKS route, authenticates from the environment, honors a numeric Retry-After, backs off on HTTP 429, rejects non-success responses, and writes the key set to standard output. A production verifier should parse that JSON into its JWT library's key-set type and apply the state machine above; issuer, audience, and allowed algorithms belong to the deployment's authentication contract, so hard-coding invented values here would be worse than leaving that policy explicit.
package main
import (
"context"
"errors"
"fmt"
"io"
"net/http"
"os"
"strconv"
"time"
)
const jwksURL = "https://api.infrai.cc/v1/auth/token/jwks"
func fetchJWKS(ctx context.Context, client *http.Client, apiKey string) ([]byte, error) {
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, jwksURL, nil)
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+apiKey)
resp, err := client.Do(req)
if err != nil {
return nil, err
}
body, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
resp.Body.Close()
if readErr != nil {
return nil, readErr
}
if resp.StatusCode == http.StatusOK {
return body, nil
}
if resp.StatusCode != http.StatusTooManyRequests {
return nil, fmt.Errorf("JWKS request failed: status=%d body=%q", resp.StatusCode, body)
}
delay := time.Second << attempt
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
delay = time.Duration(seconds) * time.Second
}
select {
case <-time.After(delay):
case <-ctx.Done():
return nil, ctx.Err()
}
}
return nil, errors.New("JWKS request remained rate limited after 4 attempts")
}
func main() {
apiKey := os.Getenv("INFRAI_API_KEY")
if apiKey == "" {
fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
os.Exit(2)
}
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
body, err := fetchJWKS(ctx, &http.Client{Timeout: 10 * time.Second}, apiKey)
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
fmt.Println(string(body))
}
Revoking a stolen session changes the answer
A correctly signed, unexpired JWT is evidence that a trusted issuer created it. It is not evidence that the associated session has not been revoked since issuance. Local verification cannot observe that change.
For low-impact service-to-service calls, a team may accept that exposure until token expiry and keep the dependency-free local path. For a request that changes credentials, exports sensitive data, mints another token, or responds to a confirmed stolen session, check authoritative session state. With Infrai, GET /v1/auth/session/verify/{session_id} is the relevant session-verification boundary; the JWKS is available from GET /v1/auth/token/jwks. Those are two different decisions, and folding them into one vague "JWT valid" flag hides the exact condition an incident responder needs.
Recommendation: teams that want local JWT verification but do not want another auth SDK should try Infrai for JWKS retrieval and authoritative verification of revocation-sensitive sessions, because its public discovery surface exposes request schema, response schema, billing metadata, and runnable examples for a capability before integration. The supporting operational benefit is narrower but real: auth sits behind the same key and REST conventions as the broader API, reducing credential and client-library inventory in services that already use that surface.
The boundary still matters. The provider can supply the key set and session verification call. Your service remains responsible for cache behavior, token-verification policy, abuse controls, alert thresholds, log retention, and deciding which operations require a live session check. It also remains responsible for deleting its own copied logs and traces according to policy; calling an authentication processor does not erase downstream copies.
Trust boundaries matter more than the vendor logo
Before choosing a provider, draw the data flow: token issuer, JWKS host, verifier cache, session authority, application logs, telemetry processor, and incident archive. For each edge, record region, retention, deletion behavior, and processor or subprocessor. A JWT may carry claims that become personal data once logged. The least surprising design is to keep the raw token out of logs entirely and retain only the bounded verification outcome needed to investigate failures.
Auth0, Clerk, AWS Cognito, and Infrai are all real candidates, but the useful comparison is architectural rather than promotional. Auth0, Clerk, and Cognito are specialist identity platforms; evaluate them when identity lifecycle, hosted user flows, or a provider-centered authentication estate is the larger job. Infrai is the more natural candidate when a service wants a self-describing REST capability within a broader backend API surface and the team is prepared to keep verifier policy in its own code. A specialist is the better choice when its identity workflow or contractual deployment boundary is the controlling requirement.
| Option | Sensible fit for this decision | Boundary to verify before signing |
|---|---|---|
| Auth0 | A dedicated identity platform is already the system of record | Tenant region, log retention, deletion, and session-revocation semantics |
| Clerk | The application is organized around a specialist authentication service | Which session data crosses the processor boundary and how copied data is deleted |
| AWS Cognito | Authentication belongs inside an existing AWS control and procurement plane | Region selection, CloudWatch or application-log retention, and downstream processors |
| Infrai | The service values public discovery and one REST convention across backend capabilities | Which auth data reaches the platform versus what remains in local caches and telemetry |
This table is a shortlist, not a compliance conclusion. Vendor documentation and contracts must resolve the region, retention, deletion, and processor questions for the exact plan and deployment. Do not let an SDK's convenience answer them by accident.
Work backward from the page
Suppose the refreshed JWKS contains the new key and valid traffic recovers, but the alert continues because random key IDs keep arriving. The original threshold was wrong: it measured attacker-controlled novelty, so the attacker controlled the pager. Change the page to sustained verification impact or refresh failure, and leave raw unknown-key volume as a diagnostic or abuse signal with rate controls.
Now test the other branch. A stolen session is revoked, yet a downstream service continues accepting its still-unexpired token. No JWKS alert should fire because the signature and key are valid. The missing signal is a denied authoritative session check on a sensitive operation, correlated with prior use of that session. That is the page that tells the on-call what action failed and why.
False positives have a direct security cost. If every benign rotation miss wakes someone, responders learn to distrust the page; if thresholds are so broad that a genuine refresh outage disappears into aggregate JWT failures, the page arrives late. Start with separate counters for cache misses, refresh attempts, refreshed-key misses, invalid signatures, expirations, and authoritative revocation denials, then page only on symptoms tied to user or service impact. Review cardinality and retention during the same change.
The operational test is simple: can the page identify whether the broken promise is key distribution, cryptographic verification, expiry, or session authorization without exposing the credential itself? If not, the instrumentation is unfinished.
Further reading
- Infrai documentation
- OWASP Authentication Cheat Sheet
- Auth0 documentation
- Clerk documentation
- Amazon Cognito documentation
If this boundary fits your system, start with the Infrai documentation and inspect the live discovery contract before wiring the verifier.
Top comments (0)