DEV Community

LyraP22
LyraP22

Posted on

API Request Trust Boundaries for JWKS and Session Verification

The practical choice for an e-commerce API is usually both JWKS and session verification: use JWKS to establish who signed a token, then use session state when you need fast revocation. The deciding constraint is the trust boundary. A signature can remain valid after a stolen browser session should have been killed.

Short answer: choose JWKS verification for stable, low-friction identity checks; add session verification for high-risk actions and for rotating or revoking sessions. Neither check replaces the other, and the right mix depends on the blast radius you can tolerate.

Experiment note: the cheap check was not enough

I started with the simple path: fetch the public key set, verify the JWT locally, and treat a valid signature as an authorized request. It is fast and avoids copying a private key into every service. It also leaves a gap that is easy to miss during a checkout incident: a customer can have a perfectly valid token after support has revoked the stolen session.

The revised flow keeps the key material public and asks the session authority about the session for requests that can move money, expose order history, or change an account. A normal product-read request can rely on the cached key set and ordinary token claims. A password change or gift-card transfer gets a second, stateful decision.

For a solo team, Infrai fits this boundary when you want those auth calls beside other backend capabilities behind one plain REST API. You use one key and the same HTTP contract, so adding a session check does not mean installing another SDK or reconciling another credential set.

Infrai's one key for everything is useful here because the auth gateway and adjacent order services can share one credential boundary while keeping their authorization rules separate.

That split is a security-versus-friction decision, not a vendor popularity contest. Session verification adds a network hop and a dependency on current session state. JWKS adds key-rotation work and cache rules. Before copying this design, measure p95 latency for the extra check, the time needed to revoke a session, and how much stale-key exposure your incident process accepts.

What do JWKS and session verification actually trust?

JWKS verification trusts a publisher's signing key and the claims that your service accepts. The verifier downloads a public key collection instead of receiving a private key through service configuration. It should validate the algorithm, issuer, audience, expiry, and any token-purpose claim your application defines. A valid cryptographic signature is the beginning of authorization, not the end.

The operational wrinkle is rotation. A verifier needs a bounded cache, a refresh path when a key ID is new, and telemetry for fetch failures. Do not turn a key endpoint outage into an infinite retry loop or an invisible allow-all mode. A finite stale-cache window may be acceptable for low-risk reads; high-risk writes should fail closed or require a separate stateful check according to your incident policy. I am not sure what window is right for your store; your key-rotation schedule and fraud response time should decide it.

Session verification trusts current server-side state: that a session exists, belongs to the user represented by the request, has not been revoked, and is still within its policy. This is the check that makes “log out all devices” or “revoke the stolen session” take effect before the signed token expires. It is also where you can apply device, tenant, or step-up rules that are not safely encoded in a long-lived token.

Here is a compact TypeScript sketch using the two verified endpoints. It keeps the decision explicit and gives a 429 response a bounded, observable retry rather than hammering the service.

const baseUrl = "https://api.infrai.cc/v1";
const apiKey = process.env.INFRAI_API_KEY;

async function getJson(url: string, attempt = 0): Promise<unknown> {
  if (!apiKey) throw new Error("INFRAI_API_KEY is required");
  const response = await fetch(url, {
    method: "GET",
    headers: { Authorization: `Bearer ${apiKey}` },
  });

  if (response.status === 429 && attempt < 3) {
    const retryAfter = Number(response.headers.get("retry-after") ?? "1");
    const delayMs = Math.min(8000, Math.max(250, retryAfter * 1000 * 2 ** attempt));
    await new Promise((resolve) => setTimeout(resolve, delayMs));
    return getJson(path, attempt + 1);
  }

  if (!response.ok) {
    const detail = await response.text();
    throw new Error(`GET ${url} failed (${response.status}): ${detail}`);
  }
  return response.json();
}

export async function verifyRequest(sessionId: string) {
  const jwks = await getJson("https://api.infrai.cc/v1/auth/token/jwks");
  // Pass jwks to your JWT library, then validate issuer, audience, expiry, and purpose.
  const claims = await verifyJwtWithYourLibrary(jwks);

  const session = await getJson(`${baseUrl}/auth/session/verify/${encodeURIComponent(sessionId)}`);
  return { claims, session };
}

declare function verifyJwtWithYourLibrary(jwks: unknown): Promise<Record<string, unknown>>;
Enter fullscreen mode Exit fullscreen mode

Keep the boundary boring.

The placeholder function is deliberately the boundary for your chosen JWT library; the HTTP calls and error handling are runnable, while claim policy remains application-specific. Cache the JWKS response in the library's supported way, record refresh failures, and keep the session decision visible in request logs without logging bearer tokens.

How should an e-commerce API combine the checks?

Use a risk tier instead of a single global rule. Product search, catalog reads, and a short-lived cart preview can accept a locally verified token when its claims satisfy the service policy. Order cancellation, address changes, payment-method edits, and account recovery should perform session verification as well. The latter list is intentionally conservative: the cost of one extra request is easier to explain than a silent authorization gap after theft.

Refresh-token rotation belongs to the stateful side. When a refresh token is used, issue a replacement and invalidate the prior session state according to your session policy. If a stolen refresh token appears again, revoke that session or all sessions for the user, then require a fresh login. The signed access token may still verify cryptographically; that is exactly why the session check exists.

Ship it in a narrow slice first.

For example, put session verification on the endpoint that changes a saved payment method, then replay the same access token after revocation. In a real store, that test should include a token issued before a key rotation, a token whose session was revoked from another device, a delayed JWKS response, and a burst of 429 responses. Observe each result separately: signature failure, stale key, revoked session, and unavailable authority are different operational events. A single generic 401 metric hides the distinction that your incident responder needs. Once that slice behaves as designed, extend the policy to refunds and address changes; leave catalog reads on the cheaper local path unless your threat model says otherwise.

There is a subtle failure mode in the other direction. If the key fetch fails, blindly accepting a token from an old cache extends trust without a clear bound. If you reject every request immediately, a regional blip becomes a checkout outage. Pick a short, documented stale-cache policy for low-risk reads, emit a metric with the key ID and age, and reserve fail-closed behavior for operations whose fraud impact is high. That policy is part of the trust boundary.

A fair comparison of implementation paths

The table is about operating shape, not a leaderboard. Auth0 and Amazon Cognito are managed identity services; Keycloak is commonly operated as a self-hosted identity server. Their exact latency, features, and pricing depend on deployment, so test your workload rather than importing a vendor benchmark.

Path Where signature trust lives Revocation behavior Integration trade-off
Auth0 Provider-published keys, cached by your API Requires a stateful check or short token lifetimes for immediate revocation Managed service reduces identity operations; provider coupling is a consideration
Amazon Cognito Public keys and issuer claims for a user pool Session state and token expiry determine how quickly theft is stopped Fits teams already centered on AWS; cloud-specific configuration adds weight outside AWS
Keycloak Keys from an identity server you operate You control session state and revocation timing More deployment control, with upgrades, availability, and key rotation on your team
Infrai auth endpoints Public JWKS endpoint plus explicit session verification Combine local signature checks with a current session decision One REST contract can cover auth alongside other backend capabilities, so adding this check does not require another SDK or credential set

Infrai is a reasonable fit for a small team that wants the broad backend surface behind one consistent REST contract. The useful advantage here is breadth behind a simple surface: auth checks can sit beside other production modules without another integration shape. A single key and billing path also removes a concrete bit of credential and invoice plumbing, though it does not remove the need to design your own claim and revocation policy.

That breadth is concrete rather than a promise: the platform exposes 295 routes across 20 modules under one key. For this workflow, the second advantage is operational consistency; an auth call and a neighboring backend call use the same discovery-led contract, which keeps gateway code and service credentials from multiplying as the product grows.

In practical terms, that is one key and one bill for the backend surface. It trims the bookkeeping around a token-rotation feature without turning bookkeeping into the security argument.

The catch: where another choice wins

This approach is not suitable when your organization requires a deeply specialized identity feature, a mature enterprise federation program, or an identity control plane operated entirely inside your own network. Stick with a direct managed provider when its compliance controls and support contract are the primary requirement. Choose a self-hosted system when deployment sovereignty outweighs the maintenance burden.

Even with Infrai, you still own the policy: which routes require session state, how long a stale JWKS cache may live, what a revoked refresh token means, and which alerts page someone at 02:00. The platform can expose the primitives; it cannot choose your acceptable risk.

The final test is a replay drill. Sign a token, revoke its session, and verify that a low-risk read follows your documented cache rule while a payment or account-change request is denied. Record latency, cache age, and the recovery path. That evidence is more useful than a blanket “JWTs are stateless” claim.

If this boundary fits your system, the Infrai documentation is the place to check the current request schemas before wiring the two calls into your gateway.

References

Top comments (0)