DEV Community

Cover image for My RAG API Never Signs Tokens or Sees Passwords
Rakesh Singh
Rakesh Singh

Posted on

My RAG API Never Signs Tokens or Sees Passwords

Anyone who could reach my FastAPI RAG service could query every document in it.

The quick fix was a /login route: check the password against Postgres, sign a JWT, return it. It would have worked. But ask one question first: if someone stole this API's config and database, who could they become? With a signing secret in the config, the answer is anyone.

So the rule I shipped: FastAPI verifies tokens; it never mints them. The issuer lives in Keycloak.

This post covers the verifier and where it breaks. Not Keycloak setup, not per-workspace permissions.

Why not just add a /login route?

Because it turns your API into an issuer, and an issuer has to hold things worth stealing.

A login route needs two things:

  • A signing secret in the config, so it can sign tokens.
  • Password hashes in the database, so it can check them.

Now replay the break-in. An attacker with the config can mint a valid token for any user ID they like. An attacker with the database has a credential dump and a brute-force target. The API was holding power it never needed to answer a question.

The wrong model is "auth is a feature of my API." It isn't. Your API needs to know who is asking. It doesn't need to be the place that decides it.

What should a FastAPI resource server actually hold?

Picture a passport office and a border guard.

The passport office confirms who you are and issues a passport that's hard to forge. The border guard checks the seal, the expiry, whether it's valid for this country, and lets you through or not. The guard can't print passports and never hears the answers you gave at the office.

In OAuth terms:

  • Keycloak (or Okta, Cognito) is the passport office: the issuer. It runs the login page, stores passwords, signs access tokens.
  • Your FastAPI service is the border guard: the resource server. It receives the token on every request and checks it.

Checking is cheap. Keycloak publishes its public keys at a JWKS URL. The API fetches them once, caches them, and verifies signature, issuer, audience and expiry locally. No call to Keycloak per request.

Replay the same break-in against this design:

  • The config: an issuer URL, an audience, a JWKS URL. All public. Nothing that signs.
  • The database: user rows keyed by Keycloak's user ID (sub). An identifier, not a credential.
  • A stolen token: one user, for about 15 minutes, still limited to what that user can access.

Steal everything inside the API and there is no one to become.

What does the verifier look like?

One dependency. Every protected route uses it.

# auth.py: this file verifies. Nothing in it can sign.
import jwt
from fastapi import Depends, HTTPException
from fastapi.security import HTTPBearer, HTTPAuthorizationCredentials

from .settings import settings  # issuer URL, audience, JWKS URL. No secret.

bearer = HTTPBearer()
jwks = jwt.PyJWKClient(settings.oidc_jwks_url, cache_keys=True)


def get_principal(
    creds: HTTPAuthorizationCredentials = Depends(bearer),
) -> Principal:
    token = creds.credentials
    try:
        key = jwks.get_signing_key_from_jwt(token).key  # Keycloak's PUBLIC key
        claims = jwt.decode(
            token,
            key,
            algorithms=["RS256"],  # pinned; never trust the token's alg header
            audience=settings.oidc_audience,
            issuer=settings.oidc_issuer,
        )
    except jwt.PyJWTError:
        raise HTTPException(status_code=401, detail="invalid token")
    return Principal(user_id=claims["sub"], token_id=claims["jti"])


@app.post("/ask")
async def ask(req: AskRequest, p: Principal = Depends(get_principal)):
    # req carries the question. p carries the identity.
    # Neither is ever derived from the other.
    ...
Enter fullscreen mode Exit fullscreen mode

The line that matters is jwks.get_signing_key_from_jwt(token).key. The only key this service ever holds is Keycloak's public key. It can check a signature. It cannot produce one.

One more detail carries the weight: the route never reads a user ID from the request body, because the body is client input. Identity is built from the verified token before the model or any retrieved document is involved, so neither can change it. That matters the moment you add agents: a search tool takes the user from p, never from arguments the model wrote.

This runs in GroundedDocs, a document Q&A service that answers only with citations from the caller's own documents. Its adversarial suite (30+ cases: missing and expired tokens, wrong workspaces, guessed document IDs) passes only when the system denies, abstains, or leaks nothing.

Code: GroundedDocs on GitHub

When does "verify, don't mint" break?

The rule holds, but it has edges:

  • Revocation lag. A verified token stays valid until it expires. Disable a user in Keycloak and they keep access for up to the token lifetime. Keep access tokens short (about 15 minutes). If you need an instant cut-off, you need introspection or a deny-list, and you've given back the "no call per request" win.
  • Key rotation. Keycloak rotates signing keys. A cached JWKS without the new kid rejects valid tokens. Your client must refetch on an unknown kid. Recent PyJWT versions do; a hand-rolled cache often doesn't.
  • JWKS or Redis is down. Fail closed. No keys means no access, never an anonymous fallback. Same for every auth-adjacent dependency: in my service, if Redis (the rate limiter) is down, /ask and upload return 503 instead of accepting unlimited traffic.
  • Scripts need tokens too. My test scripts use the password grant because a script has no browser. That's test-only. A real client uses authorization code + PKCE.
  • A valid token says who, not what. Mapping sub to workspace membership and scoping every SQL query is a separate layer, and a separate post.

What should you check in the next 20 minutes?

  1. Grep your API's config for anything that can sign (SECRET_KEY, JWT_SECRET, a private key). If it's there, your API is an issuer.
  2. Grep your API's models for password hashing (bcrypt, passlib, argon2). If it's there, you're holding credentials you don't need.
  3. Confirm algorithms is pinned, and both audience and issuer are checked.
  4. Search your routes for any user ID read from the body, query string, or a model's tool arguments.
  5. Block your JWKS URL and stop Redis locally, then hit a protected route. You want 401 or 503. Never 200.

How do you handle the gap between a 15-minute token and needing to cut someone off right now: shorter TTLs, introspection, or a deny-list?

Top comments (0)