DEV Community

Libme
Libme

Posted on

JWTs Fail Validation for Some Requests but Not Others: Clock Skew, `nbf`, and How Much Leeway Is Safe

If your JWT verification fails intermittently — the same client, the same key, most requests fine and a few rejected with "jwt not active" or "signature has expired" — the signature is almost certainly valid and the clocks are disagreeing. The issuer stamps iat/nbf/exp using its own wall clock, the validator compares them to its own, and any offset between the two machines turns a narrow slice of every token's lifetime into a dead zone. The fix is a small, explicit leeway on the validator plus monitoring on the actual clock offset, not a bigger token lifetime.

Why does this only break for some requests?

Three claims in a JWT are compared against the validator's clock:

  • nbf (not before) — reject if now < nbf
  • exp (expiration) — reject if now ≥ exp
  • iat (issued at) — informational in RFC 7519, but some libraries and maxAge checks compare it to now

Suppose the auth service's clock is 20 seconds ahead of the API server's. A token minted with nbf = issue time arrives at the API 50 ms later, but the API thinks that moment is still 20 seconds in the future, so it rejects the token as not-yet-valid. Twenty seconds later the same token verifies fine. That's the entire mechanism: a fixed offset creates a fixed-width window of failure at each end of the token's life, and whether a given request lands in that window depends on timing. Low traffic hides it; a login burst makes it look random.

The same offset in the other direction eats the tail instead: the validator's clock runs ahead and tokens get rejected as expired early. With short-lived access tokens (60–300 seconds is common machine-to-machine), 20 seconds is a large fraction of the usable life.

Takeaway: a constant clock offset produces intermittent failures, because only requests that land inside the offset window at either end of the token's life are affected.

Which error message points at which claim?

The messages are library-specific, which is why this gets misdiagnosed as a key or algorithm problem. A rough map (library APIs as of mid-2026):

Symptom / error text Claim at fault Real cause Fix
NotBeforeError: jwt not active (jsonwebtoken) nbf Validator clock behind issuer Leeway on validator; fix NTP
JWTClaimValidationFailed with claim: "nbf" (jose) nbf Same clockTolerance
ImmatureSignatureError: The token is not yet valid (nbf) (PyJWT) nbf Same leeway=
token is not valid yet (golang-jwt) nbf Same jwt.WithLeeway
TokenExpiredError: jwt expired right after issuance exp Validator clock ahead, or exp written in milliseconds-divided-wrong Check units first, then leeway
Token accepted for an absurdly long time exp exp written in milliseconds Fix issuer units
An iat-in-the-future complaint, or maxAge failing immediately iat Issuer clock ahead Leeway; don't gate on iat

Before touching leeway, rule out the unit bug — it looks like skew and no amount of tolerance fixes it. NumericDate in JWT is seconds since the epoch, not milliseconds, and Date.now() + 3600 is a token that expires 3.6 seconds after issuance. In any language whose epoch helper disagrees with the spec, this is the most common self-inflicted version of this bug.

Decode the claims without verifying and compare them to the validator's own clock:

python3 - "$TOKEN" <<'PY'
import base64, json, sys, time
part = sys.argv[1].split('.')[1]
part += '=' * (-len(part) % 4)          # base64url has no padding
claims = json.loads(base64.urlsafe_b64decode(part))
now = int(time.time())
print("host now:", now)
for k in ("iat", "nbf", "exp"):
    if k in claims:
        v = claims[k]
        unit = "SECONDS" if v < 10_000_000_000 else "MILLISECONDS (bug)"
        print(f"{k}={v} [{unit}] {v - now:+d}s relative to this host")
PY
Enter fullscreen mode Exit fullscreen mode

Run that on the machine doing the verification, not your laptop. A 10-digit value is seconds; 13 digits means the issuer is writing milliseconds.

Takeaway: if exp - iat doesn't equal your configured token lifetime in seconds, you have a units bug, not a clock bug.

How do I measure the actual offset between two machines?

Don't guess from log timestamps — those are written by the same skewed clocks. Two cheap measurements:

# 1. What does the issuer think the time is? (HTTP Date header, RFC 9110 format)
curl -sI https://auth.example.com/.well-known/openid-configuration | grep -i '^date:'
date -u

# 2. What does the local NTP client think of its own clock?
chronyc tracking | grep -E 'System time|Last offset'
timedatectl status | grep -E 'synchronized|NTP service'
Enter fullscreen mode Exit fullscreen mode

chronyc tracking prints something like System time : 0.000018 seconds fast of NTP time — that's the number you want, per host. If timedatectl says System clock synchronized: no, stop reading and fix that; everything downstream is noise.

Two environment-specific notes that cost people hours. First, containers do not have their own clock — they read the host's, so running an NTP daemon inside a container is useless and installing one is a smell. Fix the host, or on managed platforms open a ticket. Second, a laptop or VM that suspends and resumes can come back with a clock seconds-to-minutes off until the next sync; if "it only happens on my machine after lunch," that's this.

For ongoing visibility, you want the offset as a metric rather than a thing you SSH in to check. If you're already self-hosting Prometheus, node_exporter exposes node_timex_offset_seconds and a node_timex_sync_status flag, which is enough to alert on drift before it reaches token-lifetime scale — the gap is that it only covers hosts you run an exporter on, so serverless and managed runtimes stay invisible. If you'd rather not build the dashboard, Datadog ships an NTP check that reports a per-host clock offset out of the box, with the usual caveat that it's priced per host and you still have to set the alert threshold yourself. On the host side, chrony is the NTP client that handles step-vs-slew correctly after long sleeps and exposes the offset for scraping, though it assumes you actually control the OS.

Takeaway: measure the offset per host with an NTP client, and alert on it as a number — log timestamps can't reveal skew because they're produced by the clock in question.

How much leeway is safe to allow?

RFC 7519 explicitly allows "some small leeway, usually no more than a few minutes, to account for clock skew." In practice, 30 seconds is a good default and 60 seconds is defensible; anything beyond a couple of minutes means you've stopped compensating for skew and started extending token lifetime.

// Node, jsonwebtoken
const payload = jwt.verify(token, publicKey, {
  algorithms: ['RS256'],              // always pin; never trust the header's alg
  audience: 'api://orders',
  issuer: 'https://auth.example.com',
  clockTolerance: 30,                 // seconds
});
Enter fullscreen mode Exit fullscreen mode
// Node, jose
const { payload } = await jwtVerify(token, key, {
  issuer: 'https://auth.example.com',
  audience: 'api://orders',
  clockTolerance: '30s',
});
Enter fullscreen mode Exit fullscreen mode
# Python, PyJWT — leeway applies to exp, nbf and iat
payload = jwt.decode(
    token, key, algorithms=["RS256"],
    audience="api://orders", issuer="https://auth.example.com",
    leeway=30,
)
Enter fullscreen mode Exit fullscreen mode
// Go, golang-jwt v5
parser := jwt.NewParser(
    jwt.WithValidMethods([]string{"RS256"}),
    jwt.WithLeeway(30*time.Second),
    jwt.WithIssuer("https://auth.example.com"),
    jwt.WithAudience("api://orders"),
)
Enter fullscreen mode Exit fullscreen mode

Three rules I hold to. Leeway must be much smaller than the token lifetime — 30 seconds of tolerance on a 60-second token means half the window is slop, so raise the lifetime or lower the tolerance. Leeway on exp is a real security cost: a revoked-by-expiry token stays usable for that extra window, so privileged or step-up operations should verify with a tighter tolerance than general API traffic. And leeway should never be your only response — log the computed offset when a clock claim fails by less than the tolerance, so you find out that a host has drifted instead of silently absorbing it until it exceeds the tolerance too.

If your issuer is a managed provider, its clock isn't your problem but your validator's still is: Auth0, Clerk and Keycloak all mint tokens against synchronized clocks, and the offset you're fighting lives on the machine running verify.

Takeaway: set tolerance to 30 seconds, keep it an order of magnitude below your token lifetime, and log near-miss failures so drift surfaces as a signal rather than a shrug.

FAQ

Why does my JWT say "jwt not active" or "token is not valid yet"?
The token's nbf (not before) claim is in the future according to the verifying machine's clock. Either the issuing server's clock is ahead of the validator's, or the issuer deliberately back-dates nbf; allow a 30-second clockTolerance/leeway on the validator and check NTP sync on both hosts.

Is exp in a JWT in seconds or milliseconds?
Seconds. JWT's NumericDate type (RFC 7519) is seconds since 1970-01-01 UTC, so Date.now() in JavaScript must be divided by 1000 — a 13-digit exp is a bug, a 10-digit one is correct.

Should I run NTP inside my Docker container to fix clock skew?
No. Containers share the host kernel's clock, so the container can't have a different time than its host; synchronize the host (chrony or systemd-timesyncd) and leave the container image alone.

Bottom line

Intermittent JWT rejections with a valid signature are a clock problem until proven otherwise, and the first thing to check is units — a milliseconds exp is more common than real drift. Measure the offset per host with an NTP client rather than inferring it from logs, then set an explicit 30-second tolerance on every verifier and keep it well below your token lifetime. Alert on the offset metric so a drifting host shows up before it eats your auth path, and tighten tolerance for privileged operations where an extra 30 seconds of validity actually matters.

Related reading

Top comments (1)

Collapse
 
kashif_manzer profile image
Kashif Manzer •

One nuance on the monitoring section: alert on the sync status flag first and the offset second. After a VM resumes from suspend, the reported offset can look small while the clock is still unsynchronized, so an offset-only alert stays green while verification keeps failing. We alert on sync_status as critical and treat offset above a few seconds as a warning, since the 30 second leeway covers the rest. Logging the computed offset on near misses, like you suggest, is what finally showed us the pattern.