DEV Community

DimitriReed2158
DimitriReed2158

Posted on

Revoking an Abusive Tenant API Key from a Node.js Admin Endpoint Without Deploys

Revoke the key in a transaction, write an audit event in that same transaction, and make every request check the key state before doing tenant work. That is the safest way to remove an abusive tenant API key from an admin endpoint without a deploy. The important part is auditability: you need to prove which operator revoked which credential, why, and when.

Short answer: keep revocation data in your control plane, expose an authenticated, idempotent admin action, and enforce the state at the request boundary. A cache can make revocation faster, but it must never be the source of truth.

I have been paged for missed jobs and duplicate deliveries. Credential abuse has the same shape: one stale decision gets replayed across a fleet. Treat the revoke path like a runbook, not a button.

What should an admin endpoint record before revoking a tenant API key?

Start with an authorization check that is separate from the tenant key being revoked. An operator session, workload identity, or break-glass credential should carry an explicit permission such as tenant_keys.revoke; accepting the target key as proof of authority defeats the point.

The revoke record needs a stable key identifier, tenant identifier, actor identifier, reason, request identifier, and server timestamp. Store a hash or fingerprint for display, never the secret value. OWASP's Secrets Management Cheat Sheet recommends minimizing secret exposure and maintaining lifecycle controls; a revocation event is the lifecycle boundary that incident responders will inspect later.

Use a state transition with a uniqueness constraint on the key id. Repeating the same request should return the already-revoked state and create no second transition. That makes retries safe when an admin browser times out.

How can you revoke an abusive tenant API key from your own admin endpoint without a deploy?

The endpoint can be a small control-plane handler in front of your existing database. It does not need a release to invalidate a row that request middleware already consults. The example below uses Go for the handler and a generic SQL interface; the same boundary works behind a Node.js service.

package main

import (
    "context"
    "database/sql"
    "encoding/json"
    "net/http"
    "strings"
    "time"
)

type RevokeRequest struct {
    Reason string `json:"reason"`
}

type KeyStore struct{ DB *sql.DB }

func (s KeyStore) Revoke(w http.ResponseWriter, r *http.Request) {
    if !strings.HasPrefix(r.Header.Get("Authorization"), "Bearer ") {
        http.Error(w, "unauthorized", http.StatusUnauthorized)
        return
    }
    if !operatorCanRevoke(r.Context()) {
        http.Error(w, "forbidden", http.StatusForbidden)
        return
    }

    keyID := strings.TrimPrefix(r.URL.Path, "/admin/tenant-keys/")
    if keyID == "" {
        http.Error(w, "missing key id", http.StatusBadRequest)
        return
    }
    var req RevokeRequest
    if err := json.NewDecoder(r.Body).Decode(&req); err != nil || strings.TrimSpace(req.Reason) == "" {
        http.Error(w, "reason required", http.StatusBadRequest)
        return
    }

    tx, err := s.DB.BeginTx(r.Context(), nil)
    if err != nil {
        http.Error(w, "request could not be completed", http.StatusServiceUnavailable)
        return
    }
    defer tx.Rollback()

    now := time.Now().UTC()
    result, err := tx.ExecContext(r.Context(),
        `UPDATE tenant_api_keys SET revoked_at = COALESCE(revoked_at, $1)
         WHERE key_id = $2`, now, keyID)
    if err != nil {
        http.Error(w, "request could not be completed", http.StatusServiceUnavailable)
        return
    }
    if n, _ := result.RowsAffected(); n == 0 {
        http.Error(w, "key not found", http.StatusNotFound)
        return
    }
    _, err = tx.ExecContext(r.Context(),
        `INSERT INTO audit_events (event_type, key_id, actor_id, reason, occurred_at)
         VALUES ('tenant_key_revoked', $1, $2, $3, $4)`, keyID, operatorID(r.Context()), req.Reason, now)
    if err != nil {
        http.Error(w, "request could not be completed", http.StatusServiceUnavailable)
        return
    }
    if err = tx.Commit(); err != nil {
        http.Error(w, "request could not be completed", http.StatusServiceUnavailable)
        return
    }
    w.WriteHeader(http.StatusNoContent)
}

func operatorCanRevoke(context.Context) bool { return true }
func operatorID(context.Context) string      { return "operator-from-session" }
Enter fullscreen mode Exit fullscreen mode

The SQL is deliberately idempotent. COALESCE prevents a retry from moving the original revoke time, while the audit insert should be protected by an idempotency key or a unique (event_type, key_id, request_id) index in a real schema. Do not return the secret, and do not put it in logs.

Where does enforcement happen after the revoke?

The data-plane middleware must reject a revoked key before authorization reaches customer-support data. It should load the key state from the database or a cache with a bounded freshness window, then attach the tenant id to the request only after the state is active. Cache invalidation can be published after commit, but a cache miss must fail closed for sensitive operations.

A common failure mode is checking revocation only in the admin UI. Attackers do not use the UI. Another is checking it in one route and forgetting a background worker. Make the key lookup a shared function used by HTTP handlers, queue consumers, and scheduled jobs.

The incident timeline matters here. An operator sees a spike in ticket exports, identifies tenant t_482, and submits a revoke request at 09:14 UTC. The database commit at 09:14:02 is the authoritative event; a cache invalidation published at 09:14:03 is only a performance hint. Every replica should record the decision it applied, so an audit query can distinguish a request rejected by fresh state from one rejected by an expired session. If a queue consumer fetched work before the revoke, its handler must perform the same key-state check immediately before the external side effect. Otherwise the control plane looks correct while an already-enqueued job continues exporting data, which is exactly the kind of gap that makes a postmortem inconclusive.

Freeze the blast radius.

I once assumed a successful control-plane write meant the incident was over. It wasn't; the useful question was when every data-plane replica observed the state. Your mileage may vary, so measure that propagation delay and alert when it exceeds your stated bound.

How do you verify and roll back a tenant key revocation safely?

Verification should be a small, repeatable runbook. First query the audit event by key id and request id. Then send an authenticated request with the revoked key and confirm a consistent rejection without revealing whether another tenant exists. Finally, exercise a valid key for the same tenant to ensure the blast radius is one credential, not the whole account.

Keep rollback narrow. If an operator revoked the wrong key, restore only that key after a second authorization check and write a compensating tenant_key_unrevoked event; never edit or delete the original event. If abuse is ongoing, do not roll back merely to restore convenience. Issue a new scoped key and continue the investigation.

Decision point Prefer a control-plane revoke Prefer a different mechanism
Need immediate invalidation without shipping code Existing middleware reads shared key state A service that cannot consult shared state needs a gateway-level deny rule
Need tenant-level attribution Key id, tenant id, actor, reason, and request id are stored together Anonymous shared credentials cannot provide reliable attribution
Need offline operation Not suitable when data-plane nodes must run disconnected from the control plane Use short-lived signed credentials with an expiry policy

The catch is operational coupling: if your service has no centralized key state or audit store, this pattern cannot provide immediate, provable revocation. Stick with a gateway deny list or short-lived credentials until that foundation exists. Do not claim a deploy-free revoke when a worker has a hard-coded allow decision.

References

Top comments (1)

Collapse
 
raknaos profile image
Raknaos •

"The commit at 09:14:02 is the authoritative event; the cache invalidation at 09:14:03 is only a performance hint" is the sentence more postmortems need. Treating a published invalidation as the revocation is how already-enqueued jobs keep exporting while the control plane looks clean — and the shared key-lookup function across HTTP handlers, queue consumers and scheduled jobs is exactly the piece most implementations skip, because the background worker with a stale in-memory allow-decision never shows up in the admin UI.

Fail-closed on cache miss plus measuring actual propagation delay turns "eventually consistent" into a number you can alert on, which is the only version of that promise worth anything. Question on the bounded freshness window: how stale do you let cached key state get before the middleware falls back to the database? I've been sizing that window against worst-case blast radius per tenant rather than average latency, which tends to push it smaller than feels comfortable.