At 02:17, the page says one media tenant is generating an abnormal volume of requests. The on-call needs to stop that credential now, through an authenticated admin operation, without waiting for a deployment and without disabling every key owned by the publisher. The correct design is a runtime revocation record checked on every authenticated request, with an atomic state change, an immutable audit event, and billing attribution preserved under the tenant and key identifiers.
Short answer: expose a tightly authorized admin endpoint that changes one key from active to revoked, record who did it and why, and make the request authenticator reject that key before metering or workload execution. Do not delete the tenant, erase the key row, or ship a denylist in application code. Those shortcuts make an urgent security action depend on a release and leave the next access review with evidence nobody should sign.
The operational target is modest but strict: revocation should converge within a documented interval, fail closed when key state cannot be established, and leave enough durable evidence to distinguish rejected abuse from legitimate billable traffic. Four controls carry most of that load: keyed lookup, atomic revocation, explicit decision ordering, and auditable propagation.
How should an admin endpoint revoke an abusive tenant API key?
Work backward from the alert. A request reaches the media ingestion API with a bearer credential. The gateway or application resolves a one-way digest of that credential to a key record, then obtains the owning tenant, status, and stable key ID. Only an active key proceeds. The usage meter attributes accepted work to that stable identity; a rejection counter records the failed decision separately.
The page at 02:17 is late if it fires only on aggregate request volume. For a publisher platform, the more useful signal combines a per-key request rate with rejected requests, unusual route mix, and the gap between accepted requests and billable work. A tenant-level graph alone can hide which of several newsroom integrations is abusive. An IP-level graph is weaker still: addresses are shared and rotate, while the credential is the authorization object the operator can actually revoke.
Do not guess at a universal threshold. Establish a baseline per workload class, define the false-positive budget, then choose a window and threshold that can meet the abuse-containment SLO without paging on a scheduled content import. A breaking-news spike and a credential leak can look similar in the first minute. The distinction comes from correlated dimensions and a human-readable runbook, not a magically precise requests-per-second number.
No deploy. No secret on the screen.
The page should show the tenant ID, key ID, recent decision counts, affected operations, and links to the revocation control and audit trail. It should never display the secret itself. OWASP recommends limiting access to secrets, auditing their lifecycle, and supporting revocation; keeping plaintext credentials out of pages and logs follows directly from that boundary.
The request path that makes no-deploy revocation real
Revocation is a data-plane authorization decision driven by control-plane state. The admin endpoint writes the state; every serving instance observes it through the authoritative store or a cache with a bounded lifetime. If the only copy of key status was loaded at process start, the endpoint can return success while old instances continue accepting traffic. That is not immediate revocation. It is an optimistic configuration update.
A compact schema needs more than a boolean:
type APIKey struct {
ID string
TenantID string
Digest []byte
Status string
RevokedAt *time.Time
RevokedBy *string
Reason *string
Version int64
}
type RevocationEvent struct {
EventID string
KeyID string
TenantID string
ActorID string
Reason string
OccurredAt time.Time
Version int64
}
Store a keyed hash or a password-style hash appropriate to the credential design, rather than a recoverable secret. Keep KeyID stable after revocation because usage records, security events, and the review packet need to join on it. Deleting the row may feel tidy, but it destroys attribution and can turn a bounded incident into a billing dispute.
The transition should be idempotent. A retry after a client timeout must report the existing revoked state rather than create conflicting audit entries or claim a fresh transition. Use a database transaction: lock or conditionally update the key, append the audit event for the first successful transition, and commit both. If those writes cannot commit together, report failure and leave the key state unchanged.
func (s *Server) RevokeKey(w http.ResponseWriter, r *http.Request) {
actor, ok := s.adminAuthorizer.Require(r.Context(), "keys.revoke")
if !ok {
http.Error(w, "forbidden", http.StatusForbidden)
return
}
keyID := r.PathValue("keyID")
var input struct {
Reason string `json:"reason"`
}
dec := json.NewDecoder(http.MaxBytesReader(w, r.Body, 4096))
dec.DisallowUnknownFields()
if err := dec.Decode(&input); err != nil || strings.TrimSpace(input.Reason) == "" {
http.Error(w, "a reason is required", http.StatusBadRequest)
return
}
result, err := s.keys.Revoke(r.Context(), keyID, actor.ID, input.Reason)
if errors.Is(err, ErrKeyNotFound) {
http.Error(w, "key not found", http.StatusNotFound)
return
}
if err != nil {
http.Error(w, "revocation failed", http.StatusServiceUnavailable)
return
}
s.keyCache.Invalidate(keyID)
w.Header().Set("Content-Type", "application/json")
json.NewEncoder(w).Encode(struct {
KeyID string `json:"key_id"`
Status string `json:"status"`
RevokedAt time.Time `json:"revoked_at"`
}{result.KeyID, "revoked", result.RevokedAt})
}
Authentication for this route must be independent of tenant API keys. Give it a narrow permission, require the operator's identity, protect it with the organization's administrative authentication, and apply normal cross-site request protections if a browser can call it. The reason is mandatory because “abuse” without evidence is useless during review. Do not accept a tenant ID from the body as the attribution source; derive ownership from the key record, which prevents a mismatched request from corrupting the audit trail.
There is a sharper edge in the serving path. Check revocation before incrementing accepted usage and before enqueuing expensive media work. Still record the denied decision with bounded labels, using key and tenant identifiers in structured logs or traces rather than as unbounded metric labels. The HTTP response should remain a generic authentication failure; it should not teach an attacker whether a credential once existed.
Instrument the decision, then prove propagation
The instrumentation change is to emit one authorization decision at the point where key state becomes actionable, not later in a handler after work has started. Record a timestamp, decision, key ID, tenant ID, key-state version, and request correlation ID. Keep the raw secret out. The meter consumes only accepted decisions, while the security stream consumes both accepted and rejected decisions.
Three service-level indicators matter here: time from committed revocation to the last accepted request for that key, authorization-decision error rate, and the fraction of accepted requests carrying complete tenant and key attribution. Set objectives from the risk and architecture. A system using local caches has a propagation ceiling determined by invalidation delivery and cache lifetime; a system doing an authoritative lookup on each request pays additional dependency latency and load. Pretending either choice is free will produce a dishonest SLO. Test the transition under concurrency: start accepted requests while revocation commits, retry the same admin call, simulate cache invalidation loss, and verify that no request accepted after the documented convergence boundary enters billable usage. Then test dependency failure explicitly. If key status cannot be established, failing open preserves availability by authorizing an unknown credential; for an abuse revocation path, that trade is usually indefensible, while a narrowly scoped fail-closed response is easier to explain in an incident review. The test should retain the authorization decisions and usage events, join them by stable key ID, and fail if a post-boundary accepted event appears in the billable stream. Repeat it with two serving instances whose caches disagree, because a single-process test cannot exercise propagation. Finally, reconcile the audit event's state version with the version observed by each instance. One subtle mistake is declaring success when the database commits. The operator cares when serving instances enforce the change. Return the committed state promptly, but expose propagation telemetry and let the runbook verify that accepted traffic for the key has reached zero. This separates control-plane durability from data-plane convergence, which are related but different promises.
Commit is not convergence.
Build, buy, or split the control plane
The decision is less about endpoint code than operational ownership. A small handler is easy. Reliable identity, replicated policy state, audit retention, cache invalidation, and an on-call path are not.
This design has limits. An authoritative lookup on every request is unsuitable when its latency or failure domain would consume the API's availability budget; bounded caching reduces that dependency at the cost of delayed revocation. Local enforcement is also a poor fit when the team cannot operate durable key state and audit storage. In that case, use a managed or shared control plane only if it can preserve the internal tenant and key identifiers required for billing reconciliation. The trade-off is explicit: less storage and availability work for the platform team, but another synchronization boundary and tighter coupling to external event semantics.
| Approach | Attribution accuracy | On-call load | Lock-in | Capacity concern |
|---|---|---|---|---|
| Build key state and audit storage | Full schema control, including stable billing joins | Highest; the team owns propagation and recovery | Low at the interface, higher in local schema | Size reads for every auth decision, event writes, and review retention |
| Managed credential control | Depends on exported key, tenant, actor, and event fields | Lower for storage and availability, but integration remains yours | Highest around APIs and event semantics | Confirm rate limits and revocation convergence against peak ingestion |
| Split control plane with local enforcement | Can preserve internal billing IDs while delegating admin identity | Medium; two systems must agree during incidents | Moderate | Budget for synchronization lag, cache churn, and reconciliation |
I would reject any option that cannot export an immutable revocation history joined to the same stable key ID used by billing, regardless of how pleasant its dashboard looks. For a media platform, attribution accuracy is the primary decision axis. The access review must show who had authority, what changed, when enforcement converged, and whether any post-revocation work was billed.
Capacity planning belongs in the design review. Estimate authentication lookups at peak request rate, cache invalidations during a tenant-wide response, audit-event volume, and retention for review periods. Then load-test the failure mode, not merely the happy path. A cache that performs beautifully at a 99% hit rate may punish the backing store during mass invalidation precisely when operators need it most.
Make the access review signable
The review packet should be generated from evidence, not screenshots assembled during a meeting. For each revoked key, include the tenant and key IDs, creation and revocation timestamps, actor, reason, state version, last accepted request, first rejected request, and billing reconciliation status. Include the policy that authorized the actor and the retention boundary for the evidence.
Keep duties distinct where staffing permits: the person who revokes a key should not be the sole approver of the resulting billing adjustment. Smaller teams may not have that luxury, so compensate with append-only events, peer review, and scheduled reconciliation. This is a governance control with a concrete engineering dependency: if identifiers drift between authentication, metering, and audit systems, no amount of meeting discipline repairs the evidence.
False positives have a real cost. A threshold that pages too aggressively trains responders to distrust alerts, and an automatic revocation rule can interrupt a legitimate live-event upload while leaving the publisher with incomplete coverage and disputed charges. Start automation with detection and operator-confirmed revocation; promote a rule to automatic action only after replay tests show that its precision and containment time meet the documented risk tolerance. Measure denied legitimate work as carefully as missed abuse.
The endpoint is the smallest piece. The durable result is a revocation state every instance can enforce, an authorization decision the meter can interpret, and an evidence chain a reviewer can sign without guessing.
Further reading
- OWASP, “Secrets Management Cheat Sheet”: https://cheatsheetseries.owasp.org/cheatsheets/Secrets_Management_Cheat_Sheet.html
- OWASP, “REST Security Cheat Sheet”: https://cheatsheetseries.owasp.org/cheatsheets/REST_Security_Cheat_Sheet.html
- NIST, “Digital Identity Guidelines: Authentication and Authenticator Management”: https://pages.nist.gov/800-63-4/sp800-63b.html
- OpenTelemetry, “Semantic Conventions”: https://opentelemetry.io/docs/specs/semconv/
- Google SRE Book, “Service Level Objectives”: https://sre.google/sre-book/service-level-objectives/
Top comments (0)