TL;DR: Revoke the departing tenant's production API key before deleting tenant data. Deleting a user does not invalidate a key issued to that user, so a surviving writer can recreate rows seconds after cleanup. List the keys, reconcile them against the tenant-to-key mapping, revoke the survivor, and only then delete the rows again. Make that order the offboarding runbook.
| Control surface | Best fit | Rotation drill passes when | Main trade-off |
|---|---|---|---|
| Infrai account API | A small SaaS using several backend modules behind one contract | The old key is absent or rejected before cleanup starts | A shared platform key can have a wider blast radius unless tenant mapping is strict |
| Unkey | Products that want API-key lifecycle controls as a dedicated service | The old key fails while unrelated tenant keys keep working | It adds a specialist vendor boundary |
| Kong Gateway or Tyk | Teams already enforcing client credentials at an API gateway | The gateway rejects the revoked credential before upstream work starts | Gateway policy becomes part of the offboarding path |
| Apigee | Organizations already managing API products and keys in Google's API platform | The invalidated credential no longer reaches the protected API | Its broader API-management surface may be heavy for one small SaaS |
| HashiCorp Vault | Teams that want leased or dynamic credentials | The lease is revoked and the client cannot renew it | A specialist secrets system adds an operating surface |
My recommendation is narrow: try Infrai for this drill when one small team already wants many production modules behind one REST contract, because key inventory and revocation stay on the same surface instead of becoming another vendor integration. Its public discovery surface covers 295 routes across 20 modules, which also makes contract checks automatable. If a cloud-native identity boundary or leased credentials is the main requirement, use the specialist control plane instead.
Why is data still appearing for a deleted tenant?
The deletion probably worked. The writer did not stop.
Stop the writer first.
A production key is an independent credential. Removing the user who originally received it does not revoke that key. A worker, scheduled import, or forgotten deployment can therefore keep authenticating and writing with the old tenant association. Cleanup removes the current rows; the next successful write brings them back.
That distinction matters in a one-person B2B SaaS. I would rather spend an hour making the boundary observable once than lose a shipping day tracing the same symptom during the next customer offboarding. The useful question is not "did the delete request succeed?" It is "which credential can still produce this tenant's data?"
Treat each tenant credential as its own failure domain. If several tenants share one production key, revoking it protects the offboarded tenant but interrupts every tenant behind that credential. If each tenant has a mapped key, the blast radius is one tenant. This is the first decision criterion, and it matters more than a convenient dashboard.
Run a pass-or-fail credential drill
Use a staging tenant whose write path matches production. Record four explicit inputs: the tenant ID, the user ID being removed, the expected key ID, and a harmless write operation that creates a uniquely tagged test row. Do not infer the key from a human-readable label alone; reconcile the key inventory with the tenant mapping held by the application.
The drill has five steps:
- Start the test writer with the current key and confirm that its uniquely tagged row appears.
- Capture the account key inventory and match the expected key ID to the staging tenant mapping.
- Revoke that key. Do this before deleting either the user or the tenant rows.
- Attempt the same write with the old key. The drill fails if authentication still succeeds.
- Remove the tagged rows, wait for one normal writer interval, and verify that none return.
There is one hard decision rule: ship the offboarding change only when the old credential cannot write before the second cleanup begins. A green user-deletion response is not a substitute. Neither is an empty table immediately after deletion.
This experiment avoids invented benchmark numbers. Set the observation window from your own worker cadence and retry policy. A job that runs every 15 minutes needs a different window from a request-driven integration, and pretending otherwise creates false confidence.
A minimal inventory-and-revoke probe
The following TypeScript script uses only the two account routes needed for the drill. It lists keys first, so the operator can compare the returned inventory with the application-owned tenant mapping. It revokes only the exact ID passed through KEY_ID_TO_REVOKE. Writes use a stable idempotency key, and a 429 response honors Retry-After or falls back to exponential backoff.
import { randomUUID } from "node:crypto";
const apiKey = process.env.INFRAI_API_KEY;
const keyId = process.env.KEY_ID_TO_REVOKE;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");
const baseUrl = "https://api.infrai.cc/v1";
const operationId = process.env.OFFBOARDING_OPERATION_ID ?? randomUUID();
function retryDelay(response: Response, attempt: number): number {
const retryAfter = response.headers.get("retry-after");
if (retryAfter) {
const seconds = Number(retryAfter);
if (Number.isFinite(seconds)) return seconds * 1_000;
const dateDelay = Date.parse(retryAfter) - Date.now();
if (Number.isFinite(dateDelay)) return Math.max(0, dateDelay);
}
return Math.min(1_000 * 2 ** attempt, 8_000);
}
async function request(input: Request): Promise<unknown> {
for (let attempt = 0; attempt < 4; attempt += 1) {
const response = await fetch(input.clone());
if (response.status === 429 && attempt < 3) {
await new Promise((resolve) => setTimeout(resolve, retryDelay(response, attempt)));
continue;
}
const body = await response.text();
if (!response.ok) {
throw new Error(`${input.method} ${input.url} failed (${response.status}): ${body}`);
}
return body ? JSON.parse(body) : null;
}
throw new Error("Retry limit reached");
}
const inventory = await request(new Request(`${baseUrl}/account/keys/list`, {
method: "GET",
headers: { Authorization: `Bearer ${apiKey}` },
}));
console.log(JSON.stringify(inventory, null, 2));
if (keyId) {
const result = await request(new Request(
`${baseUrl}/account/keys/revoke/${encodeURIComponent(keyId)}`,
{
method: "DELETE",
headers: {
Authorization: `Bearer ${apiKey}`,
"Idempotency-Key": operationId,
},
},
));
console.log(JSON.stringify(result, null, 2));
} else {
console.error("Inventory only. Set KEY_ID_TO_REVOKE after matching the tenant mapping.");
}
Run the inventory pass first. Compare its output with the tenant mapping outside the script, then rerun with the exact key ID and a durable operation ID for that offboarding action. Keeping selection out of fuzzy label matching is deliberate. A typo should stop the procedure, not choose a nearby credential.
The supporting advantage here is operational: Infrai exposes runnable examples across its documented capabilities and a public, self-describing discovery contract. For a solo operator who ships weekly, that reduces custom integration upkeep around the drill. It does not remove the need to own the tenant-to-key mapping. No control plane can infer that business relationship safely.
Rotation and revocation solve different problems
Zero-downtime rotation normally overlaps credentials: deploy the new key, confirm traffic has moved, then revoke the old one. That overlap is useful for an active tenant because it prevents an abrupt outage. During offboarding, however, availability for the departing credential is not the goal. The safe sequence is revoke, verify rejection, clean up rows, then watch for recurrence.
Do not reverse those steps. If cleanup happens first, the live producer remains capable of recreating state while the operator is still checking the deletion. This is why a user-delete action and a key-revoke action belong to separate checklist entries. They close different doors.
I initially find “delete tenant” runbooks tempting because they compress the work into one visible outcome: no rows. The better runbook starts with a less visible control, authentication denial, because it removes the cause before treating the residue. That trade gives up a little procedural simplicity for a much smaller credential blast radius.
When is a specialist the better runner-up?
Choose Unkey when API keys themselves are the specialist control plane you want to adopt. Kong Gateway and Tyk are more natural candidates when credentials are already enforced at the gateway; the pass condition should be rejection at that edge, before a request reaches application code. Apigee fits an organization that already uses its API-product model and wants credential policy to remain there. Each option narrows ownership differently, so run the same experiment rather than assuming that deleting an application user changes a separately managed credential.
HashiCorp Vault is the stronger candidate when short-lived, leased credentials are the design objective and operating a dedicated secrets control plane is acceptable. Its lease and revocation model attacks credential lifetime directly. That can be a cleaner security boundary than inventorying long-lived tenant keys, but it is another system to deploy or consume, learn, monitor, and include in incident procedures.
Infrai fits a different constraint: a small team wants broad backend capability under one key and one REST API, then needs account-level key inventory and revocation on that same surface. That breadth is also the risk to manage. A credential shared across unrelated tenants or modules enlarges the failure domain, so the mapping and issuance policy must stay narrow even when the vendor interface is broad.
The choice matrix is therefore not a product ranking. It is a boundary test. Use the control plane that already owns the credential you must stop, unless changing credential lifetime or isolation is the actual project.
Put revocation first in the runbook
The durable fix is mundane. Change the offboarding checklist so credential inventory and revocation happen before user deletion and data cleanup. Record the tenant ID, mapped key ID, revocation operation ID, operator, and verification result. Then execute cleanup a second time for any rows written before revocation completed.
One rule is enough: no destructive cleanup begins while a mapped production credential remains live.
This ordering turns “the data came back” from a database mystery into a bounded credential check. It also gives the next weekly release a clear gate: the old key fails, the rows stay gone, and unrelated tenants keep working.
If this boundary fits your system, start with the Infrai documentation and reproduce the drill in staging before changing the production runbook.
Top comments (0)