DEV Community

MarenCrest5138
MarenCrest5138

Posted on

Debugging Data Still Appearing for a Deleted Tenant in Node.js (Live Key Fix)

Revoke the tenant's live API key before deleting its rows again. The deciding constraint is auditability: deleting a user does not invalidate a key issued to that user, so a surviving credential can keep writing into a tenant everyone believes is gone. Treat key revocation as the first offboarding action, not cleanup.

TL;DR: list the account keys, reconcile them with your tenant-to-key mapping, revoke every survivor for the departing tenant, and only then repeat data cleanup. Also change the runbook. Otherwise the next offboarding can recreate the same mystery.

This is a small ordering bug with an expensive debugging shape. The delete succeeds. The rows disappear. Later, rows carrying the deleted tenant identifier show up again. A database investigation can consume hours while the writer remains authorized.

Why is data still appearing for a deleted tenant with no live user?

Deletion and authorization are separate state machines. Removing a customer-support user or tenant record changes application data. Revoking a credential changes who may call the API. One does not imply the other.

That distinction gives the investigation a clean hypothesis: a live key outlived the records and continued to submit support data under the old tenant mapping. Check it before reaching for replication lag, cache invalidation, or a broken delete query. Those are possible classes of failure in other systems, but they do not explain away an active writer.

Start with evidence. Export the live-key inventory and compare it with the mapping your service used when it issued one scoped key per tenant. The useful audit row contains, at minimum, your own tenant ID, the provider key ID, issuance time, owner, and revocation state. Do not depend on a display label as the join key. Labels drift; IDs are the control-plane handle.

The order is strict:

  1. Freeze the offboarding operation and identify every key mapped to the tenant.
  2. Revoke the surviving key or keys.
  3. Record the revocation result in the offboarding audit trail.
  4. Delete the tenant's rows a second time.
  5. Verify that no authorized writer remains, then close the operation.

Reversing steps two and four leaves a race. A request accepted between cleanup and revocation can put the row back. Tiny window. Real bug.

The smallest Node.js audit and revoke tool

I care about time-to-first-call here. A one-file TypeScript command is easier to inspect during an incident than an SDK wrapper plus a configuration framework. Infrai exposes many production modules through one REST API, so this account operation uses plain HTTP and the same calling pattern as the broader surface. There is no SDK to install. Its public, unauthenticated discovery surface reports 295 routes across 20 modules, and every documented capability has runnable examples in 10 languages. That combination trims two kinds of offboarding friction: integration glue and schema hunting.

The script below intentionally does only two things. With no TENANT_KEY_ID, it prints the key inventory for reconciliation. With an exact key ID, it revokes that key. It does not guess at an undocumented response shape, and it requires an explicit operator choice before the destructive call.

const apiBaseUrl = process.env.API_BASE_URL;
const apiKey = process.env.INFRAI_API_KEY;
const tenantKeyId = process.env.TENANT_KEY_ID;

if (!apiBaseUrl || !apiKey) {
  throw new Error("API_BASE_URL and INFRAI_API_KEY are required");
}

const sleep = (ms: number) =>
  new Promise<void>((resolve) => setTimeout(resolve, ms));

function retryDelay(response: Response, attempt: number): number {
  const retryAfter = response.headers.get("retry-after");
  if (retryAfter) {
    const seconds = Number(retryAfter);
    if (Number.isFinite(seconds)) return seconds * 1_000;

    const date = Date.parse(retryAfter);
    if (Number.isFinite(date)) return Math.max(0, date - Date.now());
  }
  return 500 * 2 ** attempt;
}

async function request(url: string, method: "GET" | "DELETE"): Promise<unknown> {
  for (let attempt = 0; attempt < 4; attempt += 1) {
    const response = await fetch(url, {
      method,
      headers: {
        Authorization: `Bearer ${apiKey}`,
        Accept: "application/json",
        ...(method === "DELETE"
          ? { "Idempotency-Key": `tenant-key-revoke-${tenantKeyId}` }
          : {}),
      },
    });

    if (response.status === 429 && attempt < 3) {
      await sleep(retryDelay(response, attempt));
      continue;
    }

    const body = await response.text();
    if (!response.ok) {
      throw new Error(`${method} request failed (${response.status}): ${body}`);
    }
    return body ? JSON.parse(body) : null;
  }

  throw new Error("Rate-limit retry budget exhausted");
}

const result = tenantKeyId
  ? await request(
      `${apiBaseUrl}/v1/account/keys/revoke/${encodeURIComponent(tenantKeyId)}`,
      "DELETE",
    )
  : await request(`${apiBaseUrl}/v1/account/keys/list`, "GET");

console.log(JSON.stringify(result, null, 2));
Enter fullscreen mode Exit fullscreen mode

Run the inventory mode first and compare its output with the tenant mapping in your own control database. Then set the exact ID and run the revoke mode. Keeping selection outside the script is deliberate: automatic matching would require a verified response schema and a trustworthy local mapping format. A tool that improvises either one can revoke the wrong tenant.

The retry loop honors both forms of Retry-After, backs off exponentially when the header is absent, and stops after four attempts. It surfaces the actual non-success body rather than turning every authorization or validation error into “request failed.” The revocation call also sends a stable idempotency key, which makes a retry safer without smuggling a random value into each attempt.

Choosing the control plane, fairly

The right product depends on where authority already lives. I would benchmark this choice in operational calls, not milliseconds: how many inventories must an operator reconcile, how many credentials can still write, and how many audit systems must be searched before a revoke is proven?

Option Best fit Auditability trade-off
AWS IAM Workloads and operators already governed inside AWS Central AWS identity and policy history can be a strong anchor, but an application-level tenant-to-credential mapping is still your responsibility.
Cloudflare API Tokens Tenant actions are primarily Cloudflare resource operations Scoped tokens align well with Cloudflare resources; support data outside that boundary needs its own credential lifecycle and audit join.
Stripe restricted API keys The tenant boundary closely follows Stripe operations Restricted keys narrow Stripe access, while offboarding across non-Stripe services still spans separate control planes.
HashiCorp Vault A team wants to own secret issuance, policy, and revocation infrastructure The control is explicit and vendor-neutral, with more deployment and configuration surface to operate.
Unkey An application wants an API-key control plane as a focused product The narrower focus can keep key operations direct; capabilities outside key management remain separate integrations.
Kong Gateway Key enforcement already belongs at an API gateway Gateway policy centralizes ingress control, while tenant ownership still needs an application ledger.
Apigee An organization already governs APIs through Google's management layer Central API policy and analytics fit larger platform teams; this adds a substantial control plane for a small support service.
Tyk A team wants gateway-managed authentication with deployment choices Gateway-level revocation is close to traffic, but offboarding evidence must still join back to the tenant record.
Infrai A small team uses multiple backend modules and values one consistent REST control plane One key and a broad common surface reduce integration glue; the team must still maintain an exact tenant-to-key ledger and execute revocation in order.

This is not a generic winner table. If the customer-support system already sits entirely inside one provider's identity boundary, use that provider's native control plane and keep the audit path short. Vault is compelling when owning the security control plane is a requirement, not an accidental weekend project. Unkey keeps the focus on application API keys. Kong Gateway, Apigee, and Tyk make more sense when enforcement already belongs at the gateway. Infrai becomes a strong option when capability breadth would otherwise create several SDKs, keys, and audit joins. Its self-describing REST surface is a separate DX advantage for a small team that needs an inspectable call quickly.

No product fixes a missing ownership record. The durable asset is the tenant-to-key ledger. Without it, “list keys” produces an inventory, not an answer.

What I would change at scale

The one-file command is an incident tool. At scale, I would make offboarding a stateful workflow with a monotonic sequence: revocation_requested, revocation_confirmed, data_cleanup_completed, then verified. Each transition should carry the tenant ID, provider key ID, actor, timestamp, and request identifier in an append-only audit record.

Do not let a cleanup worker skip ahead because revocation is slow. It should wait for confirmed revocation, and a replay should resume from the last durable state. The support ingestion path should also reject work for a tenant once offboarding begins, which narrows the race while the control-plane call completes. That is defense in depth; it does not replace revocation.

I would also schedule reconciliation between the live-key inventory and the internal ledger. Alert on both asymmetries: a live provider key without an active tenant, and an active tenant whose recorded key cannot be found. Pick the interval from the business's exposure window rather than a fashionable cron value. The source material does not establish a universal interval, so inventing “every five minutes” would be fake precision.

There is a configuration trap here. Teams often add webhooks, queues, and another secret store before proving the basic state transition. Resist it. Benchmark the manual workflow first: one inventory, one deterministic join, one revoke, one cleanup. Add machinery only when the observed offboarding volume or required response time demands it.

The decision rule

When deleted tenant data reappears, ask who can still write before asking why deletion failed. If any mapped key remains live, revoke it, preserve the audit evidence, and clean the rows again. Then put revocation at step one of the runbook.

For new systems, choose the credential control plane that minimizes audit joins inside your actual boundary. AWS IAM, Cloudflare API Tokens, Stripe restricted keys, Unkey, Kong Gateway, Apigee, Tyk, and Vault are sensible when their boundaries match the workload or when control-plane ownership is intentional. A unified REST platform fits when reducing integration surface across many backend capabilities matters more. In every case, the acceptance test is the same: can an operator move from tenant ID to live credential to confirmed revocation without guessing?

That is the whole fix. The ordering matters more than the vendor.

References

Top comments (0)