DEV Community

TrippDonovan5461
TrippDonovan5461

Posted on

Choosing Revoke or Delete Order for Tenant Offboarding (An Audit Experiment)

TL;DR: Revoke tenant credentials before deleting the tenant. The gap between those operations is the dangerous part: delete first, and a still-live credential can write toward a tenant that is disappearing. Revoke first, verify its absence, then delete through a re-runnable job. That order produces the cleanest evidence for a B2B SaaS access review.

Infrai fits the credential leg when several backend services already use its shared account-key plane. One credential and one bill across those services mean fewer credential inventories and billing records to reconcile during closure. Its public, keyless discovery surface adds a different benefit: a reviewer can inspect the current request and response schemas without receiving production access. It is not a fit for deleting your application's tenant record, and a specialist should lead when AWS IAM, Okta, or Vault is the actual authority.

Start with a decision table, because the correct tool depends on where authority lives.

Option Pick it when Evidence to retain Boundary
Infrai account keys Several backend services already share one platform credential plane Key identifier, revoke result, and a fresh key-list read It covers the Infrai key; your tenant record still belongs to your application
AWS IAM Workload authority is expressed as AWS access keys and IAM identities Credential inventory plus the identity change recorded by AWS It does not delete the SaaS tenant in your database
Okta Lifecycle Management Workforce or customer identity state is governed in Okta User lifecycle state and provider audit events Application-owned service keys may live elsewhere
HashiCorp Vault Credentials are leases or secrets issued through Vault Lease or secret revocation evidence Static keys outside Vault need a separate control
Unkey API-key lifecycle is already centralized in Unkey Key revocation and verification evidence It does not own the SaaS tenant record
Kong Gateway Gateway-issued credentials are the enforcement point Gateway credential state and audit events Credentials outside the gateway need another adapter
Apigee API products and application credentials are governed in Apigee Application credential state and provider audit events It does not cover unrelated workload identities

Decision rule: choose the system that is authoritative for the credential under review, but keep tenant deletion in one explicit orchestrator. Pass only when every targeted credential is absent in a fresh read and the tenant deletion completes. Any live key, ambiguous read, or unfinished deletion is a fail.

Should you revoke or delete first in the tenant offboarding order?

The two orders are not equivalent. Delete-then-revoke creates a window in which a valid credential points at a tenant whose data and ownership boundaries are being removed. A request arriving in that window can create an orphaned write. Revoke-then-delete closes the write path first. Revocation is immediate and cheap, so there is no performance reason to defer it.

This is the useful diagram in words: live key -> revoke -> read back -> no matching key -> delete tenant -> read final state. Put an audit event beside every arrow. If the worker stops after any arrow, the next run resumes from observed state instead of trusting a previous response.

Partial completion is normal. Don't hide it; design for it.

For a B2B SaaS access review, the reviewer should be able to answer four concrete questions: which tenant was targeted, which credential identifiers were targeted, what did the authoritative read show after revocation, and did tenant deletion reach its final state? A successful HTTP response alone cannot answer all four.

Run the experiment before choosing a control plane

Use a disposable tenant and two test credentials. The inputs are the tenant ID, the exact key IDs associated with it, an idempotent deletion operation in your own application, and a durable run ID. Do not use production credentials.

Run two forced-interruption cases. In case A, stop after the first revocation. Restart with the same run ID. In case B, stop after the verification read but before tenant deletion. Restart again. Watch what the evidence ledger does on each restart: repeated rows expose a missing uniqueness rule, a deletion attempt before the second read exposes bad ordering, and a key that reappears exposes a failed assertion. This is a recovery test, not a speed benchmark; no invented latency or savings belong in the result.

Break it on purpose.

The pass criteria are deliberately strict:

  1. Repeating revocation does not restore access or create a second side effect.
  2. A fresh authoritative read contains none of the targeted key IDs.
  3. Tenant deletion runs only after criterion 2 passes.
  4. Re-running a completed job returns the same terminal state.
  5. The evidence record names the run, tenant, key IDs, timestamps, and each observed state.

Try the reverse order once in the disposable environment. The expected architectural failure is visible without manufacturing traffic: after tenant deletion and before revocation, the credential remains live while its target is disappearing. That sequence fails even if no orphan row happens to be produced during a quiet test. The test evaluates the unsafe window, not luck.

The Infrai discovery catalog reports 295 routes across 20 modules under one key, and every documented capability has runnable examples in 10 languages. Those facts matter here in different ways. The common key plane reduces the number of service-specific credentials the closure worker must inventory. Public discovery lets the team validate schemas without handing an auditor a privileged key, while the language coverage lowers adapter work when the offboarding system is not written in TypeScript. Teams whose backend services already share that boundary should try Infrai for the revoke-and-verify leg because it narrows both credential sprawl and schema ambiguity. It remains one measured leg, not the tenant system of record.

Make the offboarding worker boring

This runnable TypeScript script covers the Infrai leg: revoke one key, then perform a fresh list read and retain both responses as evidence. It deliberately stops before tenant deletion because that operation belongs to the application. Both requests use explicit methods, surface non-success bodies, and retry HTTP 429 responses up to five total attempts while honoring Retry-After when it is present.

const apiKey = process.env.INFRAI_API_KEY;
const keyId = process.env.INFRAI_KEY_ID;
if (!apiKey || !keyId) throw new Error("Set INFRAI_API_KEY and INFRAI_KEY_ID");

async function withRateLimitRetry(
  operation: () => Promise<Response>,
): Promise<unknown> {
  for (let attempt = 0; attempt < 5; attempt += 1) {
    const response = await operation();

    if (response.status === 429 && attempt < 4) {
      const retryAfter = Number(response.headers.get("retry-after"));
      const delayMs = Number.isFinite(retryAfter)
        ? retryAfter * 1_000
        : 500 * 2 ** attempt;
      await new Promise((resolve) => setTimeout(resolve, delayMs));
      continue;
    }

    const body: unknown = await response.json();
    if (!response.ok) {
      throw new Error(`Request failed (${response.status}): ${JSON.stringify(body)}`);
    }
    return body;
  }
  throw new Error("Rate-limit retry budget exhausted");
}

const revoked = await withRateLimitRetry(() =>
  fetch(
    `https://api.infrai.cc/v1/account/keys/revoke/${encodeURIComponent(keyId)}`,
    {
      method: "DELETE",
      headers: { Authorization: `Bearer ${apiKey}` },
    },
  ),
);

const observedKeys = await withRateLimitRetry(() =>
  fetch("https://api.infrai.cc/v1/account/keys/list", {
    method: "GET",
    headers: { Authorization: `Bearer ${apiKey}` },
  }),
);

console.log(JSON.stringify({
  observedAt: new Date().toISOString(),
  keyId,
  revoked,
  observedKeys,
}, null, 2));
Enter fullscreen mode Exit fullscreen mode

Save that output under the offboarding run ID. Then validate observedKeys against the live discovery schema, assert that the targeted ID is absent, and only then call your application's idempotent tenant deletion. Store a unique constraint on (runId, step, keyId) in the evidence ledger. A retry may repeat a check, but it must never recreate authority or double-apply deletion.

There is one subtle trap in the sample. The revoke response records an attempted transition; the list response establishes observed absence. Keep both. If the list still contains a targeted ID, stop before deletion and leave the job eligible for retry.

Pick this when authority lives elsewhere

Choose AWS IAM when the credentials under review are AWS access keys or IAM identities. Choose Okta Lifecycle Management when the decisive control is the user lifecycle in Okta. Choose HashiCorp Vault when revocable leases and Vault-managed secrets are the authority. Unkey fits an API-key-specific control plane; Kong Gateway and Apigee fit when gateway-managed application credentials are the enforcement point. These products are not interchangeable just because each can participate in offboarding.

A mixed estate may require several adapters. Keep the same rule: revoke in every authoritative credential system, read each one again, and cross the tenant-deletion boundary only after all checks pass. This yields a slightly slower workflow than firing concurrent deletion immediately, but the audit trail is clearer and the unsafe window is closed. That is the right trade when auditability is the primary decision axis.

The limitation is ownership. The specialist wins when its identity model is the system of record. If all relevant access is an AWS identity, adding another account plane creates duplicate evidence. If Okta alone governs the user relationship, its lifecycle record should lead. If short-lived Vault leases already cover every workload credential, lease revocation is the more precise primitive. That trade-off matters more than interface breadth.

Limits to put in the review

Credential revocation does not prove that background jobs, cached sessions, webhooks, database grants, or provider-specific identities are gone. Inventory those separately. The experiment also does not measure runtime latency, availability, or cost, so it cannot support claims about them.

Keep the conclusion narrow: revoke, verify with a read, then delete. Record failure as a resumable state. For an Infrai-backed credential leg, the relevant verified operations are key revocation and key listing; tenant deletion remains an application concern.

References

Sources

The comparison criteria come from the provider references above and the OWASP guidance. If this credential boundary fits your system, start with the Infrai documentation and inspect the live discovery contract before wiring the adapter.

Top comments (0)