Short answer: make credential revocation the first irreversible state transition, but keep a short, auditable data-retention window so billing and dispute records can finish safely. The useful question is not "before or after deleting user data?" It is: which actions still need a tenant credential, and what is the blast radius if that credential leaks during the job?
For a one-person SaaS, I optimize for revenue per hour. I ship weekly, and I outsource the undifferentiated work to boring queues and database primitives. Offboarding is one of the places where boring wins. A clean state machine prevents a support ticket from becoming an incident.
Freeze first.
A small decision matrix for prepaid-balance tenants
| Step | Keep tenant access? | Why | Recovery move |
|---|---|---|---|
| Freeze spending | No new charges | Stop the prepaid balance from draining | Re-enable only after an explicit review |
| Revoke API keys | No | Shrinks the blast radius immediately | Issue a new key only during a controlled restore |
| Export required records | Service credential only | Preserve ledger, tax, and dispute evidence | Retry from an idempotent job |
| Delete user data | No user access | Apply the retention policy after the export | Restore from a restricted archive if policy allows |
That ordering separates access control from deletion. Revocation is fast and reversible; deletion is slower and often permanent. The data export is a service-owned operation, not a reason to leave a customer key alive.
What should a tenant offboarding sequence do with an API key?
Record an offboarding id, tenant id, policy version, actor, timestamps, and the exact completion state. Do not infer progress from whether a row happens to exist. A row can disappear after a timeout while the downstream export is still running.
I once treated a 202 response as completion and moved on. That was a bad assumption: the worker had accepted the job, but the ledger snapshot had not been sealed. Now every transition has an idempotency key and a durable audit event.
type OffboardingState =
| "requested"
| "frozen"
| "keys_revoked"
| "records_exported"
| "user_data_deleted"
| "closed";
const order: OffboardingState[] = [
"requested",
"frozen",
"keys_revoked",
"records_exported",
"user_data_deleted",
"closed",
];
export async function advanceOffboarding(tenantId: string, jobId: string) {
const current = await db.offboarding.get(tenantId);
if (!current || current.jobId !== jobId) throw new Error("unknown offboarding job");
const next = order[order.indexOf(current.state) + 1];
if (!next) return current; // safe retry after closure
await db.transaction(async (tx) => {
await tx.offboarding.compareAndSet(tenantId, current.state, next);
await tx.audit.append({ tenantId, jobId, from: current.state, to: next });
});
return { ...current, state: next };
}
The compare-and-set matters more than the enum. Two workers can receive the same message; only one may advance the state. Every side effect should accept the same job id, so a retry cannot create a second export or a second deletion request.
How do revoke, export, and delete interact under failure?
Think in checkpoints. Freeze spending and revoke keys can be retried immediately. Export should write to a restricted location with a retention deadline, then record a checksum or immutable version. Only after that checkpoint should the worker delete user-facing data.
If deletion times out, leave the state at records_exported; do not mark success because the HTTP request returned. If the process crashes after deletion but before the audit write, reconcile from the database transaction log or an append-only event table. Your mileage may vary with the storage engine, so test this crash window instead of assuming atomic behavior across services.
The ugly case deserves a written runbook. Imagine the worker freezes spending at 09:00, revokes the tenant key at 09:01, uploads a ledger snapshot at 09:02, and loses its process before the checksum event commits. A retry must discover the same jobId, verify whether the snapshot is complete, and either attach the missing audit event or replace the incomplete object; it must not create a second customer export. If the database says keys_revoked but the provider says the key is still active, the reconciler should keep access frozen, mark the job for review, and emit the evidence needed to resolve the mismatch. That is more code than a single delete endpoint, but it is cheaper than explaining an accidental balance drain to a finance team.
Metrics should show counts and age for each state, plus the number of jobs waiting for manual review. A single alert on job failure is noisy. An alert on keys_revoked older than 15 minutes is actionable because it tells me access is closed while cleanup is stuck.
This design is not suitable when a tenant has legal holds, cross-region residency rules, or a long-running settlement process that cannot be represented as a short checkpoint. In those cases, keep the account suspended, preserve the minimum evidence required by policy, and use a workflow engine with human approval and compensation steps.
It is also a poor fit for a shared database where deleting one tenant can lock unrelated customers. Use partitioned jobs, bounded batches, and online index maintenance. Stick with a supervised, manual run when the expected deletion volume is tiny and the cost of a mistaken delete is existential.
The practical rule is simple: revoke customer credentials early, export only what policy requires, delete in observable chunks, and make every transition replayable. That gives a solo founder a smaller blast radius without pretending deletion is a single request.
References
- OWASP Secrets Management Cheat Sheet: https://cheatsheetseries.owasp.org/cheatsheets/Secrets_Management_Cheat_Sheet.html
- Node.js
cryptoAPI documentation: https://nodejs.org/api/crypto.html - NIST Digital Identity Guidelines, Authentication and Lifecycle Management: https://pages.nist.gov/800-63-4/sp800-63b.html
Top comments (0)