DEV Community

AlaricCross6851
AlaricCross6851

Posted on

6 Production Rules to Delete User Accounts and Revoke Sessions Safely

Treat account deletion as a durable, one-way state machine, and make credential invalidation the first externally visible transition. For an education platform that accepts Google and GitHub sign-in, the practical rule is: block new logins, revoke local sessions and API keys, erase owned data through retryable steps, and retain only the minimum tombstone needed to make retries harmless.

TL;DR: use six stages: requested, locked, credentials_revoked, data_erased, provider_unlinked, and completed. A repeated request must return the same operation identifier; a crashed worker must resume from durable state; and completion must never imply that an upstream identity-provider account was deleted. The page worth trusting is "a locked account still authenticated," not a green dashboard saying the queue is healthy.

Why can a deleted student still sign in?

Social sign-in separates identity proof from the application account. Google or GitHub can still issue a valid identity assertion after the learning-platform profile is gone, because deleting the local profile does not delete the person's upstream account. If the callback treats every valid assertion as permission to create a profile, the next login can resurrect an erased student with the same provider subject.

That is the failure mode I would design the page around. The callback must check a deletion tombstone keyed by the issuer and provider subject before provisioning anything. It must also reject an account whose deletion state has reached locked, even when an older browser cookie remains cryptographically valid. Signature verification answers whether a token was issued by the expected authority; it does not answer whether this application still permits the account to act.

Bot resistance changes the front door, too. Do not let an anonymous caller use the deletion endpoint as an account-enumeration oracle. Require a recently authenticated session, apply the same response shape regardless of whether a background operation already exists, and step up authentication before destructive action when the session is old or the risk signal is high. OWASP recommends reauthentication for sensitive features and after risk events; the deletion confirmation is exactly that boundary.

The six states are deliberately unglamorous. requested records intent and an idempotency key. locked prevents fresh sessions and provisioning. credentials_revoked records the local session-family and key cutoff. data_erased means the application-owned deletion plan has finished. provider_unlinked removes the local federation binding. completed is the stable response for every later retry.

How should a Node.js job delete a user account and revoke sessions?

The request handler should not synchronously walk course submissions, comments, uploaded files, audit stores, and analytics exports. That produces a long partial-failure window, and client retries can run two destructive traversals at once. Keep the transaction narrow: authenticate, lock one account, create or find one operation, advance the credential epoch, and enqueue through an outbox written in the same database transaction.

This sketch uses interfaces rather than a framework because the invariant belongs below the routing layer. The code is Go, but the transaction boundary is the useful part for a Node.js service as well.

package erasure

import (
    "context"
    "errors"
    "time"
)

var ErrRecentAuthenticationRequired = errors.New("recent authentication required")

type Request struct {
    AccountID      string
    IdempotencyKey string
    AuthenticatedAt time.Time
}

type Operation struct {
    ID      string
    Account string
    Stage   string
}

type Store interface {
    InTransaction(context.Context, func(Tx) error) error
}

type Tx interface {
    FindOperation(context.Context, string, string) (Operation, bool, error)
    LockAccount(context.Context, string, time.Time) error
    AdvanceCredentialEpoch(context.Context, string) error
    CreateOperation(context.Context, string, string, string) (Operation, error)
    AppendOutbox(context.Context, string, string) error
}

func RequestDeletion(ctx context.Context, db Store, now time.Time, r Request) (Operation, error) {
    if now.Sub(r.AuthenticatedAt) > 10*time.Minute {
        return Operation{}, ErrRecentAuthenticationRequired
    }

    var op Operation
    err := db.InTransaction(ctx, func(tx Tx) error {
        found, ok, err := tx.FindOperation(ctx, r.AccountID, r.IdempotencyKey)
        if err != nil {
            return err
        }
        if ok {
            op = found
            return nil
        }

        if err := tx.LockAccount(ctx, r.AccountID, now); err != nil {
            return err
        }
        if err := tx.AdvanceCredentialEpoch(ctx, r.AccountID); err != nil {
            return err
        }
        op, err = tx.CreateOperation(ctx, r.AccountID, r.IdempotencyKey, "locked")
        if err != nil {
            return err
        }
        return tx.AppendOutbox(ctx, op.ID, "erase-account")
    })
    return op, err
}
Enter fullscreen mode Exit fullscreen mode

The 10*time.Minute window is an application policy in this example, not a standard or a universal recommendation. Pick a window through threat modeling, document it, and test the clock boundary. More important is the ordering: the account lock and credential epoch change commit with the operation record. A worker outage can delay data cleanup, but it cannot leave an account usable merely because the queue stalled.

Store the idempotency key under a uniqueness constraint scoped to the authenticated account. Do not use the key as authorization, do not accept an account identifier from the request body as identity, and do not log bearer tokens or the key itself. A random operation ID is suitable for status lookup only when that lookup is also authorized.

The main limitation is coordination cost. A credential epoch requires every resource service to consult fresh enough account state, while an outbox, checkpoints, and tombstones add storage and operational machinery. A small monolith with only opaque server-side sessions may reasonably delete its session rows in the opening transaction instead; a distributed platform should accept the extra machinery because independent services otherwise disagree about whether deletion has taken effect. That is a real trade-off, not free reliability.

Revoke capabilities before erasing their evidence

Session revocation has two layers. Delete server-side session records where they exist, then increment a per-account credential epoch or set a valid_after timestamp checked by every protected request. Stateless access tokens minted before the cutoff remain signed correctly, so resource servers must compare a token claim with current account state, or use short-lived access tokens backed by a revocable session family. API keys need the same treatment: mark their hashed records revoked before removing profile data that explains who owned them.

Do this early.

For federated login, the platform should remove its stored provider tokens and linkage, but it must not claim to revoke every upstream session. OAuth 2.0 token revocation defines a revocation endpoint for tokens and says an authorization server may invalidate related tokens; OpenID Connect defines relying-party initiated logout separately. Provider support and semantics vary, so local denial remains the control the application owns.

Erasure itself should be a plan of independently repeatable handlers. A handler may delete a private profile row, detach an uploaded object, or replace an author reference where a course record must be retained for another lawful purpose. GDPR Article 17 supplies the right to erasure and its exceptions; it does not justify promising that every byte in backups disappears during the request transaction. Retention decisions need a documented legal basis and lifecycle, not an improvised DELETE CASCADE.

package erasure

import "context"

type Step func(context.Context, string) error

type Plan struct {
    RevokeSessions Step
    RevokeKeys     Step
    EraseOwnedData Step
    UnlinkIdentity Step
    WriteTombstone Step
}

func (p Plan) Run(ctx context.Context, accountID string) error {
    steps := []Step{
        p.RevokeSessions,
        p.RevokeKeys,
        p.EraseOwnedData,
        p.UnlinkIdentity,
        p.WriteTombstone,
    }
    for _, step := range steps {
        if err := step(ctx, accountID); err != nil {
            return err
        }
    }
    return nil
}
Enter fullscreen mode Exit fullscreen mode

Each step must treat "already absent" as success. That requirement sounds small until a worker dies after deleting an object but before recording its checkpoint. On retry, absence is evidence that the desired state may already hold; turning it into a fatal error creates a permanently stuck operation. Conversely, broad error suppression is dangerous. Authentication failure, storage timeout, and malformed ownership data are not "already absent."

What should wake the on-call engineer?

Queue depth alone is a weak signal. A dashboard can show a modest backlog while one tenant's deletion has retried for hours, or show a spike caused by an expected batch while every operation still meets its objective. Ask what page fired.

The actionable page is an invariant violation: any request accepted after locked, any new federation link created for a tombstoned issuer-subject pair, any credential accepted with an epoch older than the account epoch, or any operation exceeding the documented completion objective. Failed steps should emit structured events with the operation ID, stage, attempt count, dependency class, and next retry time, but without email addresses, provider tokens, lesson content, or raw idempotency keys.

A useful operational view separates progress from correctness:

Signal Page or ticket? Reason
Locked account successfully calls a protected API Page The security boundary failed
Tombstoned identity creates a fresh profile Page Erasure was reversed
One retryable dependency error Observe Backoff may recover without intervention
Oldest operation breaches its completion objective Page A user-facing obligation is stuck
Queue depth rises while operation age stays bounded Ticket or observe Capacity may need adjustment, but correctness holds

Metrics need bounded labels. Stage and dependency class are useful; account ID and operation ID belong in authorized logs or traces, not metric labels. This distinction is boring until cardinality makes the monitoring system fail during the incident it was meant to explain.

Verify recovery, not the happy path

Test at the boundaries where state disagrees. Start two requests with the same idempotency key and assert one operation exists. Start two requests with different keys for the same account and assert the account still has one active erasure workflow. Kill the worker after every external side effect, restart it, and verify convergence. Present an access token minted before the credential cutoff to each resource service. Replay an old browser cookie. Attempt Google and GitHub sign-in for a tombstoned binding and verify that neither path reprovisions the account.

Then test hostile traffic. Rate-limit by more than a user-supplied email address, because a bot can rotate addresses and use the endpoint to burn mail or identity-provider capacity. Keep denial responses uniform enough to avoid disclosing whether an account or operation exists. Record abuse signals separately from the durable erasure state so a noisy defense system cannot silently cancel a valid request.

Rollback is asymmetric. Before locked, an uncommitted request can disappear. After the lock commits, rolling back by re-enabling credentials is usually the wrong recovery action; repair the worker or dependency and continue forward. If policy permits a cancellation grace period, model it as an explicit pre-erasure state with a deadline, require recent authentication to cancel, and never call the account erased during that window.

The production checklist is short enough to use during review:

  1. Recent authentication gates the destructive request, while rate limits and neutral responses resist enumeration and automation.
  2. A database uniqueness constraint makes request retries return one durable operation.
  3. Account lock, credential cutoff, operation creation, and outbox insertion share one transaction.
  4. Every protected service enforces the lock or credential epoch; deleting cookies in one browser is insufficient.
  5. Erasure steps are independently idempotent, checkpointed, and explicit about retained records.
  6. Federation tombstones prevent automatic reprovisioning, and provider unlinking is not misrepresented as deletion of the upstream identity.
  7. Failure injection covers crashes between side effects and checkpoints, while alerts target invariant violations and excessive operation age.
  8. Logs and metrics avoid secrets and personal data, and access to operation status is authorized.

Completion means the local account can no longer authenticate, its governed erasure plan has converged, and retries cannot recreate it. That is a claim an incident responder can test at 3 a.m.; a progress bar and a successful job count are not.

References

Top comments (0)