Short answer: make account deletion a three-gate workflow: record the user's confirmed intent, revoke every active session and one-time-code challenge, then remove or retain data according to a documented retention policy. For a developer-tools app adding phone OTP login, the delete request is not complete when the user row disappears; it is complete when no credential or session can re-enter the account and the audit trail can explain what happened.
I treat this as a runbook problem. A queue can redeliver a deletion command, a deploy can interrupt it, and a support engineer can receive a second request while the first is still running. The worker must make repetition boring. I was initially tempted to put a single DELETE FROM users behind an endpoint. Then I mapped the data graph and found phone challenges, refresh tokens, device sessions, consent receipts, and billing references outside that table. The row delete was the easy part.
What should an account deletion workflow preserve about consent, sessions, and user removal?
Start with an explicit state machine: requested, confirmed, revoking, purging, completed, and blocked. Store a deletion request ID, subject ID, policy version, confirmation timestamp, actor, and reason. The request ID is correlation data, not proof that work finished. A second command with the same ID should return the recorded state instead of starting an unrelated purge.
Consent cleanup needs precision. Delete or anonymize product-consent records only when the policy permits it; retain a minimal legal or security record when a regulation or dispute process requires evidence. Keep the purpose, policy version, and timestamp, but remove the phone number and other direct identifiers where retention does not require them. Do not call every retained audit event “consent.” A receipt that proves a deletion request was confirmed is a different record from a marketing opt-in.
Session revocation is its own gate. Mark server-side sessions and refresh-token families revoked, invalidate password-reset links, and invalidate outstanding OTP challenges for the subject. Access tokens that are already issued may remain cryptographically valid until expiry, so APIs should also reject a token whose session or subject status is revoked. Keep the expiry short enough to match the application's risk tolerance.
One sentence matters here.
Delete twice.
User removal comes last because downstream jobs need a stable subject reference while revocation propagates. Use a tombstone or opaque deletion ID for joins, then erase or tokenize profile data, phone numbers, uploaded content, and search indexes. In the developer-tools app, that means checking API keys issued for CLI access, webhook signing secrets, team memberships, comments, and cached search documents as well as the login tables. Make an inventory from schema ownership and event consumers, not from whatever tables happen to be in the authentication repository; the second list is usually longer. A foreign key that prevents deletion is a dependency to resolve deliberately, not a reason to skip the workflow.
How can a Go service make phone OTP revocation idempotent?
The handler should authenticate the current session, require a recent OTP confirmation for a sensitive delete, and create a durable request before enqueueing work. Never infer confirmation from a phone number supplied in the request body. Rate-limit OTP attempts, avoid revealing whether a phone is registered, and keep the code lifetime and retry limits aligned with your threat model. OWASP's Authentication Cheat Sheet recommends generic responses for account enumeration and careful handling of authentication errors; those rules apply during deletion too.
The worker can use a transaction for each gate. The conditional updates below make retries safe: a row already marked revoked or purged produces no second transition.
package deletion
import (
"context"
"database/sql"
"fmt"
)
func RevokeSessions(ctx context.Context, tx *sql.Tx, subjectID, requestID string) error {
result, err := tx.ExecContext(ctx, `
UPDATE sessions
SET revoked_at = COALESCE(revoked_at, now()),
revoked_by_request = COALESCE(revoked_by_request, $2)
WHERE subject_id = $1`, subjectID, requestID)
if err != nil {
return fmt.Errorf("revoke sessions: %w", err)
}
_, _ = result.RowsAffected() // emit the count as an audit metric
_, err = tx.ExecContext(ctx, `
UPDATE otp_challenges
SET consumed_at = COALESCE(consumed_at, now()),
consumed_reason = COALESCE(consumed_reason, 'account_deletion')
WHERE subject_id = $1 AND consumed_at IS NULL`, subjectID)
if err != nil {
return fmt.Errorf("invalidate otp challenges: %w", err)
}
return nil
}
The purge step should use an allowlisted table plan, not dynamically concatenate table names from client input. For each table, record a count before and after the operation, and make the operation resumable. A process restart after table three must continue at table four, not repeat an irreversible action blindly. If a child service owns data, send a signed deletion event with the request ID and wait for its acknowledgement or an operator review state.
Which signals prove the deletion actually finished?
A green HTTP 202 only proves that a request was accepted. The useful SLO is the age of the oldest non-completed deletion request, split by state. Log state transitions with request ID, subject pseudonym, policy version, actor type, and counts; never log the OTP itself or a raw phone number. Alert on requests stuck in revoking or purging, a growing retry count, and any post-completion authentication attempt that is accepted.
Build a verification job that samples completed requests and checks three invariants: no active session for the subject, no unconsumed OTP challenge, and no addressable profile row in the application database. The checker should use the tombstone ID for joins and should not recreate personal data in its report.
I am not sure your legal retention window matches the technical deletion deadline. Have counsel and the data owner sign off on that boundary, then pin the policy version in the request. Your mileage may vary across regions, but an undocumented exception will become an on-call surprise.
Where does this workflow not fit?
The catch is that synchronous deletion is not suitable when the account owns large objects, cross-region replicas, or records managed by several teams. Use an asynchronous workflow with a visible pending state, and give support a safe status lookup. If a regulated system requires a human approval before erasure, keep the request in blocked until that approval is recorded.
A hard delete is also the wrong choice when financial or abuse investigations require immutable evidence. Retain the narrowest pseudonymous record required by policy, document who can access it, and set an expiry for that retention. Stick with a longer-lived account suspension when the user asks only to stop login or when an active dispute makes immediate erasure unsafe.
Do not promise that revocation erases an already delivered email, an exported report, or a token cached by another system. Promise the control you can enforce: future requests are denied, owned data is removed or anonymized within the stated policy, and every exception has an owner.
Rollout and rollback checks
Ship the workflow behind a feature flag and run it against synthetic subjects first. Test duplicate clicks, two browser sessions, an OTP replay, a worker crash after revocation, a queue redelivery during purge, and a downstream service that responds late. The expected result is one request record, revoked access, and a resumable purge.
Rollback means stopping new purge work, not restoring deleted personal data from an ad-hoc backup. Keep the state machine able to finish in-flight requests, quarantine records that violate an invariant, and page an owner with the request ID. After the flag is disabled, run the verification job until no request remains in an unsafe intermediate state.
A deletion workflow earns trust through evidence: consent intent is attributable, sessions and OTPs are revoked, removal is policy-aware, and retries do not create a second identity event. That is the standard I would apply before calling phone OTP login production-ready.
Top comments (0)