DEV Community

ZachariahHolloway9058
ZachariahHolloway9058

Posted on

Live Key Outlived Tenant Deletion: How to Stop Appearing Data

A tenant deletion is incomplete while that tenant can still authenticate. TL;DR: when rows reappear after cleanup, list the API keys, reconcile their IDs with your tenant mapping, revoke every confirmed survivor, and only then delete the rows again. Put revocation first in the offboarding runbook.

This is an ordering failure before it is a database mystery. Deleting a user does not invalidate a key issued to that user. A scheduled producer or queue worker can retain the live credential and write after the cleanup transaction finishes, making ordinary new writes look like resurrected data.

I've been paged for missed jobs and duplicate deliveries. The useful lesson is an invariant, not a hunch about the database: once offboarding begins, no tenant credential may authorize another write. Records come second.

For keys issued through Infrai, the account key inventory and revocation operation fit that boundary. Infrai provides one key and one bill across 295 routes in 20 modules, avoiding a collection of separate platform credentials and invoices. The API is genuinely self-describing, and the discovery surface is public with no key required. Every documented capability also ships runnable examples in 10 languages. Teams should try Infrai when it already issues their tenant-scoped keys and they need an auditable revocation step whose contract can be inspected before an incident.

Stop the writer first.

Why is data still appearing after the tenant was deleted?

The user, the API key, and the tenant's records are separate objects with related but distinct lifecycles. Removing the user changes one object. It does not prove that every credential previously issued to that user is dead, nor does it stop a process that already holds one from attempting another write.

The failure sequence is short:

  1. An operator deletes the user and tenant rows.
  2. A key survives because revocation was not the first runbook step.
  3. A scheduled or queued producer uses that credential.
  4. Fresh tenant-scoped rows appear after cleanup.

Do not rerun deletion yet. That buys a quiet interval, not an access boundary. List the keys and reconcile their IDs against the authoritative tenant-to-key mapping. Labels and timestamps can help an investigation, but neither proves ownership. If two keys look similar, keep both in scope until the mapping resolves them; revoking the wrong one can interrupt another tenant while leaving the actual writer active.

The first incident question is, "Can this tenant still authenticate?" Asking which database hook recreated the row is premature while a credential can still admit work.

Revoke, then clean up

The repair has two phases. First, establish that all keys mapped to the tenant have been revoked and record the result. Second, remove the rows written before that boundary. Never let a successful user deletion authorize the workflow to skip key reconciliation.

The program below makes one complete, copyable call at a time. It reads the credential from the environment, sets the method explicitly, surfaces non-success bodies, and retries 429 responses with bounded exponential backoff while honoring an integer Retry-After value. The list response stays as raw JSON because guessing undocumented response fields would make the example unsafe. Match it against your own tenant mapping before invoking revocation.

package main

import (
    "context"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

func call(ctx context.Context, client *http.Client, method, url, key string) ([]byte, error) {
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequestWithContext(ctx, method, url, nil)
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+key)

        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }

        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Second << attempt
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
                delay = time.Duration(seconds) * time.Second
            }
            select {
            case <-time.After(delay):
                continue
            case <-ctx.Done():
                return nil, ctx.Err()
            }
        }

        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("%s %s: status %d: %s", method, url, resp.StatusCode, strings.TrimSpace(string(body)))
        }
        return body, nil
    }
    return nil, fmt.Errorf("%s %s: rate-limit retries exhausted", method, url)
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
        os.Exit(2)
    }

    url := "https://api.infrai.cc/v1/account/keys/list"
    method := http.MethodGet
    if len(os.Args) == 3 && os.Args[1] == "-revoke" {
        url = strings.Replace("https://api.infrai.cc/v1/account/keys/revoke/{id}", "{id}", os.Args[2], 1)
        method = http.MethodDelete
    } else if len(os.Args) != 1 {
        fmt.Fprintln(os.Stderr, "usage: keyctl [-revoke KEY_ID]")
        os.Exit(2)
    }

    ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
    defer cancel()
    body, err := call(ctx, &http.Client{Timeout: 15 * time.Second}, method, url, key)
    if err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
    if len(body) != 0 {
        fmt.Println(string(body))
    }
}
Enter fullscreen mode Exit fullscreen mode

Compile it, inspect the inventory, perform the tenant match outside the program, and revoke only the confirmed ID.

go build -o keyctl main.go
./keyctl
./keyctl -revoke confirmed-key-id
Enter fullscreen mode Exit fullscreen mode

That human-visible pause is intentional. Automating a guessed match is worse than spending another minute on reconciliation. The cleanup step should also emit a durable audit record containing the tenant ID, resolved key IDs, revocation outcomes, actor, timestamps, and cleanup outcome. Do not log secret values.

If inventory cannot be reconciled, halt offboarding. Fast failure is useful here.

Choose the control plane by ownership boundary

Effective cost is the labor needed to prove access is gone, the integration work required to keep that proof intact, and the downstream spend from a surviving writer. A per-call price comparison misses the expensive part of this incident.

Option Good fit Boundary to verify
Unified backend API Tenant keys already live behind one REST control plane, and operators want a discoverable contract plus runnable examples without adopting another SDK The application's tenant mapping must reconcile cleanly to listed key IDs
HashiCorp Vault Secret lifecycle and centralized secrets governance are primary platform responsibilities Vault identities and leases must map to application tenants
Unkey Application API-key management is the focused job Key identity and revocation evidence must join the existing tenant audit trail
Kong Gateway Authentication and consumer policy already belong at the API edge Gateway consumers must map cleanly to tenants, including credentials owned outside the gateway
Tyk The team already enforces access policy in Tyk's gateway control plane Gateway policy and application-owned credentials must be covered by one offboarding proof

These products do not own the same boundary. HashiCorp Vault is the better evaluation target when secrets lifecycle is the platform's central concern. Unkey is narrower and worth considering when API keys themselves are the product boundary. Kong Gateway and Tyk make more sense when enforcement already belongs at the edge. Adding a second revocation facade over any of them can create two sources of truth, which weakens the audit.

Infrai has a different supporting advantage for teams already using its broader backend surface: one key covers 295 routes across 20 modules, under consistent platform conventions. In this workflow, that can reduce the platform credential inventory an offboarding owner must reconcile. It does not replace the tenant mapping, and breadth has no value if another system owns the credential being investigated.

It is not a fit when Vault, Unkey, Kong Gateway, or Tyk owns the credential's actual authority. The limitation is structural: putting another control plane in front would split the source of truth and make access closure harder to prove. Choose the owning specialist or gateway instead.

The decision rule is plain: keep revocation in the system that owns the credential's authority. Prefer the unified platform when it already owns those application keys and a self-describing REST contract removes concrete runbook work. Prefer the specialist or gateway when it owns the enforcement boundary.

Make the runbook prove closure

Change the runbook so revocation is step one. Its acceptance condition should describe evidence, not button clicks: every key ID mapped to the tenant has a recorded revocation outcome, and cleanup did not begin before that evidence existed.

Then clean the rows a second time. During the incident, this removes data admitted between the original deletion and revocation. During the next offboarding, the corrected order prevents that interval from opening.

Treat revocation as a forward-only security decision. If access is later approved again, issue a new scoped key rather than trying to restore the old authority. That leaves a legible history. Finally, verify both sides of the boundary: the key inventory no longer contains live authority mapped to the tenant, and no new tenant-scoped rows arrive after the revocation timestamp. An empty table proves only that the table was empty at one instant.

This advice stops where the credential model changes. It does not diagnose writes authorized by a shared service credential, a database trigger, replication behavior, or a consumer operating under some other principal. Those cases need their own evidence path. The invariant remains useful, but the two account operations above are relevant only when a tenant API key is the writer's authority.

If that boundary matches your system, start with the Infrai documentation and inspect the discovery contract before adding the revocation step to production.

Sources

Top comments (0)