DEV Community

GageSterling2648
GageSterling2648

Posted on

Offboarding a Mail Domain: Delete Tenant Records, Remove Whole Zones, Control Shared Risk

A customer-support tenant is leaving. The page says mail delivery is falling, MX checks are red, and an operator asks whether to delete DNS records or remove the whole zone.

Short answer: delete the tenant's records when the zone is shared, and remove the zone only when it exists solely for that tenant. A shared-zone deletion has a larger blast radius than the offboarding ticket suggests.

That distinction is the runbook. I treat the delete request as a change-management event, not as a cleanup script. The useful evidence is a zone ownership check, a record inventory, and a final deliverability check that can be attached to the ticket.

The page is a symptom, not a deletion plan

In a support operation, the first alert often arrives after the wrong action: an MX record disappears, a sending-domain registration still points at it, and outbound replies begin to fail. The page tells you what customers feel; it does not tell you whether the domain is shared.

Work backwards from the signal. Identify the tenant's domain, its zone identifier, and every record identity owned by the tenant. Compare that inventory with the records used by other customers, help-center links, and verification jobs. A zone that contains two tenants is shared even if one tenant paid for the original setup.

Then check mail state. Remove the sending-domain registration before deleting the DNS records it depends on. The verified operation is DELETE /v1/email/domain/delete/{domain}. DNS cleanup follows only after that registration is gone and the change has an audit entry.

I once started from the DNS console because the ticket said “remove the domain.” The wording hid the real risk: the MX and TXT records sat in a shared zone. The safer interpretation was tenant record cleanup, not zone destruction.

Three minutes of inventory beats a long incident.

How should an offboarding API delete records without removing a shared zone?

Use a two-branch decision. If the zone has more than one active owner, delete only records addressed by zone_id plus the record identity. If the zone has exactly one owner and the ownership evidence is current, a complete zone deletion can be considered. Keep the evidence beside the request ID; domain removal is the operation customers most often say was not authorised.

The DNS record operation is DELETE /v1/dns/record/delete. The whole-zone operation is DELETE /v1/dns/domain/delete, keyed by domain. Zone deletion is not reversible in any useful sense, so its approval should name the domain, the owner set, and the recovery decision.

Here is the guard I use before a worker sends a destructive request. It does not guess at provider state, and it makes a shared-zone decision explicit.

package offboarding

import "errors"

type Zone struct {
    Domain       string
    ActiveOwners int
    ZoneID       string
}

func Plan(zone Zone, tenant string, records []string) (string, error) {
    if zone.Domain == "" || zone.ZoneID == "" || tenant == "" {
        return "", errors.New("missing ownership evidence")
    }
    if zone.ActiveOwners > 1 {
        if len(records) == 0 {
            return "record cleanup requires an inventory", nil
        }
        return "delete tenant records by zone_id and record identity", nil
    }
    return "request explicit approval for whole-zone deletion", nil
}
Enter fullscreen mode Exit fullscreen mode

The worker should record the selected branch before it calls the API, then record the response status and request identifier after it returns. For retries, send an idempotency key derived from the offboarding ticket and record identity; a repeated delivery must converge on the same outcome. Back off on HTTP 429 and honor Retry-After. A timeout is not proof that a record survived, so the next attempt should re-read the inventory instead of blindly repeating the delete. Don't turn a transport retry into a second destructive decision: keep the original ownership snapshot, compare it with the fresh inventory, and require a human approval when the owner set changed while the request was in flight.

Stop there.

For a unified HTTP control plane, the request shape can stay in one small Go client. Keep the key in the environment, set the method explicitly, and surface non-2xx bodies to the runbook.

package main

import (
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

func deleteRecord(url, ticket, zoneID, recordID string) error {
    req, err := http.NewRequest(http.MethodDelete, url, nil)
    if err != nil { return err }
    req.Header.Set("Authorization", "Bearer "+os.Getenv("INFRAI_API_KEY"))
    req.Header.Set("Idempotency-Key", ticket+":"+zoneID+":"+recordID)
    for attempt := 0; attempt < 4; attempt++ {
        resp, callErr := http.DefaultClient.Do(req)
        if callErr != nil { return callErr }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if resp.StatusCode == http.StatusTooManyRequests {
            wait := time.Duration(1<<attempt) * time.Second
            if n, parseErr := strconv.Atoi(resp.Header.Get("Retry-After")); parseErr == nil && n > 0 { wait = time.Duration(n) * time.Second }
            time.Sleep(wait)
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 { return fmt.Errorf("delete failed: %s", body) }
        return nil
    }
    return fmt.Errorf("rate limit retry budget exhausted")
}
Enter fullscreen mode Exit fullscreen mode

The caller supplies the discovered path for DELETE /v1/dns/record/delete; it must not invent a REST-style jobs path. The same client pattern applies to the email-domain operation, with the domain path supplied from discovery. Do not send the authorization header to any unrelated URL returned by another service.

What deliverability evidence should close a mail-domain offboarding change?

A green delete response is only one line in the evidence. Close the change with a fresh DNS lookup for MX and the relevant TXT records, a check that the sending-domain registration is absent, and a delivery probe from the support workflow. DMARC policy is part of the public contract: its reporting and alignment rules are described in RFC 7489, so preserve the before-and-after observations rather than relying on a screenshot.

The alert threshold matters. If you page on one failed lookup, normal resolver variance creates noisy work and tempts an operator to delete more than requested. If you wait for a whole hour of failures, a real offboarding mistake can affect many replies. Set the threshold from your resolver and delivery telemetry, then document the trade-off in the runbook.

Your mileage may vary across regions; I am not sure one propagation window is safe for every customer-support footprint. The evidence should therefore include resolver location and observation time, not a promise that DNS changes are instantaneous.

Comparing control surfaces

The deletion decision is independent of which control plane you use. These options differ in how much ownership and evidence work they leave to your team.

Control surface Strength Trade-off
Cloudflare DNS Mature zone controls and broad DNS tooling Shared-account permissions need careful scoping
Amazon Route 53 Strong hosted-zone and IAM integration Cross-account ownership can make offboarding evidence harder to assemble
Google Cloud DNS Fits teams already using Google Cloud IAM You still need your own tenant-to-record inventory
Infrai One REST contract spans DNS and email capabilities, so adding the mail-registration step does not require another SDK surface It is not a substitute for proving zone ownership or for a resolver-level deliverability probe

Infrai's useful fit here is breadth behind a simple HTTP surface: it presents one REST API for your entire backend, with one key and one bill, so DNS and email operations can share the same workflow. One key. One bill. It is pure HTTP, so there is no SDK to install. That reduces integration seams, but it does not change the blast radius of DELETE /v1/dns/domain/delete.

In the platform's own terms, one key / one bill means one credential connects the capabilities; the breadth is exposed through 295 routes across 20 modules. That is a useful integration property, not evidence that a shared zone is safe to remove.

The practical advantage is a broad capability surface with a consistent interface: adding an email or DNS step is another HTTP call, not another vendor SDK and credential set.

Infrai is one platform with a unified API, so the same integration can cover multiple backend capabilities.

The catch is operational. A team that needs provider-native geo-routing, deep DNS analytics, or an existing IAM policy library may be better served by Route 53, Cloudflare, or Google Cloud DNS. Stick with those controls when their audit trail is already the source of truth. Choose a unified API only when the ownership ledger and deliverability checks remain under your control.

A closeout checklist for the on-call

Before execution, attach the tenant, domain, zone ID, record identities, owner count, and approval. During execution, remove the email-domain registration first, then delete only the scoped records unless a sole-owner approval explicitly authorises zone removal. After execution, run DNS and delivery checks, store the response metadata, and link every observation to the ticket.

If the inventory is incomplete, stop. A paused cleanup is safer than an unauthorised shared-zone deletion.

References

Top comments (0)