Short answer: offboard a custom domain by treating the DNS zone and tenant binding as one deletion boundary, removing only records owned by that binding, and leaving unrelated tenant records untouched; delay the final cutover until SPF, DKIM, and DMARC evidence has been collected.
This is an architecture decision, not a clever delete query. DNS propagation makes a fast-looking cutover deceptive, while an over-broad zone delete can break another tenant's mail. I optimize for a reversible, auditable change that can survive a reconciliation run days later.
Scope first.
The decision record: zone scope is the safety boundary
The invariant is simple: a record can be deleted only when its normalized owner name, record type, and tenant binding match the offboarding request. The zone is necessary scope, but it is not sufficient scope. A shared zone may contain a customer's apex records, a platform verification TXT record, and a second tenant's delegated hostname.
I keep three identifiers in the offboarding record: tenant_id, the zone name, and a domain ownership token. The token is not a secret substitute for authorization; it is evidence that the request refers to the same binding that was onboarded. Every mutation gets an idempotency key and an append-only audit event with the before-image. That gives reconciliation something concrete to compare with intent.
| Option | What it protects | Where it fails | Use it when |
|---|---|---|---|
| Delete the whole zone | Fast cutover | Destroys unrelated tenant and operational records | A single tenant owns an isolated zone and the contract explicitly permits zone destruction |
| Delete by record name only | Limits the blast radius somewhat | A reused name can still collide across types or tenants | A provider exposes strong per-record ownership and the zone is not shared |
| Delete by zone, tenant binding, name, and type | Preserves other tenants and supports replay | Requires an ownership index and a verification pass | Shared zones, delegated mail hosts, and regulated audit trails |
The third option is the default. The first option is tempting during an urgent offboarding, but it turns propagation delay into a multi-tenant incident. The catch is that a precise delete cannot remove records that the system never registered as its own; those need a manual ownership review, not a wider query.
How should a custom domain offboarding delete records by zone without touching other tenants?
Start with a read phase. Resolve the request to one canonical zone, load the ownership index, and compare the provider's current snapshot with the intended records. Do not infer ownership from a string prefix alone: mail.example.com and mail.example.com. must normalize to the same DNS name, while mail2.example.com must not match it.
Then plan the mutation. The plan is a sorted list of exact record keys (zone, owner, type, tenant_id), each carrying its observed value and version. A repeated request should produce the same plan and a no-op result after the first successful commit. That is the exactly-once mindset applied to an API that may actually deliver at-least-once retries.
Here is the critical path in Go. The provider interface is deliberately generic; the safety property lives in the selection predicate and the audit trail, not in a vendor-specific endpoint.
package offboard
import (
"context"
"fmt"
"sort"
)
type Record struct {
Zone, Name, Type, TenantID, Value, Version string
}
type DNS interface {
List(ctx context.Context, zone string) ([]Record, error)
Delete(ctx context.Context, r Record, idempotencyKey string) error
}
type Audit interface {
Append(ctx context.Context, event string, r Record) error
}
func Offboard(ctx context.Context, dns DNS, audit Audit, tenantID, zone, key string) error {
records, err := dns.List(ctx, zone)
if err != nil {
return err
}
plan := make([]Record, 0, len(records))
for _, r := range records {
if r.Zone == zone && r.TenantID == tenantID {
plan = append(plan, r)
}
}
sort.Slice(plan, func(i, j int) bool {
return fmt.Sprintf("%s|%s|%s", plan[i].Name, plan[i].Type, plan[i].Version) <
fmt.Sprintf("%s|%s|%s", plan[j].Name, plan[j].Type, plan[j].Version)
})
for _, r := range plan {
if err := audit.Append(ctx, "dns.record.delete.planned", r); err != nil {
return err
}
if err := dns.Delete(ctx, r, key); err != nil {
return err
}
if err := audit.Append(ctx, "dns.record.delete.committed", r); err != nil {
return err
}
}
return nil
}
Then stop.
In production, the read and delete phases also need optimistic concurrency: reject a record whose version changed after the plan, then re-read it. Never silently delete a newer value written by another tenant or operator. I also persist a tombstone for each key so a retry can distinguish “already removed” from “never owned.”
SPF, DKIM, and DMARC make propagation part of the contract
Offboarding mail is not complete when an API returns success. SPF is evaluated from TXT policy, DKIM depends on a selector-specific TXT key, and DMARC evaluates alignment and publishes reporting or enforcement policy. RFC 7489 describes DMARC policy discovery and aggregate (rua) and forensic (ruf) reporting; those reports are useful evidence that the old identity is no longer sending, but they arrive on a delayed schedule.
That delay changes the cutover rule. Mark the domain as pending_offboard, stop issuing new credentials, and keep the old records until the observation window has elapsed. During the window, query authoritative and recursive resolvers, record their answers, and compare them with the planned state. A low TTL can improve future changes, but it cannot erase answers already cached under a previous TTL.
Consider a concrete cutover: at 09:00 the zone answers with the old DKIM key, at 09:04 an authoritative query shows the new state, and at 09:08 one mailbox provider still verifies a cached answer. If the offboarding job declares victory at 09:04, the audit trail says “deleted” while the delivery path still trusts the old selector. The safe state machine keeps the binding pending, records both observations, and only closes it after the configured evidence threshold is met; if a later query reveals an unexpected value, reconciliation reopens the case instead of issuing a blind restore.
The subtle failure is deleting a DKIM selector that another tenant still uses because both tenants share the same zone. Selector ownership must therefore be recorded just like an invoice line: exact name, exact type, exact tenant, exact value hash. If a selector is shared by policy, the correct action is to detach the binding and retain the record.
Reconciliation, rollback, and evidence
I separate intent from observed DNS. The intent store says which tenant should own which record; periodic reconciliation reads the provider and emits drift events. A missing record, an unexpected value, and an extra record are different states with different runbooks. Treating all three as “delete and recreate” is how an offboarding job becomes a write storm.
Rollback is also scoped. Restore only the before-images whose tenant binding still exists and whose ownership token has not been reassigned. If the domain has been transferred, restoring the old record would be an unauthorized write. This is where an audit trail earns its keep: an operator can prove what was removed, why, and under which authorization without guessing from current DNS.
There is a valid rejected option: hand the entire zone to a registrar export/import workflow. It is quick for a single-tenant zone, and it may be the least operationally expensive path for a small team. It is not suitable when several tenants share the zone, when DMARC reporting must remain attributable, or when compliance requires record-level deletion evidence. Stick with zone destruction only when isolation is contractual and independently verified.
Your mileage may vary on the observation window. I'm not sure any fixed number is correct across recursive resolvers and mail receivers; the right value comes from the longest TTL you publish, the reporting cadence you observe, and the risk of an old sender remaining active. A 300-second TTL is a configuration choice, not proof of a five-minute cutover.
The decision rule is therefore narrow: delete by canonical zone plus tenant binding, preserve before-images, wait for propagation evidence, and remove only the records the ownership index can prove. That protects other tenants while giving the offboarded customer a defensible end state.
Top comments (0)