DEV Community

LyraP22
LyraP22

Posted on

Domain Exit Controls — Shared-Zone Record API Safety for Game Onboarding

Short answer: for a game studio retiring a verification domain, delete only the records that its onboarding ownership proof created; remove a whole zone only after proving that no other tenant, service, or delegated child depends on it.

The deciding constraint is drift: the request may say one domain is leaving while the published zone still carries records with several owners. DNS is shared infrastructure more often than a product screen suggests. A single zone can carry player-support mail, a download hostname, campaign redirects, and domain-verification records for separate teams. Removing the zone because one onboarding record is obsolete turns an identity cleanup into an outage-shaped risk.

The practical choice is to treat offboarding as a scoped reconciliation operation, not a delete button. It is a narrow distinction, but it changes the data model, review path, and rollback story.

What should a domain offboarding API delete: records or a whole shared zone?

Use the zone as an administrative container, not proof of exclusive ownership. A deletion plan should begin with the domain being offboarded, expand to the exact record-set identities it owns, then check whether the zone itself has independent records, delegations, or attached services. If any of those checks finds another owner, the plan deletes records only.

A record-set identity is more precise than a hostname string: name, type, and routing policy together identify what is being changed. For example, a verification TXT value at _game-proof.launch.example might belong to the onboarding workflow, while an MX record at launch.example belongs to support mail. They are in the same zone and plainly should not share a lifecycle.

The catch is that a record inventory alone cannot establish intent. A useful system stores an ownership manifest when onboarding creates a record: tenant or game ID, requested domain, record-set key, creation timestamp, approval reference, and an expiry or review date. At offboarding time, compare that manifest with a fresh authoritative-zone read. The manifest answers what the workflow meant to create; DNS answers what is currently published. Neither is enough on its own.

Keep the plan explicit. A reviewer or an automation rule should be able to say why every candidate is in scope and why every nearby record is excluded. This is a permissions problem before it is an API problem.

Published item Manifest owner Offboarding action Reason
_game-proof.launch.example TXT game-418/onboarding Delete record set Exact manifest key and expected value
launch.example MX support-platform Preserve Different owner and service purpose
cdn.launch.example CNAME release-team Preserve Not created by onboarding
launch.example zone multiple owners Preserve The zone is not exclusively owned

The two-source comparison also catches a subtle case: someone may have replaced a verification value manually after onboarding. A plan that deletes by name and type alone can remove the replacement. Compare the intended value as well, and escalate a mismatch for review instead of guessing.

Stop there.

The experiment: build a plan before a destructive mutation

The tempting implementation is simple: receive a domain-exit event, list matching records, and issue deletions. It feels efficient because there is one event and a small number of calls. It fails because the list query describes proximity, not ownership. A zone's apex and its descendants are a shared namespace; matching a suffix is not authorization.

The chosen approach has two phases. First, calculate a dry-run plan from an immutable manifest and a current zone snapshot. Second, execute only the approved plan using conditional checks where the DNS control plane provides them. Persist the snapshot identifier, plan identifier, and each result so that a retry can distinguish an already-applied deletion from a new request.

Consider a game whose onboarding request named arena.example, while its current zone snapshot contains the expected _game-proof.arena.example TXT record, an MX record installed by the support team, and a CNAME used by the release pipeline. The manifest owns only the TXT record. The plan therefore proposes one deletion, two explicit preservations, and zero zone actions. Before execution, it records the value it expects to delete, the snapshot used to make the decision, and the reviewer who approved it. If the TXT value has changed by execution time, the system does not reinterpret that change as permission. It pauses the item for review, because a new value can mean a later onboarding attempt or a separate verification flow. Once the approved TXT deletion is applied, the verifier checks the authoritative source again and logs the observed absence with the original plan ID. A retry sees that completed result instead of attempting a broader cleanup. This extra bookkeeping is ordinary operational work, but it prevents a stale event from becoming authority over DNS that another team now owns.

Here is a focused TypeScript example. The repository deliberately has no deleteZone method in the normal offboarding path; making whole-zone removal a separate privileged operation prevents a broad action from hiding inside a record cleanup.

type RecordSet = {
  name: string;
  type: "TXT" | "CNAME" | "MX";
  values: string[];
  owner: string;
};

type ManifestEntry = {
  recordKey: string;
  expectedValues: string[];
  owner: string;
};

const keyFor = (record: Pick<RecordSet, "name" | "type">) =>
  `${record.name}|${record.type}`;

export function makeOffboardingPlan(
  records: RecordSet[],
  manifest: ManifestEntry[],
  leavingOwner: string,
) {
  const intended = new Map(
    manifest
      .filter((entry) => entry.owner === leavingOwner)
      .map((entry) => [entry.recordKey, entry]),
  );

  return records.flatMap((record) => {
    const entry = intended.get(keyFor(record));
    const valuesMatch = entry?.expectedValues.join(",") === record.values.join(",");

    if (record.owner === leavingOwner && entry && valuesMatch) {
      return [{ action: "delete-record-set" as const, record }];
    }

    if (record.owner === leavingOwner) {
      return [{ action: "review" as const, record, reason: "manifest mismatch" }];
    }

    return [];
  });
}
Enter fullscreen mode Exit fullscreen mode

This example is intentionally conservative. Array ordering may not be meaningful for every record type, so a production comparator should normalize values according to the control plane's record-set semantics before it makes the equality decision. I’m not sure a cross-provider abstraction can safely hide every policy variant; the evidence needed to resolve that is documented conditional-update behavior and a test zone that exercises it.

A zone removal deserves a different approval path with stronger predicates: no records outside the departing owner's approved manifest, no child-zone delegation, no active DNSSEC material managed by another lifecycle, and no attached service declaration. The exact inventory varies by control plane, but the decision rule does not. If the system cannot prove exclusivity, it preserves the zone.

Small batches help.

Delete one record set, verify the new authoritative state, then continue. For a large game portfolio, batch size should be tuned against control-plane rate limits and the cost of a paused rollout, not against the apparent convenience of one bulk request.

Verification records need a separate retention rule

Domain ownership proof is commonly represented by a TXT record, and DMARC itself uses DNS TXT records at a well-defined name. RFC 7489 makes that dependency concrete: receivers discover a policy at _dmarc under the organizational domain. The standard is a useful reminder that a record that looks like disposable text may participate in mail policy, reporting, or another security workflow.

For onboarding proof, record deletion should be gated on a terminal lifecycle event and a retention decision. A game studio might retain an audit reference after removing the TXT record, while retaining neither the token nor the old domain value in application logs. The point is traceability without keeping credentials or stale proof material longer than necessary.

Also check negative caching behavior. Resolvers can cache an absence as well as an answer, so a verification retry immediately after a change may not observe the expected state everywhere. Measure the authoritative result separately from resolver observations, record TTLs in the plan, and make the product workflow honest about the wait. Don't promise instant global convergence.

What to measure before copying this pattern

Measure the share of offboarding plans that resolve automatically, the count escalated for manifest drift, the time from approved plan to authoritative verification, and the number of protected records excluded from each plan. Those metrics expose the real trade-off: a strict manifest reduces accidental deletion but increases review work when teams create DNS records outside the onboarding system.

Track reversibility too. A deletion log with the prior record-set value, actor, approval reference, and verification outcome lets an operator restore a record through the same controlled workflow if business intent changes. It should not become a hidden bypass around access controls.

This method is not suitable when the organization has no reliable record ownership data and cannot establish it during a review window. In that case, stick with a manual, dual-reviewed change process until future onboarding writes the manifest. Automation without an ownership boundary only makes uncertainty faster.

The final decision is deliberately boring: delete an owned record set when intent and published state agree; preserve the zone unless exclusive ownership is proven. That policy keeps a departing game's domain cleanup from changing unrelated player-facing infrastructure.

References

Top comments (0)