TL;DR: Treat a DNS change as a typed publication with three identities: environment, customer, and zone. Resolve all three from the request, compare them with independently loaded expectations, and stop before any write when they disagree. If staging records appeared in a production zone, the useful question is not whether the record value looked valid. It is how a shared configuration value crossed the boundary without proving which zone it named.
For a customer-support product, this matters at the exact moment a customer points help.customer.example at the service. A syntactically valid record can still be published to the wrong place. The smallest defensible fix is an explicit plan step, followed by a checked apply step.
How did staging DNS records reach the wrong production zone?
The record and the destination are separate inputs. Teams often inspect the host, type, and value, then treat zoneId as plumbing. Shared configuration makes that assumption dangerous: the same key can be present in staging and production while referring to different destinations. Nothing about a plausible identifier proves its environment.
This is an identity bug. Frame it that way. The intended tuple is (environment, customer, zone), while the published tuple came from a record request combined with a shared zone value. Debugging starts by printing both tuples from the plan, without credentials, and finding the first point where their sources diverge.
Do not begin by editing the record. First freeze automated writes, retain the requested hostname and change identifier, and inspect the plan that selected the zone. Then verify the destination through an independent environment-specific mapping. If the writer and the verifier read the same shared key, the check is theater.
That last detail bites.
Build the smallest checked publisher
A narrow interface keeps this testable and avoids dragging provider configuration through the application. The request names intent. The resolver supplies the target. A separate policy supplies the expected identity. The publisher writes only after equality has been established.
type Environment = "staging" | "production";
type DomainRequest = {
environment: Environment;
customerId: string;
hostname: string;
target: string;
};
type Zone = {
id: string;
customerId: string;
environment: Environment;
};
type DnsWriter = {
upsertCname(zoneId: string, hostname: string, target: string): Promise<void>;
};
function assertPublicationIdentity(request: DomainRequest, zone: Zone): void {
const mismatches = [
request.environment !== zone.environment && "environment",
request.customerId !== zone.customerId && "customer"
].filter(Boolean);
if (mismatches.length > 0) {
throw new Error(`DNS publication identity mismatch: ${mismatches.join(", ")}`);
}
}
async function publishCustomerDomain(
request: DomainRequest,
zone: Zone,
writer: DnsWriter
): Promise<void> {
assertPublicationIdentity(request, zone);
await writer.upsertCname(zone.id, request.hostname, request.target);
}
This code deliberately does less than a full DNS integration. It establishes the boundary that matters. The caller may resolve a zone from a database, a repository-owned map, or another source, but it must return identity metadata alongside the opaque ID. Passing a bare string erases the evidence needed for the check.
I would also make the plan printable and the apply operation consume that exact plan. This is a trade-off: it adds one data structure and a little ceremony, but it removes hidden configuration reads between review and execution. Config bloat is still config bloat. One immutable plan is enough.
type PublicationPlan = {
request: DomainRequest;
zone: Zone;
};
function planSummary(plan: PublicationPlan): Record<string, string> {
return {
environment: plan.request.environment,
customerId: plan.request.customerId,
hostname: plan.request.hostname,
zoneId: plan.zone.id
};
}
The summary exposes routing decisions, not secrets. Log it with a stable change identifier so an operator can connect intent, approval, and publication without reconstructing state from scattered configuration. Avoid logging credentials or unrelated customer data.
Test the boundary, not the happy path
The high-value test is the one that reproduces the class of failure: a staging request paired with a production zone. It should prove that no writer call occurs. A second mismatch on customer identity prevents one tenant's custom-domain request from borrowing another tenant's zone.
class RecordingWriter implements DnsWriter {
calls: Array<{ zoneId: string; hostname: string; target: string }> = [];
async upsertCname(zoneId: string, hostname: string, target: string): Promise<void> {
this.calls.push({ zoneId, hostname, target });
}
}
const request: DomainRequest = {
environment: "staging",
customerId: "support-team-17",
hostname: "help.customer.example",
target: "tenant.staging.service.example"
};
const wrongZone: Zone = {
id: "zone-production-42",
customerId: "support-team-17",
environment: "production"
};
const writer = new RecordingWriter();
await publishCustomerDomain(request, wrongZone, writer).catch(() => undefined);
if (writer.calls.length !== 0) throw new Error("unexpected DNS write");
Also test an unknown environment, a missing mapping, and a customer mismatch at the configuration boundary. The desired behavior is boring: reject before network I/O. A warning is inadequate because the next line can still publish.
For deployed systems, compare the approved plan with the observed record after publication. That detects drift between intent and published state, while the preflight identity check prevents the specific wrong-zone write. These controls answer different questions and should not be collapsed into one vague DNS validation flag.
What changes at larger scale
At low volume, an environment-owned map plus the typed guard may be sufficient. As customer-domain count grows, move the mapping into a source with explicit ownership, review history, and uniqueness constraints, but keep the runtime contract unchanged because more storage machinery should not leak into every call site. Separate planning from applying: planning can run on each proposed customer-domain change and emit the four useful fields, environment, customer, hostname, and zone ID, while applying accepts that approved plan rather than resolving shared configuration again. This narrows the interval in which intent can drift. Recovery deserves its own operation as well. Removing an accidental record is a new DNS change with a reviewable target, not an ad hoc inverse of the original request; before removal, confirm that the record still matches the accidental value, or a cleanup job could overwrite a later legitimate change. There is an email boundary too. If the customer-support domain participates in mail handling, DNS changes can intersect with domain-level email policy. DMARC is defined in RFC 7489. Treat those records as separately owned policy inputs rather than casually bundling them into the custom-hostname publisher. One deploy should not gain authority merely because both concerns live in DNS.
The trade-off I would keep: the three identity fields are intentionally redundant. The zone ID may already imply a destination inside a provider, yet the application should not have to infer business intent from an opaque token. Carrying environment and customer identity next to it creates a cheap comparison at the write boundary.
The cost is extra mapping data and failed deployments when metadata is stale. That is preferable to publishing a valid staging record in a production zone. Measure the workflow by time to a trustworthy first call, not time to an unchecked first call. The fastest path is a small plan, one fatal identity guard, and a writer that never consults shared ambient configuration.
Top comments (0)