DEV Community

YancySterling6529
YancySterling6529

Posted on

Healthtech DNS Configuration: Intended and Current State Converge Within Ownership Boundaries

Choose the zone owner before trying to make intended state and current state converge. In a healthtech DNS configuration moving away from a registrar-specific API, customer-owned zones should produce instructions and verification results; platform-owned zones may permit an automated writer. Mixing those paths turns an ordinary record mismatch into an ownership incident.

TL;DR: intended state is the versioned DNS snapshot the application wants, while current state is a fresh authoritative observation. Drift is the typed difference between them. A safe reconciler repeatedly observes, plans, applies only records inside its ownership boundary, and verifies from authoritative servers until the allowed subset matches. It never treats a cached recursive answer as proof, and it never replaces an entire customer zone to fix one application record.

How should intended and current DNS configuration state converge?

The data flow is small enough to say plainly. A deployment produces desired records. An observer queries the authoritative DNS path and normalizes the answers. A planner compares matching owner names, types, values, and TTL policy. The executor either writes an approved change to a platform-owned zone or emits a customer action for a customer-owned zone. A later observation confirms the result.

This distinction matters in healthtech because one domain can carry unrelated responsibilities: an application endpoint, patient notifications, staff mail, and third-party verification records. DNS itself does not label a record as belonging to one internal team. The controller therefore needs an explicit scope, such as a set of exact owner-name and record-type pairs. Anything outside that scope is visible context, not writable inventory.

Desired does not mean globally authoritative. It means authoritative for the narrow slice this controller is allowed to manage. For a delegated platform zone such as app.clinic.example, that slice may be broad. For a customer apex such as clinic.example, it is usually a few named records and no more.

The customer-owned versus platform-owned decision should be stored beside the zone, not inferred from which credentials happen to work today. Credentials change during a registrar migration. Ownership should not. This is the part that makes the usual explanation of drift incomplete: convergence is constrained by authority, so a difference can be real and still be off-limits to the writer. In the clinic example, an MX answer may be useful evidence that the observer reached the correct zone, but it cannot become a deletion merely because the intended application snapshot omitted it. The controller has to preserve that asymmetry in its data model, plan output, audit event, and retry behavior. A boolean such as managedZone is too coarse when a customer delegates one subdomain but retains the apex; store ownership at the narrowest boundary the executor can enforce.

Scope first.

A small planner before the trade-offs

The following TypeScript example deliberately stops at planning. Provider adapters can translate the plan later, while customer-owned zones can render the same plan as instructions. Values are normalized by record type before comparison, which prevents harmless presentation differences from looking like drift.

type RecordType = "A" | "AAAA" | "CNAME" | "TXT" | "MX";

type DnsRecord = {
  name: string;
  type: RecordType;
  ttl: number;
  values: string[];
};

type ManagedKey = `${string}|${RecordType}`;
type Change =
  | { kind: "create"; desired: DnsRecord }
  | { kind: "replace"; current: DnsRecord; desired: DnsRecord }
  | { kind: "delete"; current: DnsRecord };

const keyOf = (record: Pick<DnsRecord, "name" | "type">): ManagedKey =>
  `${record.name.toLowerCase().replace(/\.$/, "")}|${record.type}`;

function normalizedValues(record: DnsRecord): string[] {
  const domainLike = record.type === "CNAME" || record.type === "MX";
  return record.values
    .map((value) => {
      const trimmed = value.trim();
      return domainLike ? trimmed.toLowerCase().replace(/\.$/, "") : trimmed;
    })
    .sort();
}

function equivalent(current: DnsRecord, desired: DnsRecord): boolean {
  return current.ttl === desired.ttl &&
    JSON.stringify(normalizedValues(current)) ===
      JSON.stringify(normalizedValues(desired));
}

function planChanges(
  desired: DnsRecord[],
  observed: DnsRecord[],
  managed: Set<ManagedKey>,
): Change[] {
  const wanted = new Map(desired.map((record) => [keyOf(record), record]));
  const found = new Map(observed.map((record) => [keyOf(record), record]));
  const changes: Change[] = [];

  for (const managedKey of managed) {
    const next = wanted.get(managedKey);
    const current = found.get(managedKey);

    if (next && !current) changes.push({ kind: "create", desired: next });
    else if (!next && current) changes.push({ kind: "delete", current });
    else if (next && current && !equivalent(current, next)) {
      changes.push({ kind: "replace", current, desired: next });
    }
  }
  return changes;
}

const managed = new Set<ManagedKey>([
  "portal.clinic.example|CNAME",
  "_dmarc.clinic.example|TXT",
]);

const desired: DnsRecord[] = [
  { name: "portal.clinic.example", type: "CNAME", ttl: 300,
    values: ["edge.health.example."] },
  { name: "_dmarc.clinic.example", type: "TXT", ttl: 3600,
    values: ["v=DMARC1; p=none; rua=mailto:dmarc@clinic.example"] },
];

const observed: DnsRecord[] = [
  { name: "portal.clinic.example.", type: "CNAME", ttl: 600,
    values: ["old-edge.health.example."] },
  { name: "clinic.example", type: "MX", ttl: 3600,
    values: ["10 mail.clinic.example."] },
];

console.log(planChanges(desired, observed, managed));
Enter fullscreen mode Exit fullscreen mode

The unlisted MX record is ignored. That single property is more important than making the adapter clever: the plan cannot delete what it does not own. The missing DMARC record becomes a create action, and the portal mismatch becomes a replacement. On a customer-owned zone those are proposed actions; on a platform-owned zone they can enter an approval and execution path.

The sample compares TTL because this policy declares TTL part of the desired snapshot. Another system may treat TTL as advisory. Pick one rule and test it. Otherwise the controller can issue perpetual updates because its planner and writer disagree about whether TTL differences count. A 300-second desired TTL and a 600-second observed TTL are concrete drift under this policy, even when the CNAME target matches.

That choice has a cost.

Drift is more than a string difference

DNS records are sets, yet APIs often return arrays with unstable ordering. Names may be absolute with a trailing dot or relative without one. Domain-name values are case-insensitive, while arbitrary TXT payloads must not be casually lowercased. MX and SRV values contain multiple fields, so production normalization should parse their wire-level meaning rather than apply one string rule to every type. RFC 1034 and RFC 1035 define the core concepts and record behavior.

There is another trap: current state is a timed observation. Recursive resolvers cache answers according to TTL, and negative answers can also be cached. Querying one recursive resolver immediately after a write can report the old value even when the authoritative data has changed. Verification should query the relevant authoritative servers, retain the observation time and source, and require agreement appropriate to the zone's rollout policy.

Fast loops don't fix stale evidence.

The planner also needs to distinguish absence from failure. NXDOMAIN means the queried name does not exist; NODATA means the name exists but lacks the requested type. A timeout or server failure proves neither. Converting all three into an empty record set creates destructive plans during a transient lookup problem. In that case the correct change count is zero and the observation is marked unhealthy.

DMARC makes the ownership boundary concrete. RFC 7489 places policy in a TXT record at the _dmarc label and defines policy and reporting tags. A DNS controller may manage that exact TXT record without owning the domain's MX records or every other TXT value. Because email policy changes have operational consequences, the desired value should remain reviewable as data rather than hidden inside adapter code.

Convergence without a registrar dependency

A portable controller separates four interfaces: desired-state storage, authoritative observation, planning, and mutation. Only the mutation adapter should know a registrar or DNS host's request model. The desired snapshot stays provider-neutral, with a stable zone identifier, ownership mode, managed keys, revision, and records.

For platform-owned zones, attach the expected prior state or provider revision when the API supports conditional changes. If the state changed after planning, observe again instead of overwriting a newer edit. RFC 2136 supplies a standard DNS UPDATE mechanism with prerequisite tests, although deployments vary in whether and how they expose it. A provider-specific API can offer equivalent concurrency protection; the architectural requirement is compare-before-write, not a particular transport.

For customer-owned zones, convergence is cooperative. Generate an exact diff, show the expected values, and poll authoritative DNS for completion. Do not ask for broad account credentials merely to automate two records. This path is slower and adds support work, but it preserves the customer's administrative boundary and avoids coupling onboarding to one registrar's API. The trade-off is explicit: less automatic control in exchange for a smaller credential and ownership surface.

No shortcut changes that boundary.

DNSSEC adds a related check. Signed zones publish DNSKEY and RRSIG records, while DS records create the parent-side link described by RFC 4034. Moving authoritative service without coordinating that chain can leave validating resolvers unable to accept the zone. Treat delegation and DNSSEC state as migration gates, not ordinary application records for the reconciler to rewrite.

A converged plan is an empty authorized diff, not a claim that the entire zone matches a local file. That definition survives provider moves because it depends on record semantics and ownership, not an API's object shape.

Operating the loop after cutover

Run observation and planning in dry-run mode before enabling writes. Record the desired revision, observed snapshot hash, authoritative server, proposed operations, execution result, and verification result under one correlation ID. Metrics should separate observed drift, pending customer action, rejected mutation, and verification timeout; collapsing them into a single failure counter makes ownership delays look like broken automation.

Set a change budget. A plan that unexpectedly touches many managed keys should require review, as should any delegation or DNSSEC change. Retry reads with bounded backoff, but do not blindly replay a mutation whose outcome is unknown. Observe first. If the desired state is already present, the retry has nothing to do.

Testing needs more than mocked provider responses. Table-driven planner tests should cover record ordering, trailing dots, case rules, missing types, unrelated records, and failed observations. An integration environment should exercise an authoritative server so caching and negative answers are visible. Finally, migration rehearsal should export the old registrar state, translate it into the neutral model, compare it against authoritative observations, and leave the original service available until delegation and signed-zone checks pass.

The operating checklist is brief when written as a decision sequence. Confirm who owns the zone, freeze the managed key set, capture a fresh authoritative snapshot, review the planned diff, apply through the correct ownership path, and verify from authority. Keep the previous desired revision for an explicit rollback, but produce another reviewed plan before applying it. No full-zone replacement should be hiding in that sequence.

This approach does cost more engineering time than calling one registrar API directly. The payoff is controlled scope: customer records remain customer records, platform automation stays useful where it is authorized, and the next provider migration changes an adapter instead of redefining the source of truth.

References

Top comments (0)