DEV Community

DimitriReed2158
DimitriReed2158

Posted on

Healthcare DNS Verification: What Control Actually Proves Before Record Changes

Treat domain verification as evidence that a party could complete one DNS challenge, not as authorisation to manage every record under that name. For a healthtech admin console, keep verification, identity, policy, and publication as separate decisions; then continuously compare approved intent with authoritative DNS. That boundary prevents a valid token from becoming a permanent, overbroad write credential.

TL;DR: A DNS challenge can demonstrate practical control of the place where the verifier asked for a value. It does not identify the human, prove organizational approval, establish legal ownership, or grant an unlimited right to change mail, application, or delegation records. Verification should expire, authorization should be explicit, and every publication should be reconcilable to an approved change.

What does domain verification actually prove about DNS control?

The common proof is straightforward: generate an unpredictable value, ask the claimant to publish it at a specified owner name, and query DNS until the expected value appears. Observing the value supports a narrow conclusion: someone in the claimant's path could cause that response to be published at that time.

That path may include a registrar account, a DNS hosting account, an infrastructure pipeline, a delegated child zone, or an operator who received a one-time request. DNS itself does not tell the verifier which person acted or why. RFC 1034 describes a distributed namespace built around zones and delegation; those mechanisms answer naming questions, not employment, legal ownership, or clinical-system authorization questions.

Scope matters. Proof at _verify.portal.example should not silently become permission to replace MX records at example, edit api.example, or change the zone's name-server delegation. A parent-zone operator may publish a token for a delegated service without owning the service workflow. Conversely, an application team may control a delegated child while having no access to the parent.

The proof is also temporal. Credentials can be revoked, staff can change roles, and delegations can move after a successful check. A token observed last year is historical evidence, not a current access decision.

Nothing more is proved.

Why is a verified badge an unsafe write policy?

In a healthcare environment, the blast radius crosses systems quickly. A request to connect a patient-notification hostname can look local in an admin console while touching records used for web routing or mail authentication. DMARC illustrates the distinction well: RFC 7489 defines a policy and reporting mechanism tied to DNS-published records. The presence of a verification token does not authorize another process to alter that policy.

The dangerous implementation is a single Boolean such as domain_verified=true followed by a generic “edit DNS” capability. It collapses four questions:

  1. Was the challenge satisfied for the exact requested name?
  2. Is the actor authenticated and still associated with the tenant?
  3. Has policy approved this record type, owner name, value, and lifetime?
  4. Does published DNS still match the approved intent?

Those answers age differently. Challenge evidence might remain useful for hours or days. A user session should be shorter. Approval may cover one change only. Publication can drift minutes later because another controller writes to the same zone.

This is where an SRE reflex helps: model the operation as a reconciliation loop, not a form submission. After being paged by missed jobs and duplicate deliveries, I assume every queue message may arrive late or twice. DNS changes deserve the same idempotency rule. The desired record set gets a stable change ID and version; workers compare before writing; retries converge on the same state.

One bit cannot carry that meaning.

Build a narrow authorization envelope

Store challenge evidence separately from the proposed record mutation. An approval should name the tenant, zone, owner name, record type, intended values, actor, approver, expiration, and immutable change ID. If the console supports only a defined subtree, enforce that boundary server-side. Do not infer it from what the UI happened to display.

The trade-off is more state and more review work. A small team that changes one manually administered zone twice a year may reasonably keep a ticket, require two-person review, and use registrar or DNS-host controls instead of building a reconciler. This design earns its complexity when an internal console serves multiple tenants, accepts recurring requests, or delegates publication to queued workers. It also has a hard limitation: reconciliation can detect a mismatch, but it cannot reveal the intent of an out-of-band editor. The operator still needs ownership boundaries and an escalation path.

The following Go sketch keeps the decision explicit. It omits provider-specific transport and treats the publisher as a narrow interface. The important part is that verification is one input, never the authorization result itself.

package dnschange

import (
    "context"
    "errors"
    "strings"
    "time"
)

type Change struct {
    ID           string
    TenantID     string
    Zone         string
    Owner        string
    Type         string
    Values       []string
    VerifiedName string
    VerifiedAt   time.Time
    ApprovedBy   string
    ExpiresAt    time.Time
}

type Publisher interface {
    Current(ctx context.Context, zone, owner, recordType string) ([]string, error)
    Replace(ctx context.Context, changeID, zone, owner, recordType string, values []string) error
}

func Apply(ctx context.Context, now time.Time, c Change, p Publisher) error {
    if c.ID == "" || c.ApprovedBy == "" {
        return errors.New("missing change identity or approval")
    }
    if now.After(c.ExpiresAt) {
        return errors.New("approval expired")
    }
    if c.Owner != c.VerifiedName && !strings.HasSuffix(c.Owner, "."+c.VerifiedName) {
        return errors.New("owner is outside verified namespace")
    }
    if c.Type != "TXT" && c.Type != "CNAME" {
        return errors.New("record type is outside policy")
    }

    current, err := p.Current(ctx, c.Zone, c.Owner, c.Type)
    if err != nil {
        return err
    }
    if equalSet(current, c.Values) {
        return nil
    }
    return p.Replace(ctx, c.ID, c.Zone, c.Owner, c.Type, c.Values)
}

func equalSet(a, b []string) bool {
    if len(a) != len(b) {
        return false
    }
    counts := make(map[string]int, len(a))
    for _, value := range a {
        counts[value]++
    }
    for _, value := range b {
        counts[value]--
    }
    for _, count := range counts {
        if count != 0 {
            return false
        }
    }
    return true
}
Enter fullscreen mode Exit fullscreen mode

That suffix check is illustrative, not sufficient for arbitrary user input. Normalize fully qualified names before policy evaluation, reject malformed names, and account for the zone apex explicitly. The DNS wire format and textual presentation have edge cases; RFC 1035 is the baseline, and a maintained DNS library is preferable to home-grown parsing. I default to rejecting a name the policy layer cannot normalize confidently because a delayed change is easier to recover from than a write outside the approved subtree.

Keep verification tokens single-purpose. Use a dedicated owner name, bind the server-side challenge to the tenant and requested namespace, require an exact value match, set an expiry, and consume or revoke the challenge after success. Query authoritative servers when making the verification decision so a recursive cache does not obscure which version is currently published. Even then, record the observed answer and time as evidence, not as identity.

Reconcile intent against publication

The source of truth should be the approved desired state. Published DNS is observed state. The reconciler compares them and reports three outcomes: converged, pending within the allowed window, or drifted. “Verified” is not one of those publication states.

Use a per-name, per-type comparison. DNS record sets are unordered, so comparing presentation order creates false drift. Normalize values according to the record type, but retain the raw observed response for audit. A changed time-to-live may also matter; decide that in policy instead of accidentally ignoring it.

Concurrency needs a rule. Before replacing a record set, compare the current observation with the value captured during approval. If it changed, stop and require a fresh plan instead of overwriting an operator's newer edit. Where the publication mechanism supports prerequisites, RFC 2136 defines conditional update behavior that can guard against stale assumptions. Otherwise, implement the compare immediately before the write and accept that the transport may not offer an atomic guarantee.

Fail closed.

Track a small set of operational signals: challenge age, verification failures by reason, approval latency, publish attempts by change ID, time to convergence, and drift count by owner and type. Alert on sustained drift and expired approvals still being attempted. A single retry is ordinary; repeated writes for an already-converged change point to broken idempotency or stale observation.

Verify and roll back without creating a second incident

Before rollout, test the negative paths: a token under a sibling name, an expired challenge, a user from another tenant, an unapproved MX change, a duplicate queue delivery, and a record modified between approval and publication. The success test should query authoritative DNS and confirm the complete intended record set, not merely find one expected string among unexpected values.

Deploy the reconciler with writes disabled first. Compare proposed actions with observed records and have the owning team review the diff. Then enable a tightly scoped record type and subtree, with a concurrency limit that respects the publication system. Expanding scope should be a policy change with review, not a side effect of another successful verification.

Rollback means another reviewed desired state, not replaying an old API response. Capture the previous record set when the plan is approved, validate that the currently published set is still the one introduced by the change, and restore only then. If it differs, pause. Someone or something has written after you, and a blind rollback would erase that work.

The runbook decision is concise: accept domain verification as temporary evidence for one namespace; require current identity and explicit policy for each mutation; publish idempotently; and continuously detect drift. That gives the healthtech console a defensible permission boundary without asking DNS to prove something it was never designed to express.

References

Top comments (0)