DEV Community

callumreed2198
callumreed2198

Posted on

Health Domains and Records: 4 Controls for Services That Depend on Them

A credential that can change customer DNS is a production control plane, even when the API call looks like routine configuration. TL;DR: keep customer-owned zones under customer control by default, delegate only the narrow names the health platform must operate, and make every dependent-service change a reconciled, idempotent workflow. Use a platform-owned zone when the platform truly owns the namespace and its failure domain. One key can simplify authentication, but it must not erase ownership boundaries.

This matters for a healthtech product that lets a clinic serve a patient portal from its own domain. The visible request is small: point a name at the product. The operational reality includes record creation, domain verification, certificate readiness, and possibly mail authentication. Those steps do not become one atomic change because one credential can authorize them.

I have been paged by missed jobs and duplicate deliveries in cron and queue infrastructure. The invariant I carried away is blunt: a scheduled task may run zero times or more than once from the application's point of view, so correctness has to live in persisted desired state and repeatable reconciliation. DNS automation deserves the same treatment.

No shortcut changes that.

How should domains, records, and services that depend on them share access?

The credential answers who may request a change. It does not answer who owns the zone, which records the workflow may touch, or which downstream service is ready. Treating those questions as equivalent creates a wide blast radius: a portal onboarding worker intended to manage portal.clinic.example may gain the practical ability to alter unrelated names in clinic.example.

That coupling is especially uncomfortable around mail. DMARC defines a DNS-published policy and reporting mechanism for message authentication. A record change can therefore influence behavior outside the web endpoint that motivated domain onboarding. The safe architectural conclusion is not that mail is special; it is that every record set needs an explicit owner and purpose, including records consumed by a different service.

I use four controls:

  1. Scope authorization to a zone or delegated child name, then constrain the allowed record names and types in application policy.
  2. Store desired state and a stable operation ID before enqueueing work.
  3. Reconcile observed records against that state; do not model retries as repeated imperative mutations.
  4. Gate dependent-service activation on observed prerequisites, with separate status and audit events for each transition.

One key may enter the control plane, but it should reach these controls before it reaches a provider adapter. The trade-off is deliberate: more policy checks and state transitions create implementation work, while narrower authority and explicit readiness make a failed onboarding diagnosable. After pages for missed and duplicate queue work, I stopped treating successful dispatch as proof of successful completion. The same distinction applies here. A completed credential check proves authorization; it does not prove that the desired record is observable or that a dependent service should be activated.

Choose ownership before designing the workflow

Customer-owned and platform-owned zones solve different problems. For a customer-owned zone, the clinic remains the authority for the surrounding namespace. The platform should accept proof that the requested name is configured, or receive narrowly delegated authority for a child name. This preserves a clean exit path: the customer can redirect or remove the name without transferring its entire DNS estate.

A platform-owned zone is appropriate when the product defines and owns the namespace, such as a generated tenant name beneath a platform domain. Operations are simpler because one team controls policy and lifecycle. The trade-off is equally plain: the customer does not own that public name, so it should not be presented as equivalent to bringing a customer domain.

Decision signal Customer-owned zone Platform-owned zone
Namespace authority Customer Platform
Credential reach Narrow integration or delegation Platform zone scope
Offboarding control Customer redirects or removes records Platform removes tenant mapping
Best fit Branded clinic or patient portal Product-assigned tenant hostname

Do not choose ownership because one credential makes a broader path convenient. Choose it from the namespace contract, then shape credentials and automation around that decision.

Make retries boring

The preventative code path is a reconciler. It accepts desired state, reads observed state through a generic interface, and applies a change only when the two differ. A durable store would also enforce uniqueness for the operation ID; the abbreviated example keeps that boundary visible without pretending an in-memory map is production storage.

package domains

import (
    "context"
    "fmt"
)

type DesiredRecord struct {
    OperationID string
    Zone        string
    Name        string
    Type        string
    Value       string
}

type RecordControl interface {
    Read(ctx context.Context, zone, name, recordType string) (string, error)
    Upsert(ctx context.Context, record DesiredRecord) error
}

type OperationStore interface {
    Claim(ctx context.Context, operationID string) (bool, error)
    Complete(ctx context.Context, operationID string) error
}

func Reconcile(ctx context.Context, api RecordControl, ops OperationStore, want DesiredRecord) error {
    claimed, err := ops.Claim(ctx, want.OperationID)
    if err != nil || !claimed {
        return err
    }

    got, err := api.Read(ctx, want.Zone, want.Name, want.Type)
    if err != nil {
        return fmt.Errorf("read observed record: %w", err)
    }
    if got != want.Value {
        if err := api.Upsert(ctx, want); err != nil {
            return fmt.Errorf("apply desired record: %w", err)
        }
    }
    return ops.Complete(ctx, want.OperationID)
}
Enter fullscreen mode Exit fullscreen mode

The important part is not the interface shape. Persist OperationID before dispatch, make Claim atomic, and leave incomplete work eligible for later reconciliation. Record the requested owner, normalized name, intended value, observed value, attempt count, and terminal state. Alert on age and lack of progress, not on every retry.

Activation belongs after observation. A worker that has submitted a record change has evidence of submission, not evidence that every prerequisite for the patient portal is satisfied. Keep statuses such as requested, observed, and active distinct. Short paths hide incidents.

Two bad states are enough to justify the extra bookkeeping: a job disappears after submission but before a durable completion marker, or a retry arrives after the first attempt already changed the record. The first needs reconciliation from desired state. The second needs an atomic claim plus a comparison with observed state. Neither is fixed by handing the worker a more powerful credential.

Where this pattern stops

A reconciler cannot repair an incorrect ownership decision. It also should not silently manage records outside its declared scope, even if the credential permits the call. Reject that request, preserve the audit evidence, and require an explicit ownership change.

The advice changes for a platform-assigned hostname inside a zone operated solely by the platform. There, broad zone automation may match the actual ownership boundary. Even then, stable operation IDs and observed-state checks still protect queue redelivery and scheduled repair runs.

For customer domains, the runbook should begin with ownership and end with reversal: identify who can change the name, which exact records the workflow owns, how dependent services expose readiness, and how the customer exits. The best one-key design is not the one with the fewest calls. It is the one whose authority, state transitions, and rollback path remain legible during an incident.

Sources

Top comments (0)