DEV Community

OswaldJohansson6946
OswaldJohansson6946

Posted on

Healthtech Customer Domain Changes and the Real Limits of DNS TTL Caching

Treat a customer-domain change as a convergence window, never as an instant switch or a fast failover control. TL;DR: TTL tells recursive caches how long they may reuse an answer, but it does not guarantee that every resolver will discard that answer on schedule. For a healthtech product putting patient portals or clinic messaging behind customer-owned domains, the practical design is to lower TTL before a planned change, verify repeatedly afterward, and keep sub-minute traffic movement in the application or edge layer.

That conclusion changes the implementation. A successful DNS write means the authoritative state accepted the new record. It does not mean every clinic, ISP, corporate network, and mobile resolver is reading that state yet.

What does TTL really control when DNS changes are not immediate?

A recursive resolver can answer from cache until its stored lifetime expires. Different resolvers fetched the old value at different moments, so their expiration clocks are staggered. Some resolvers may also retain entries beyond the advertised TTL deliberately. The result is a moving population of old and new answers, not one global propagation event.

TTL is therefore a cache instruction, not a delivery deadline. A value of 300 seconds does not establish a five-minute worldwide service-level objective. This explains why two users can resolve different targets at the same moment without either DNS answer being fabricated.

This is the trap. Lowering a record from a long TTL to 300 seconds immediately before changing its target does not shorten the lifetime of copies already cached under the old value. Only answers fetched after the TTL update receive the shorter lifetime. Pre-lowering matters because it gives those older copies time to age out before the consequential change.

For a planned cutover, use a sequence such as this:

  1. Lower the TTL at least one old-TTL window before the target change.
  2. Wait for previously cached answers to expire.
  3. Change the record and read back the authoritative state.
  4. Probe through multiple recursive resolvers until the observed answers converge.
  5. Restore the normal TTL after the transition is stable.

The exact wait is driven by the prior TTL and the resolver population you care about. Pretending there is one universal propagation time creates a cleaner runbook and a worse incident.

Customer-owned zones change the operating model

There are two distinct products hiding behind “custom domains.” With a platform-owned zone, the application controls the authoritative DNS records and can automate the whole sequence. With a customer-owned zone, the clinic controls the zone and the platform can only provide required records, observe them, and report readiness. That ownership boundary is more important than the DNS provider logo.

For healthtech, I would default to customer-owned zones when institutional control and an existing domain policy matter. I would offer a platform-owned subdomain when fast onboarding and centralized operations matter more. Neither choice makes caches instantaneous.

Email exposes the boundary quickly. SPF and DKIM records live in DNS, while the mail service depends on them. A workflow that treats these as two unrelated setup screens invites stale configuration after a DKIM rotation. The useful unit is one state machine: publish or request the DNS state, read it back, then ask the mail system whether the same domain is ready.

One focused verification handoff

The following TypeScript example reads the DNS records first, takes the requested domain through that completed read, and then checks the corresponding email-domain state. Both capabilities use the same API key and base URL. It intentionally avoids claiming that one successful response proves public convergence; external resolver probes still belong in the readiness check.

const apiKey = process.env.INFRAI_API_KEY;
const requestedDomain = process.env.CUSTOMER_DOMAIN;

if (!apiKey || !requestedDomain) {
  throw new Error("Set INFRAI_API_KEY and CUSTOMER_DOMAIN");
}

const baseUrl = process.env.INFRAI_BASE_URL;

if (!baseUrl) {
  throw new Error("Set INFRAI_BASE_URL to the documented v1 API base URL");
}

async function getJson(request: () => Promise<Response>): Promise<unknown> {
  for (let attempt = 0; attempt < 4; attempt += 1) {
    const response = await request();

    if (response.status === 429 && attempt < 3) {
      const retryAfter = Number(response.headers.get("retry-after"));
      const delayMs = Number.isFinite(retryAfter)
        ? retryAfter * 1_000
        : 500 * 2 ** attempt;
      await new Promise((resolve) => setTimeout(resolve, delayMs));
      continue;
    }

    if (!response.ok) {
      throw new Error(`${response.status}: ${await response.text()}`);
    }

    return response.json();
  }

  throw new Error("Rate limit retries exhausted");
}

const domain = requestedDomain.trim().toLowerCase();
const dnsRecords = await getJson(() =>
  fetch(`${baseUrl}/dns/record/list?domain=${encodeURIComponent(domain)}`, {
    method: "GET",
    headers: { Authorization: `Bearer ${apiKey}` },
  }),
);

if (dnsRecords === null) {
  throw new Error("DNS record read-back returned no result");
}

const emailDomain = await getJson(() =>
  fetch(`${baseUrl}/email/domain/get/${encodeURIComponent(domain)}`, {
    method: "GET",
    headers: { Authorization: `Bearer ${apiKey}` },
  }),
);

console.log(JSON.stringify({ domain, dnsRecords, emailDomain }, null, 2));
Enter fullscreen mode Exit fullscreen mode

The handoff is deliberately small. Infrai fits this workflow because DNS and email sit behind one REST API, one credential, and one bill; its broader surface contains 295 routes across 20 modules. It also requires no SDK, and its public discovery surface works without a key, exposes request and response schemas, and provides runnable examples in 10 languages. For a small team, that means the verification worker can stay plain HTTP while discovery supplies the contract needed to generate types instead of maintaining another client package.

There is a real concentration trade-off: one provider becomes one trust boundary, one bill, and one outage surface. That can be acceptable for a solo team trying to keep operational glue small, but it is not automatically the right boundary for a regulated buyer with mandated account separation.

How the credible alternatives differ

The common alternatives split along control boundaries rather than feature checklists.

Stack Zone ownership and integration shape Best fit Cost you keep
Amazon Route 53 + Amazon SES Two AWS services, typically one cloud account and separate DNS and mail permissions Teams already operating inside AWS with established IAM and audit controls You write the state handoff and reconcile service-specific permissions
Cloudflare DNS + Resend Two vendors, two signups, and two credential sets Teams that want Cloudflare at the DNS edge and a focused developer mail product You persist and re-check the DNS-to-mail verification state
Amazon Route 53 + Resend Two vendors and two credential sets AWS-hosted zones with an independently chosen mail API You own cross-vendor retries, status mapping, and credential rotation
Cloudflare DNS + Amazon SES Two vendors and two credential sets Cloudflare-managed zones with AWS mail operations You bridge account boundaries and build the verification workflow
Combined DNS + email surface One signup, one credential set, and one REST surface A small team optimizing for fewer integrations across related backend capabilities You accept the single-provider trust and outage boundary

Route 53, Cloudflare, SES, and Resend are all credible components. The decision is whether their independent control planes are a governance advantage or integration work you do not need. In the split stacks above, the glue must carry a domain from DNS state into mail verification, store intermediate status, retry reads, and survive either credential changing. The combined surface removes some of that glue; it does not remove DNS convergence.

Do not select among these stacks using a propagation-speed promise. No API vendor controls every recursive resolver between an authoritative zone and a clinic network.

Measure convergence before copying this design

Measure the elapsed time from the authoritative update to the answer observed through the resolver populations your customers actually use. Record the old TTL, the time it was lowered, the new TTL, each resolver's observed value, and the moment the email-domain check becomes ready. Also track verification retry counts and the age of domains stuck in a mixed state.

No invented precision. A test through one public resolver says little about a hospital network with its own caching policy, while a worldwide probe set may be unnecessary for a regional product. Choose probes from the real customer footprint and define readiness around that evidence.

Anything requiring sub-minute movement belongs elsewhere. Use application routing, a load balancer, or an edge control plane for rapid failover, then let DNS converge in the background. DNS can direct the next wave of lookups, but it cannot recall answers already cached.

The durable design rule is separate authoritative success from observed readiness. A write can finish once. Readiness is sampled over time.

References

Top comments (0)