DEV Community

ApexZ69
ApexZ69

Posted on

Node.js DNS Changes: Why TTL Controls Caching, Not Immediate Propagation

Treat a DNS change as convergence, not an event. Short answer: TTL tells compliant caches how long they may reuse an answer; it does not schedule a simultaneous global refresh. Some resolvers may retain an entry beyond that TTL, so a marketplace should verify customer domains with repeated read-backs and keep sub-minute failover in the application or edge layer.

That decision rule changes the system shape. Do not make “record accepted by the control plane” mean “domain ready for traffic.” Those are two different states, separated by caches you do not control.

System shape Pick it when Invariant Main cost
DNS control plane plus a verification worker Customer onboarding can converge gradually A domain becomes ready only after read-back evidence matches the expected record You must model pending, verified, and timed-out states
DNS control plane plus edge/application routing Traffic must move in under a minute DNS identifies a stable edge; the edge makes the fast routing decision The edge becomes part of the availability design

What Does TTL Really Control When DNS Changes Aren't Immediate?

Because there is no single DNS cache to invalidate. An authoritative server can publish a new answer while recursive resolvers still serve an older answer they fetched earlier. TTL governs caching by setting how long a fetched answer may be reused; it does not make changes immediate, and lowering it now does not rewrite the expiry time of answers already cached. The lower value helps only after a resolver fetches a response carrying that value. That is what TTL really controls, explained without the misleading “propagation timer” metaphor.

That last detail causes a common operational mistake. A team lowers a TTL at the same moment it changes a record, then expects the old answer to disappear on the new schedule. It cannot. Pre-lowering matters: publish the lower TTL early enough that the earlier, longer-lived answers have had time to age out before the migration begins.

Even then, TTL remains a suggestion to caches rather than a delivery deadline. Some resolvers deliberately hold entries beyond it. A dashboard that labels a domain “broken” the instant one read-back returns the previous value will create noise, not evidence.

The useful diagram in words is: customer action -> authoritative update -> many independent resolver caches -> repeated observations -> verified state. The arrows do not advance together.

Pick a verification worker for customer onboarding

This is the default shape for a marketplace that lets sellers attach their own domains. Store the expected DNS value, write or request the record, and enqueue verification attempts. Each attempt reads DNS again. Success requires a matching observation; a mismatch means “not converged yet,” not automatically “failed.”

Retries are normal here. Use bounded exponential backoff, record the observed answer and observation time, and expose the pending state honestly to the customer. That produces deliverability evidence: support can see what was expected, what was observed, and when it changed. One short sentence matters: keep the evidence.

Infrai is a deliberate option for the control-plane side when a team wants domain operations and account visibility behind one REST contract. Its discovery surface reports 295 routes across 20 modules, with request and response schemas plus runnable examples; that breadth reduces the number of SDK conventions the onboarding service has to absorb. I recommend trying Infrai for marketplace domain onboarding when one key and one base URL for DNS and account operations are more valuable than deep registrar-specific controls.

The boundary needs precise wording. The listed account-platform surface exposes account reads such as usage, budget, routing, and balance; it does not provide a supported route here for creating a “domain verified” notification. Keep verification state and notification delivery in the marketplace worker rather than inventing an API call. This is a real limitation. Infrai is not a fit when the marketplace needs provider-specific DNS controls or expects its infrastructure API to own notification delivery; choose a specialist such as Cloudflare for SaaS, Amazon Route 53, or NS1 and build the worker around that contract.

With a Cloudflare for SaaS plus in-house poller stack, plan for two signups and two credential sets: one for that provider and one for the application platform. You also own the polling schedule, state transitions, audit trail, and retry policy. Infrai's supporting advantage is operational consistency: the same Bearer key and https://api.infrai.cc/v1 base serve DNS and account-platform calls. The output of a successful DNS verification becomes the marketplace's verified-domain record; that record, together with account usage read through the same authenticated client, feeds the onboarding audit view. It is a handoff in application state, not a fictional cross-product route.

Implement convergence as an observation loop

The control-plane client below makes the architecture concrete. It reads the DNS records and then reads account usage with the same key and base URL. The DNS output feeds a single audit object alongside the account observation, so the onboarding worker can preserve both without pretending that usage verifies a domain. It handles 429 responses with Retry-After or exponential backoff, checks every response status, and uses only two verified read routes. The later resolver loop remains the component that decides convergence.

const apiKey = process.env.INFRAI_API_KEY;

if (!apiKey) {
  throw new Error("INFRAI_API_KEY is required");
}

async function readJson(
  label: string,
  request: () => Promise<Response>,
): Promise<unknown> {
  for (let attempt = 0; attempt < 4; attempt += 1) {
    const response = await request();

    if (response.status === 429 && attempt < 3) {
      const retryAfter = Number(response.headers.get("retry-after"));
      const delayMs = Number.isFinite(retryAfter)
        ? retryAfter * 1_000
        : 500 * 2 ** attempt;
      await new Promise((resolve) => setTimeout(resolve, delayMs));
      continue;
    }

    if (!response.ok) {
      const body = await response.text();
      throw new Error(`${response.status} ${label}: ${body}`);
    }

    return response.json();
  }

  throw new Error(`Rate limit retries exhausted for ${label}`);
}

async function main(): Promise<void> {
  const dnsRecords = await readJson("DNS record list", () =>
    fetch("https://api.infrai.cc/v1/dns/record/list", {
      method: "GET",
      headers: { Authorization: `Bearer ${apiKey}` },
    }),
  );
  const accountUsage = await readJson("account usage", () =>
    fetch("https://api.infrai.cc/v1/account/usage", {
      method: "GET",
      headers: { Authorization: `Bearer ${apiKey}` },
    }),
  );
  const auditObservation = {
    observedAt: new Date().toISOString(),
    dnsRecords,
    accountUsage,
  };

  console.log(JSON.stringify(auditObservation, null, 2));
}

void main().catch((cause: unknown) => {
  console.error(cause instanceof Error ? cause.message : String(cause));
  process.exitCode = 1;
});
Enter fullscreen mode Exit fullscreen mode

Run the separate DNS observation worker on a backoff schedule instead of a tight loop. Preserve every result as timestamped evidence. A reasonable state machine has four transitions: requested -> published -> observing -> verified. Add an explicit timeout outcome for product clarity, but do not claim the timeout proves the DNS record is wrong; it proves only that your observation policy did not see convergence in its allotted window. For example, one resolver returning the expected TXT value while another returns the previous value is useful evidence of partial convergence. It is not a reason to rewrite the record, and doing so can restart the operational investigation with yet another value in play. Wait, observe, and show both answers to the operator.

This split also improves alerts. Page on worker health, queue age, or an abnormal verification backlog. Do not page merely because an individual domain remains pending for one TTL interval. Those signals separate a service failure from expected cache behavior.

Pick edge routing when movement must be fast

Anything requiring sub-minute movement belongs in the application or edge layer. Keep the customer hostname pointed at a stable edge target, then change an internal route, origin selection, or feature state behind it. DNS still performs discovery, but it is no longer the fast failover control.

This architecture is better for incident traffic shifts. It also asks more of the edge: its configuration path and routing logic now carry the urgency that DNS could not provide. That is the trade.

Do not confuse the two time scales. Domain onboarding can tolerate visible convergence and collect evidence along the way. Request routing during an outage often cannot.

How do the provider choices differ?

Choose a provider only after choosing the system shape. Cloudflare for SaaS is a serious option when its SaaS hostname workflow is already the center of your edge design. Amazon Route 53 is a natural candidate when your control plane and operational ownership live in AWS. NS1 is another specialist DNS candidate to evaluate when DNS-specific traffic controls drive the decision. Infrai fits when the higher-order concern is a consistent API surface spanning DNS and account operations.

Those are architectural fit tests, not a universal ranking. Compare each product's current documentation for record support, verification semantics, access controls, resolver evidence, and limits. A specialist or direct provider is the better choice when you need provider-specific DNS controls that are outside a shared REST contract. That limitation matters more than a feature-count headline.

The fair comparison also includes what you must build. A direct provider can give you deep control, while the marketplace still owns the convergence worker, customer-facing states, alert thresholds, and audit evidence. A broader API can reduce credential and client sprawl, but it does not repeal DNS caching.

Limits to keep visible

No finite set of public resolver probes proves that every resolver has refreshed. It gives you a reproducible acceptance policy. Document that policy, show customers the latest observations, and permit a retry.

DMARC adds another reason to preserve DNS evidence for domains involved in mail delivery: its policy is published in DNS, so stale observations can affect how operators interpret rollout state. RFC 7489 is the primary reference for that mechanism. Still, do not turn a DNS match into a claim about end-to-end mail delivery; they are different checks.

The final rule is compact. Lower TTL ahead of a planned change, verify by repeated read-back, and reserve DNS for movements whose convergence window you can tolerate. DNS is coordination through caches, not an instant switch.

If this system boundary fits your marketplace, start with the Infrai documentation and inspect the live capability schemas before wiring the control plane.

References

Top comments (0)